Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

4.2 KiB

CODE_REVIEW: add-loop-templates-onboarding

Reviewed Files

  1. scripts/loop-runner.py -- _resolve_prompt function (lines ~248-303), _invoke_harness signature change (lines ~306-338), cmd_tick call site updates (lines ~626, ~638, ~669)
  2. prompts/loop-implement.md -- new file
  3. prompts/loop-verifier.md -- new file
  4. prompts/loop-orchestrate.md -- new file
  5. templates/loops/ci-triage/loop.json -- roles updated
  6. templates/loops/self-improvement/loop.json -- new file
  7. tests/test_loop_templates.py -- new test file (18 tests)
  8. tests/test_loop_runner.py -- prompt ref renames
  9. tests/test_blast_radius.py -- prompt ref renames
  10. tests/test_goal_mode.py -- prompt ref renames
  11. tests/test_framework_self_consistency.py -- exclusion set update
  12. README.md -- Loop Engineering onboarding section
  13. CHANGELOG.md -- task 6 entry
  14. design/loops/technical.md -- section 8 prompt resolution docs

Findings

1. _resolve_prompt -- token substitution correctness

The function correctly handles the two-stage search (loop-local then framework), reads the file, substitutes tokens, writes to outputs/, and returns the temp path. The fallback to raw prompt_ref when the file is not found preserves backward compatibility.

Concern: token injection. The str(value) substitution via content.replace("{" + key + "}", str(value)) is safe for the current token set (all values are controlled: task_brief from SPEC.md, acceptance_criteria from loop.json, etc.). No user-supplied input flows into these tokens without being read from a file first. Acceptable for v1.

Verdict: PASS

2. _invoke_harness signature change

The new loop_path and tick_num parameters are optional with defaults (None and 0). Existing callers that don't pass them get the old behavior (raw prompt_ref passed through). This is backward compatible.

Verdict: PASS

3. cmd_tick call site updates

All three call sites (implement, verify, orchestrate) now pass loop_path=loop_path and tick_num=tick_num where tick_num is computed once as state.get('iteration_count', 0) + 1. This is correct -- the tick number should be consistent across all three role invocations in the same tick.

Verdict: PASS

4. Prompt file content

  • loop-implement.md: has all required tokens, ALLOWED/FORBIDDEN sections, no auto-approve. Correct.
  • loop-verifier.md: has strict JSON output format, score rubric, artifact_content token. Correct.
  • loop-orchestrate.md: has verdict token, phase transition logic, no-edit rule. Correct.

All three prompts are excluded from the self-consistency stop-condition check since they are role prompts, not delivery prompts. This is consistent with how orchestrate.md is already excluded.

Verdict: PASS

5. Template updates

  • ci-triage/loop.json: roles filled with {"prompt": "loop-implement.md"} etc. All other fields unchanged. Correct.
  • self-improvement/loop.json: has work_source: audit, use_worktree: true, file_scope with 4 paths, max_iterations: 10, score_plateau_window: 3. Matches technical.md section 9. Correct.

Verdict: PASS

6. Test infrastructure updates

Renaming prompt refs from "loop-implement.md" to "test-impl.md" (and similar) in existing tests is the correct approach. These tests don't test prompt resolution -- they test other runner behavior. Using non-existent prompt refs ensures _resolve_prompt falls back to the raw string, preserving the old argv contents that the test assertions depend on.

Verdict: PASS

7. Edge cases

  • Empty prompt_ref: _resolve_prompt returns prompt_ref or "" at line 260. Safe.
  • Missing outputs dir: out_dir.mkdir(parents=True, exist_ok=True) at line 300. Safe.
  • Missing artifact file for {artifact_content}: caught by try/except OSError, returns empty string. Safe.
  • Loop-local prompt override: searched first, allows per-loop customization without modifying framework prompts. Good design.

Verdict: PASS

Summary

All 7 review areas pass. The implementation is correct, backward compatible, and well-tested. 18 new tests cover the prompt resolution, prompt file content, template updates, and tick integration. Full suite: 393 passed.

Overall verdict: APPROVED