Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

3.7 KiB

CODE_REVIEW: add-goal-mode

Reviewed against SPEC.md R1-R8.

R1-R8 checklist

Req Status Notes
R1 find_work dispatch PASS _find_work dispatches on work_source.kind; missing/unknown falls back to single with WARNING
R2 audit work_source PASS _find_work_audit calls --audit --json, sorts by severity, creates task via --create-task when no task field
R3 backlog work_source PASS _find_work_backlog reads design/<area>/BACKLOG.md, picks top - [ ], slugifies bold heading
R4 verifier tokens PASS {task_brief}, {acceptance_criteria}, {next_hint} in extras; substituted via _substitute
R5 truncate_tokens PASS 4 chars/token heuristic; marker appended; caps at 4000/2000/1000
R6 next_hint loop PASS _next_hint_text reads last_verdict.next_hint; empty on first tick; fed into implement and verify
R7 loop.json schema PASS ci-triage template has explicit work_source + acceptance_criteria; technical.md updated
R8 audit --json PASS cmd_audit emits JSON line with violations/loops/total_tasks/untracked_tasks

Edge cases checked

  1. Missing work_source field -- falls back to single with no WARNING (only unknown kinds warn). Backward compat with ci-triage template preserved. PASS
  2. Unknown work_source.kind -- WARNING logged, falls back to single. PASS
  3. Audit with no violations -- returns None from _find_work_audit; skip reason no_work; does not increment iteration_count. PASS
  4. Audit violation with null task -- slugifies message, calls --create-task, returns slug. PASS
  5. Audit violation with existing task -- returns task name directly, no create-task call. PASS
  6. Backlog with all items checked -- returns None; skip no_work. PASS
  7. Backlog with no bold marker -- falls back to _slugify(line_body). PASS
  8. Empty task_brief / acceptance_criteria / next_hint -- _truncate_tokens("") returns ""; substitution replaces with empty string; no KeyError. PASS
  9. last_verdict is None -- _next_hint_text checks isinstance(last, dict); returns "". PASS
  10. --audit --json with no violations -- emits {"violations":[], ...}; runner sees empty list, skips. PASS
  11. --audit --json output pickable by _run_json -- single JSON line on stdout; _run_json takes splitlines()[-1]. PASS

Code-quality observations

  1. _find_work_audit subprocess timeout=15 for --create-task -- reasonable; if create-task hangs, the runner swallows it and returns the slug anyway. The task dir may not exist yet, but the orchestrator will handle it on the next tick. Acceptable for v1.
  2. _slugify used for both audit and backlog -- consistent slug derivation. The regex [^A-Za-z0-9._-]+ -> - is reasonable.
  3. _task_dir_for duplicates status.py _task_dir logic -- documented as intentional (no cross-script imports per technical.md). If the task dir layout changes, both need updating. Acceptable for v1.
  4. Token substitution only works if harness command contains the placeholder -- the default command ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"] does not include {task_brief} etc. Custom harness configs must add them explicitly. This is by design (SPEC R4: "tokens absent from the prompt stay literal").
  5. _acceptance_criteria_text handles both string and list -- list joined with newlines. If the value is a dict or other type, str(raw) is called. Defensive enough.
  6. _find_work_backlog reads from design/<area>/BACKLOG.md -- uses project_dir == AUTOMATON_DIR check to pick framework vs project path. Consistent with _task_dir_for pattern.

Verdict

APPROVE. Ready for bug_find.