- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
3.7 KiB
3.7 KiB
CODE_REVIEW: add-goal-mode
Reviewed against SPEC.md R1-R8.
R1-R8 checklist
| Req | Status | Notes |
|---|---|---|
| R1 find_work dispatch | PASS | _find_work dispatches on work_source.kind; missing/unknown falls back to single with WARNING |
| R2 audit work_source | PASS | _find_work_audit calls --audit --json, sorts by severity, creates task via --create-task when no task field |
| R3 backlog work_source | PASS | _find_work_backlog reads design/<area>/BACKLOG.md, picks top - [ ], slugifies bold heading |
| R4 verifier tokens | PASS | {task_brief}, {acceptance_criteria}, {next_hint} in extras; substituted via _substitute |
| R5 truncate_tokens | PASS | 4 chars/token heuristic; marker appended; caps at 4000/2000/1000 |
| R6 next_hint loop | PASS | _next_hint_text reads last_verdict.next_hint; empty on first tick; fed into implement and verify |
| R7 loop.json schema | PASS | ci-triage template has explicit work_source + acceptance_criteria; technical.md updated |
| R8 audit --json | PASS | cmd_audit emits JSON line with violations/loops/total_tasks/untracked_tasks |
Edge cases checked
- Missing
work_sourcefield -- falls back tosinglewith no WARNING (only unknown kinds warn). Backward compat with ci-triage template preserved. PASS - Unknown
work_source.kind-- WARNING logged, falls back tosingle. PASS - Audit with no violations -- returns
Nonefrom_find_work_audit; skip reasonno_work; does not increment iteration_count. PASS - Audit violation with null task -- slugifies message, calls
--create-task, returns slug. PASS - Audit violation with existing task -- returns task name directly, no create-task call. PASS
- Backlog with all items checked -- returns
None; skipno_work. PASS - Backlog with no bold marker -- falls back to
_slugify(line_body). PASS - Empty task_brief / acceptance_criteria / next_hint --
_truncate_tokens("")returns""; substitution replaces with empty string; no KeyError. PASS last_verdictis None --_next_hint_textchecksisinstance(last, dict); returns"". PASS--audit --jsonwith no violations -- emits{"violations":[], ...}; runner sees empty list, skips. PASS--audit --jsonoutput pickable by_run_json-- single JSON line on stdout;_run_jsontakessplitlines()[-1]. PASS
Code-quality observations
_find_work_auditsubprocess timeout=15 for--create-task-- reasonable; if create-task hangs, the runner swallows it and returns the slug anyway. The task dir may not exist yet, but the orchestrator will handle it on the next tick. Acceptable for v1._slugifyused for both audit and backlog -- consistent slug derivation. The regex[^A-Za-z0-9._-]+->-is reasonable._task_dir_forduplicatesstatus.py_task_dirlogic -- documented as intentional (no cross-script imports per technical.md). If the task dir layout changes, both need updating. Acceptable for v1.- Token substitution only works if harness command contains the placeholder -- the default command
["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]does not include{task_brief}etc. Custom harness configs must add them explicitly. This is by design (SPEC R4: "tokens absent from the prompt stay literal"). _acceptance_criteria_texthandles both string and list -- list joined with newlines. If the value is a dict or other type,str(raw)is called. Defensive enough._find_work_backlogreads fromdesign/<area>/BACKLOG.md-- usesproject_dir == AUTOMATON_DIRcheck to pick framework vs project path. Consistent with_task_dir_forpattern.
Verdict
APPROVE. Ready for bug_find.