- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
3.3 KiB
3.3 KiB
Code Review: add-status-brakes
Reviewed against SPEC.md R1–R10. All requirements implemented; no functional gaps found.
R1–R10 checklist
| Req | Status | Notes |
|---|---|---|
R1 .state.loop schema |
✅ | All 13 defaults present; atomic write via tmp+rename |
R2 --create-loop |
✅ | kebab/Dup/template validation; name patching |
R3 --version, --approve --loop |
✅ | version parses ## Framework Version; approve only clears halt; resumed_count++ |
R4 --can-continue |
✅ | Correct boolean: status == "running" only |
R5 --check-gate (6 gates) |
✅ | Order matches SPEC; first failure halts; JSON structured |
R6 --install-schedule |
✅ | Triple dispatch Darwin/Linux/Windows; stubs generated; pause disables (best-effort) |
R7 --can-edit --loop [--loop-worktree] |
✅ | Root residency + file_scope; refuses outside root |
R8 --transition halt refusal |
✅ | Owned-task scan; points user at --approve --loop |
R9 --audit/--loop-list |
✅ | Cat-6 runs even with no tasks; untracked/halted flagged; missing current_task flagged |
R10 .state.log |
✅ | ISO timestamps; tested for PAUSED/RESUMED/APPROVED/HALT |
Defensive coding observations
- Atomic
.state.loopwrites — tmp+replace(). Crashes mid-write cannot corrupt state. - Best-effort schedule disable — wrapped in
try/exceptso a non-existent cron/plist on a dev box cannot crash--pause-loopor the halt path..state.loopremains source of truth; the OS unit reads it on next wake and self-skips. - No new pip deps — stdlib only (
platform,subprocess,json,re,datetime). Per project constraints. - Harness-agnostic — every gate is reachable via
status.pysubprocess +--json. No harness-specific code. Works with opencode, any other harness, or a raw shell. --approve --loopis the only halt-clear — D4 enforced;--resume-loopexplicitly refuses halted loops and tells the user to approve.- R8 ownership scan —
_loop_owning_taskis O(loops) per transition; loops are few, so fine. Could be cached later if needed.
Edge cases checked
- Empty project (no tasks) —
--auditstill runs Cat-6 (R9 fix; was originally early-return). - Loop with no
loop.json—--install-scheduleexits 2 with clear message. - Loop with no
.state.loop— every--loopcommand refuses with the_loop_untracked_hint. --check-gateon a paused loop —_gate_loop_statusreturns thepaused:reason (not a halt, since the user paused it; harness checks separately via--can-continue).- Budget informational when
max_budget_usd == null— gate skipped, returns None. - Score plateau with too-short history — gate skipped.
- Worktree missing —
_gate_worktree_drifttreats as no-drift (runner will recreate). git difffailure — warning logged to stderr, drift gate skips. Not a halt; per "best-effort portable" principle (D13).
Things deliberately NOT in this task (per scope)
loop-runner.pyitself — task 3.- Verifier role / graded JSON — task 4.
- Worktree creation plumbing — task 5.
- Full
templates/loops/ci-triage/content (prompts, README) — task 6. --upgrade-loopsfor stray pre-state-loop dirs —audit just flags them. Refactor in v1.1.
Verdict
APPROVE. No blocking issues. Ready for bug_find.