Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

3.3 KiB
Raw Permalink Blame History

Code Review: add-status-brakes

Reviewed against SPEC.md R1–R10. All requirements implemented; no functional gaps found.

R1–R10 checklist

Req Status Notes
R1 .state.loop schema ✅ All 13 defaults present; atomic write via tmp+rename
R2 --create-loop ✅ kebab/Dup/template validation; name patching
R3 --version, --approve --loop ✅ version parses ## Framework Version; approve only clears halt; resumed_count++
R4 --can-continue ✅ Correct boolean: status == "running" only
R5 --check-gate (6 gates) ✅ Order matches SPEC; first failure halts; JSON structured
R6 --install-schedule ✅ Triple dispatch Darwin/Linux/Windows; stubs generated; pause disables (best-effort)
R7 --can-edit --loop [--loop-worktree] ✅ Root residency + file_scope; refuses outside root
R8 --transition halt refusal ✅ Owned-task scan; points user at --approve --loop
R9 --audit/--loop-list ✅ Cat-6 runs even with no tasks; untracked/halted flagged; missing current_task flagged
R10 .state.log ✅ ISO timestamps; tested for PAUSED/RESUMED/APPROVED/HALT

Defensive coding observations

  1. Atomic .state.loop writes — tmp+replace(). Crashes mid-write cannot corrupt state.
  2. Best-effort schedule disable — wrapped in try/except so a non-existent cron/plist on a dev box cannot crash --pause-loop or the halt path. .state.loop remains source of truth; the OS unit reads it on next wake and self-skips.
  3. No new pip deps — stdlib only (platform, subprocess, json, re, datetime). Per project constraints.
  4. Harness-agnostic — every gate is reachable via status.py subprocess + --json. No harness-specific code. Works with opencode, any other harness, or a raw shell.
  5. --approve --loop is the only halt-clear — D4 enforced; --resume-loop explicitly refuses halted loops and tells the user to approve.
  6. R8 ownership scan — _loop_owning_task is O(loops) per transition; loops are few, so fine. Could be cached later if needed.

Edge cases checked

  • Empty project (no tasks) — --audit still runs Cat-6 (R9 fix; was originally early-return).
  • Loop with no loop.json — --install-schedule exits 2 with clear message.
  • Loop with no .state.loop — every --loop command refuses with the _loop_untracked_hint.
  • --check-gate on a paused loop — _gate_loop_status returns the paused: reason (not a halt, since the user paused it; harness checks separately via --can-continue).
  • Budget informational when max_budget_usd == null — gate skipped, returns None.
  • Score plateau with too-short history — gate skipped.
  • Worktree missing — _gate_worktree_drift treats as no-drift (runner will recreate).
  • git diff failure — warning logged to stderr, drift gate skips. Not a halt; per "best-effort portable" principle (D13).

Things deliberately NOT in this task (per scope)

  • loop-runner.py itself — task 3.
  • Verifier role / graded JSON — task 4.
  • Worktree creation plumbing — task 5.
  • Full templates/loops/ci-triage/ content (prompts, README) — task 6.
  • --upgrade-loops for stray pre-state-loop dirs —audit just flags them. Refactor in v1.1.

Verdict

APPROVE. No blocking issues. Ready for bug_find.