- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
50 lines
3.3 KiB
Markdown
50 lines
3.3 KiB
Markdown
# Code Review: add-status-brakes
|
||
|
||
Reviewed against SPEC.md R1–R10. All requirements implemented; no functional gaps found.
|
||
|
||
## R1–R10 checklist
|
||
|
||
| Req | Status | Notes |
|
||
|-----|--------|-------|
|
||
| R1 `.state.loop` schema | ✅ | All 13 defaults present; atomic write via tmp+rename |
|
||
| R2 `--create-loop` | ✅ | kebab/Dup/template validation; name patching |
|
||
| R3 `--version`, `--approve --loop` | ✅ | version parses `## Framework Version`; approve only clears halt; `resumed_count++` |
|
||
| R4 `--can-continue` | ✅ | Correct boolean: `status == "running"` only |
|
||
| R5 `--check-gate` (6 gates) | ✅ | Order matches SPEC; first failure halts; JSON structured |
|
||
| R6 `--install-schedule` | ✅ | Triple dispatch Darwin/Linux/Windows; stubs generated; pause disables (best-effort) |
|
||
| R7 `--can-edit --loop [--loop-worktree]` | ✅ | Root residency + file_scope; refuses outside root |
|
||
| R8 `--transition` halt refusal | ✅ | Owned-task scan; points user at `--approve --loop` |
|
||
| R9 `--audit`/`--loop-list` | ✅ | Cat-6 runs even with no tasks; untracked/halted flagged; missing current_task flagged |
|
||
| R10 `.state.log` | ✅ | ISO timestamps; tested for PAUSED/RESUMED/APPROVED/HALT |
|
||
|
||
## Defensive coding observations
|
||
|
||
1. **Atomic `.state.loop` writes** — tmp+`replace()`. Crashes mid-write cannot corrupt state.
|
||
2. **Best-effort schedule disable** — wrapped in `try/except` so a non-existent cron/plist on a dev box cannot crash `--pause-loop` or the halt path. `.state.loop` remains source of truth; the OS unit reads it on next wake and self-skips.
|
||
3. **No new pip deps** — stdlib only (`platform`, `subprocess`, `json`, `re`, `datetime`). Per project constraints.
|
||
4. **Harness-agnostic** — every gate is reachable via `status.py` subprocess + `--json`. No harness-specific code. Works with opencode, any other harness, or a raw shell.
|
||
5. **`--approve --loop` is the only halt-clear** — D4 enforced; `--resume-loop` explicitly refuses halted loops and tells the user to approve.
|
||
6. **R8 ownership scan** — `_loop_owning_task` is O(loops) per transition; loops are few, so fine. Could be cached later if needed.
|
||
|
||
## Edge cases checked
|
||
|
||
- Empty project (no tasks) — `--audit` still runs Cat-6 (R9 fix; was originally early-return).
|
||
- Loop with no `loop.json` — `--install-schedule` exits 2 with clear message.
|
||
- Loop with no `.state.loop` — every `--loop` command refuses with the `_loop_untracked_hint`.
|
||
- `--check-gate` on a paused loop — `_gate_loop_status` returns the `paused:` reason (not a halt, since the user paused it; harness checks separately via `--can-continue`).
|
||
- Budget informational when `max_budget_usd == null` — gate skipped, returns None.
|
||
- Score plateau with too-short history — gate skipped.
|
||
- Worktree missing — `_gate_worktree_drift` treats as no-drift (runner will recreate).
|
||
- `git diff` failure — warning logged to stderr, drift gate skips. Not a halt; per "best-effort portable" principle (D13).
|
||
|
||
## Things deliberately NOT in this task (per scope)
|
||
|
||
- `loop-runner.py` itself — task 3.
|
||
- Verifier role / graded JSON — task 4.
|
||
- Worktree creation plumbing — task 5.
|
||
- Full `templates/loops/ci-triage/` content (prompts, README) — task 6.
|
||
- `--upgrade-loops` for stray pre-state-loop dirs —audit just flags them. Refactor in v1.1.
|
||
|
||
## Verdict
|
||
|
||
APPROVE. No blocking issues. Ready for bug_find. |