- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
4.1 KiB
Adversarial Bug Report: add-status-brakes
Adversarial pass: attack the brakes layer as a hostile agent would, looking for ways to escape enforcement.
Attack vectors tried
A1 — Can a halted loop be silently resumed by re-running --create-loop with the same name?
cmd_create_loop refuses duplicates (exit 2). Attacker cannot overwrite a halted loop's state via create. Defended.
A2 — Can a hostile harness skip --check-gate and just call --transition?
Not if the loop owns the task — R8 refuses the transition until --approve --loop clears the halt. Defended.
BUT: if the loop never current_task-claimed the task, _loop_owning_task returns None and the transition proceeds. The agent can edit a task that isn't claimed by any loop. That is correct behavior (humans and ad-hoc agents can still work), but it means a hostile agent could race the loop runner to claim a task. Mitigation: loop runner should call a --claim-loop-task (not in v1) or set current_task atomically before transitioning. Not a brakes-layer bug — it's a runner-layer concern (task 3). Note for task 3.
A3 — Can --pause-loop be used to bypass the worktree drift gate?
No — pausing just stops the scheduler; .state.loop remains the source of truth. --check-gate still runs all gates; a paused loop's gates return paused: (not ok). The runner will not tick. Defended.
A4 — Can the agent hand-write .state.loop to fake "running"?
Yes — .state.loop is a plain JSON file. A fully local adversary with filesystem access can edit it. This is an accepted threat model: the framework assumes the harness is cooperative for state files but enforces via the pre-edit hook (--can-edit) for source edits. .state.loop is not in any task's file_scope, so it's never editable by a loop agent. Defended by file-scope design.
A5 — Race: two concurrent --check-gate invocations both halt the loop
Both call _halt_loop which uses atomic tmp+rename. Last writer wins. Both write the same halt_reason (deterministic from gate), so the result is consistent. No corruption. Defended.
A6 — Can --approve --loop be called while the loop is mid-tick?
--approve does tmp+rename. If a tick is concurrently writing iteration_count, the approve's write wins and the tick's increment is lost. Window is small (subprocess boundary). Acceptable for v1; the next tick re-reads and re-increments. Not a corruption vector. Note for v1.1: file-locking (fcntl) on .state.loop would close this race. Add to BACKLOG.
A7 — Can --install-schedule be pointed at a different project than the loop?
--install-schedule uses _find_project_dir(args.project) and writes the stub at loop_path / run-tick.*. The stub cds into the project root and invokes the runner with the loop name. An attacker could swap the loop_name in the stub after generation, but that's just running an arbitrary loop — not a privilege escalation. Not an attack.
A8 — Can the schedule wake the loop after it's halted?
Yes — the OS unit fires run-tick on schedule. run-tick invokes loop-runner.py --mode tick --loop NAME, which must call --check-gate first and exit 1 if not ok. The runner's contract (task 3) is: gate first, then work. The OS unit itself cannot refuse. So a halted loop's schedule will fire run-tick, which will no-op via the runner's gate check. The --pause-loop best-effort disable is belt-and-braces. Defended by runner contract (must be enforced in task 3).
Hardening recommendations (for BACKLOG)
fcntlfile-lock on.state.loopfor tick/approve race (A6) — v1.1.--claim-loop-taskto atomically setcurrent_taskbefore a runner touches the task (A2) — task 3._enable_scheduleLinux parity with Darwin/Windows (O4) — task 5 / v1.1.blast_radius.base_branchparameterization for drift diff (O3) — task 5.
Verdict
PASS — no exploitable escape from the brakes layer. All adversarial vectors are either defended today or have explicit runner-contract mitigations landing in tasks 3/5. Hardening items routed to design/loops/BACKLOG.md.