- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
9.7 KiB
Add Status Brakes
Implement the loop-aware extension to status.py per design/loops/technical.md §3 and §4. This is the second-tier enforcement layer that the loop runner (task 3) will call. Brakes live inside status.py so they cannot be routed around by the harness.
Goal
Make status.py aware of loops. Add .state.loop files, on-disk loop folder layout, gate-check commands, schedule-unit installers, and the --approve --loop resume path. No runtime/runner code in this task — task 3 (add-loop-runner) wires loop-runner.py to call these commands. This task only ships the enforcement surface.
Requirements
R1. Loop directory layout
Each project gets .automaton/loops/<name>/ containing:
loop.json— copied fromtemplates/loops/<template>/loop.json(template files themselves are task 6's deliverable;--create-loopworks against any existing template dir).state.loop— JSON state file (R2 schema).state.log— append-only tick log, seeded empty on creationworktree/— created lazily on first worktree-needing tick (task 3's runner creates it;--create-loopdoes NOT set up worktree)run-tick.sh— generated by--install-schedule(R6); not present at--create-looptime
Loops without .state.loop are UNTRACKED — mirror of v2.0 task .state rule. All --loop commands refuse to operate on an untracked loop and emit the upgrade hint: Run --upgrade-loops to bootstrap. (--upgrade-loops is not in this task; future bootstrap work. Provided only as the hint target.)
R2. .state.loop schema
JSON:
{
"schema_version": 1,
"name": "<loop-name>",
"status": "running",
"halt_reason": null,
"iteration_count": 0,
"resumed_count": 0,
"last_tick_at": null,
"last_verdict": null,
"score_history": [],
"current_task": null,
"worktree_branch": null,
"worktree_path": null
}
status ∈ {"running", "halted", "paused", "complete"}. halt_reason ∈ the five deaths + null. last_verdict is the most recent verdict JSON or null. score_history is capped at score_plateau_window (from loop.json), FIFO.
_write_state_loop() helper mirrors _write_state()'s atomic-tmp-then-replace pattern.
R3. New flags on status.py
All route through one argparse parser to keep harness integration single-point.
status.py --create-loop <name> --from-template <template> [--project <p>]
status.py --install-schedule <name> [--interval N] [--project <p>]
status.py --pause-loop <name> [--project <p>]
status.py --resume-loop <name> [--project <p>]
status.py --approve --loop <name> [--project <p>]
status.py --can-continue <name> [--project <p>] (--json supported)
status.py --check-gate <name> [--task <t>] [--project <p>] (--json supported)
status.py --can-edit --project <p> [--task <t>] [--file <path>] [--loop <name>] [--loop-worktree]
status.py --loop-list [--project <p>]
status.py --version
--approve --loop is the only way to clear a halt. --resume-loop only clears paused (user-initiated pause), never a halt — refuses with "loop is halted, use --approve --loop to clear halt".
--version reads the ## Framework Version section of config.md and prints as automaton <version>\n. Exit 0 always (matches POSIX convention for --version). When the section is missing, prints automaton (unknown version)\n and still exits 0.
R4. --check-gate JSON return
Returns JSON to stdout (last line, pre-encoded). Exit code 0 on ok:true; exit code 1 on ok:false (HALTED/PAUSED/COMPLETE etc.); exit code 2 on error (loop untracked / not found).
Shape (from technical.md §4):
{
"ok": false,
"reason": "halted:verifier_failed",
"halt_reason": "verifier_failed",
"remaining_iterations": 0,
"remaining_budget_usd": null,
"task_phase": "implement",
"task_in_halt_loop": true,
"out_of_scope_files": []
}
Gate checks execute in order: loop status → iteration count → budget → task phase → worktree drift → score plateau. The first failing check halts and sets halt_reason atomically. Worktree drift requires git diff --name-only main...HEAD scoped to loop.json.blast_radius.file_scope (uses subprocess.run best-effort; on no-git environments, drift check is skipped with a stderr warning, not a halt).
R5. --can-continue shorthand
Returns {"ok": true/false, "status": "running|halted|paused|complete"} — used by schedulers/CI to decide run-tick.sh shouldn't proceed. More general than --check-gate (which is the pre-tick gate). --can-continue is the "is the loop alive at all" check.
R6. --install-schedule platform dispatcher
Detect platform.system():
Darwin→ write~/Library/LaunchAgents/com.automaton.loop.<name>.plistwithStartInterval = interval_seconds. Also writesrun-tick.sh(chmod +x) into the loop dir for the plist'sProgramArguments.Linux→ readcrontab -l, strip any existing# automaton-loop:<name>block, append a new block tagged# automaton-loop:<name>\n*/N * * * * <run-tick.sh>, andcrontab -back. Also writesrun-tick.sh.Windows→schtasks /create /tn "AutomatonLoop_<name>" /tr <run-tick.sh> /sc minute /mo <N_minutes> /f. Also writesrun-tick.bat(Windows uses.bat, not.sh, but the runner is still Python).- Other → refuse with exit 2 and an unsupported-OS message.
run-tick.sh content is locked by technical.md §6:
#!/usr/bin/env bash
cd "<project_root>"
python3 "<framework>/scripts/loop-runner.py" --mode tick --loop "<name>"
--pause-loop:
- Darwin → rename plist to
.disabled. - Linux → strip the
# automaton-loop:<name>block from crontab. - Windows →
schtasks /change /tn "AutomatonLoop_<name>" /disable.
--resume-loop is the inverse; refuse with halt-state error per R3.
R7. --can-edit --loop extension
Existing --can-edit semantics preserved. New flags:
--loop <name>adds a worktree-scope clause: edits allowed only if file is inside<loop_dir>/worktree/(or, when--loop-worktree, against<project>/.automaton/loops/<name>/worktree/).--loop-worktree(requires--loop) switches the file-scope anchor to the worktree path instead of the project root.
Exit codes/host-side output unchanged; only the ALLOWED/DENIED response shifts.
R8. --transition refuses when a halted loop owns the task
status.py --transition <phase> --task <t> already operates on tasks. New behavior: when the task's current_task field is set in any loop whose status is halted and whose halt_reason is one of the five deaths, transitions are refused with "task is bound to halted loop '<name>' (halt_reason=<reason>). --approve --loop <name> to resume.". Exit 1.
When the loop is running or paused, transitions proceed normally (the loop will see the new phase at next tick).
R9. --audit extension
--audit output gains a Loops section listing every loop with (name, status, halt_reason, iteration_count, started_at). Loops in halted state are flagged with an audit warning.
New flag --loop-list provides the same data as --audit's loop section but standalone.
R10. Tests
New file tests/test_status_brakes.py covering:
cmd_create_loop— creates dir + loop.json + .state.loop with default state; refuses on duplicate; refuses on missing template..state.loopschema initialization — all R2 fields present.--approve --loopclears halt, incrementsresumed_count, refuses on running loop, refuses on untracked loop.--resume-loopclears paused, refuses on halted.--pause-loopinvalidates--can-continue.--check-gateJSON for: clean running, halted on iterations, halted on verifier_failed (flat score), halted on drift, paused loop.--can-edit --loopallowed when file under worktree, denied when outside.--transitionrefused when owning loop halted; allowed when running/paused.--install-schedulewritesrun-tick.sh(and a stub plist on Darwin using tmp_path monkey-patching ofPath.home()).--versionprints "automaton " reading from a fixtureconfig.md.
Acceptance Criteria
--create-loopproduces a valid.state.loopwith R2 fields; duplicate-name returns exit 2.--approve --loopincrementsresumed_count, clearshalt_reason, returns status torunning. Refuses on a running loop.--resume-loopclearspausedonly; refuses onhalted.--check-gate --jsonemits the §4 JSON shape; returns exit 1 when not-ok.--can-edit --loop --file <outside>exit 1; the same file inside the worktree exit 0.--transition --task <t>exit 1 when an owning loop is halted.--versionwritesautomaton <version>\nto stdout fromconfig.md's## Framework Versionsection.--install-schedulewritesrun-tick.shand the OS-native schedule unit (Darwin plist / Linux crontab block / Windows schtasks invocation) using a tmp_path fixture.--auditincludes a Loops section.tests/test_status_brakes.pypasses.- Pre-existing framework tests still green:
pytest tests/ -q.
Non-Goals
- No
loop-runner.pyin this task (task 3). - No verifier prompt contents (task 6 templates).
- No worktree creation logic for live ticks (task 5 —
--create-loopmakes the dir but not the worktree). - No
--upgrade-loopscommand (referenced only in error messages; bootstrap path remains manual for v1). - No parallel mode (D6 stays opt-in; not implemented in v1).
Dependencies
- Task 1 (
fix-context-sizing) — DONE.--check-gatebudget check relies onvram_detect.py --loop-modeJSONavailable_context_kb >= 16000. Theloop_mode_eligiblefield is available.
Out of Scope (deferred)
--upgrade-loops(bootstrap pre-2.0 loops) — not blocking v1; manual create-loop is the path.- Dashboard "Loops" panel — v1.1.