- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
10 KiB
SPEC: add-loop-runner
Implements scripts/loop-runner.py --mode tick (and --mode daemon opt-in). The runner is the per-tick engine that calls the brakes, spawns the three session roles (Implement / Verify / Orchestrate), parses the graded verifier verdict, and updates .state.loop. It is the runtime partner of the brakes layer landed in task add-status-brakes.
Goal
A single Python entry point that any OS scheduler (launchd / cron / schtasks) or human can invoke as:
python3 <framework>/scripts/loop-runner.py --mode tick --loop <name> --project <p>
It must:
- Be idempotent in the failure case -- a crash mid-tick does not advance
iteration_countor corrupt.state.loop. - Never invoke an LLM directly. All role sessions are external subprocesses against the user's configured harness, dispatched from
loop.jsonharness.command. - Refuse to run when
--check-gatereturns not-ok, and exit 0 (clean exit; do not crash the scheduler) so the OS unit's retry backoff stays calm. - Apply all six brake gates indirectly via
--check-gate(no duplicated gate logic in the runner).
Requirements
R1 -- Entry point and CLI shape
--mode {tick,daemon}required.--loop NAMErequired.--project PATHoptional (forwarded tostatus.py).--jsonoptional -- emit machine-readable tick summary as the last line.--interval SECONDSfor--mode daemononly (default: read fromloop.jsonschedule.interval_seconds, else 3600).- Unknown
--mode→ exit 2. - Unknown loop (no
.state.loop) → log SKIP, exit 0 (not 2; the runner never escalates a missing loop to a hard error, because the OS scheduler must keep firing).
R2 -- Tick flow (per technical.md §7)
In order:
- Load: read
.state.loopandloop.jsonfrom<loops>/<name>/. Treat missing.state.loopasuntrackedSKIP (R1). - Gate:
subprocess.run([python, status.py, "--check-gate", NAME, "--project", P, "--json"]). Parse JSON. Ifok == false: appendSKIP reason=…to.state.log, exit 0. - Find work (v1: only
singlework_source):current_task = state["current_task"]. If null: SKIPno_current_task.audit/backlogwork_sources are stubbed for v1 (return SKIP) and fleshed out in tasks 4 and 6. - Worktree: deferred to task
add-blast-radius-scheduler. The runner usesstate["worktree_path"]if set elseproject_rootas cwd. If worktree configured but missing, write aworktree_missingwarning to.state.logand SKIP (human_interventionhalts are owned by--check-gate, not the runner). - Spawn Implement: build harness command from
loop.jsonharness.commandwith{prompt}=roles.implement.prompt,{cwd}= resolved cwd,{output}= unique artifact path under<loop>/outputs/<tickN>-<role>.json. Invoke viasubprocess.run. Capture stdout. Do not block on harness timeout; v1 has no timeout (the harness owns its own timeout policy). - Spawn Verify: same as Implement, with
{prompt}=roles.verify.prompt. Add{artifact}substitution token (pointing at Implement's output path). Capture stdout -- this must parse as JSON (verdict). - Parse verdict: accept either raw JSON or ```json fenced blocks or JSON with leading
// / #line comments. Strict keys:pass(bool, required),score(float 0.0–1.0, required),reasons(list of strings, optional),next_hint(string, optional). On parse failure → halt asverifier_failed, writeHALT verifier_failed:unparseableto.state.log, exit 0. - Append score: push
verdict["score"]tostate["score_history"], capped atbrakes.score_plateau_window(drop oldest beyond window). - Spawn Orchestrate:
{prompt}=roles.orchestrate.prompt, plus inject{verdict}(JSON-serialized) and{current_task}and{current_phase}as substitution tokens. The orchestrator's stdout is captured but not parsed in v1 -- the orchestrator is the actor that callsstatus.py --transition/--approveitself (no auto-approve path). - Update state (the runner's own writes -- never overlap with orchestrator writes):
state["iteration_count"] += 1state["last_tick_at"] = iso8601_nowstate["last_verdict"] = verdict- Atomic write via tmp+rename (same helper as status.py -- duplicate the small writer rather than import across scripts).
- Tick log: append
TICK pass=<bool> score=<f> iter=<N>to.state.log. - Exit 0.
Order of failure-mode Halt writes (all delegated to status.py via _disable_schedule best-effort, but the halt itself is a direct .state.loop write from the runner):
- Parse failure →
verifier_failed(R7 above).
The runner does not check iterations / budget / drift / task-phase gates itself -- --check-gate (R2 step 2) already did. The runner is responsible only for verifier_failed (verdict parse) and for verifier_failed (score plateau) indirectly via the next tick's --check-gate.
R3 -- --mode daemon
time.sleep(interval)loop callingcmd_tick().KeyboardInterrupt→ exit 0 cleanly with aDAEMON_STOPPEDlog entry.--max-iterations N(optional) caps daemon loop count. 0 / unset = unbounded.
R4 -- Harness command substitution
loop.json harness.command is a list of strings. The runner walks each element, replacing {prompt}, {cwd}, {output}, {artifact}, {verdict}, {current_task}, {current_phase} with values from the tick context. Missing tokens stay literal (so configurations can opt out of, say, the {output} token by simply not including it).
Default harness.command (when loop.json doesn't specify one) is ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"], matching the user's primary harness (D8 -- never inspect model capability).
R5 -- Context-floor guard (D13)
Before invoking the Implement role, the runner calls vram_detect.py --loop-mode --json. If the JSON loop_mode_eligible == false, the runner halts the loop with human_intervention and writes HALT human_intervention:context_below_floor. Existing shell: a "context too small" loop cannot burn tokens through a harness call that would fail anyway.
This guard is implemented in the runner (not in --check-gate) because --check-gate is per-tick and the available-context value is hardware-state, not loop-state -- we don't want it cached in .state.loop between ticks.
R6 -- Idempotence
- State writes are atomic (tmp+rename).
- The Implement / Verify / Orchestrate invocations do not mutate state; only step 10 writes.
- Verifier parse failure short-circuits before step 10, so a tick that fails to parse its verifier does not increment
iteration_count. The harness retry on next tick starts from the samecurrent_taskanditeration_count. - A
KeyboardInterruptorSIGTERMbetween steps 5 and 10 leaves.state.loopunchanged. The harness subprocess may be left running (the runner does not own process groups in v1).
R7 -- Tests (tests/test_loop_runner.py)
Required by AGENTS.md. All harness calls are stubbed via monkeypatch.setattr(subprocess, "run", fake_run). No live LLM calls in CI.
test_tick_pass-- fixture loop with acurrent_taskinimplement, mock--check-gatereturns ok, mock verifier returns{"pass": true, "score": 0.9}. Assertiteration_count == 1,last_verdict["pass"] is True,.state.loghasTICK pass=True score=0.9 iter=1.test_tick_skip_when_halted-- pre-halt.state.loop, mock--check-gatereturns not-ok. Assertiteration_countunchanged,.state.loghasSKIP reason=halted:….test_tick_skip_when_untracked-- no.state.loop. Assert exit 0,.state.loghasSKIP untracked.test_tick_skip_no_current_task--.state.loophascurrent_task: null. Assert SKIPno_current_task.test_verifier_parse_failure_halts-- mock verifier returns garbage. Assert loop halted asverifier_failed,last_verdictis null,iteration_countunchanged (R6 idempotence).test_score_history_capped-- loop withscore_plateau_window: 3, run 5 ticks with mock verifier returning scores 0.5, 0.4, 0.4, 0.4, 0.4. Assertscore_historylength is 3 (the last three).test_json_output_mode----jsonprints a structured tick summary on the last line.test_daemon_mode_runs_n_iterations----mode daemon --max-iterations 3runscmd_tickthree times then exits 0.test_context_floor_refuses-- mockvram_detect.pyreturnsloop_mode_eligible: false. Assert loop haltedhuman_intervention, harness subprocess never invoked.test_unknown_mode_rejected----mode bogusexits 2.test_unknown_loop_skip_clean_exit----loop ghostexits 0 (R1).test_harness_command_substitution-- fixture loop.json with customharness.commandcontaining{prompt},{cwd},{output}. Assert stubsubprocess.runsaw the substituted values verbatim.test_orchestrator_invoked_after_verifier-- assert subprocess invocations happen in order: gate → implement → verify → orchestrate. Capture argv patterns to confirm.
R8 -- Out of scope (other tasks)
- Live harness adapter -- provided by user as
harness.command; no new adapter code. auditwork_source -- task 4 (goal-mode / verifier session) and task 6 (self-improvement template).backlogwork_source -- task 7 (self-improvement loop) anddesign/<area>/BACKLOG.mdintegration.- Worktree creation plumbing -- task
add-blast-radius-scheduler. - Verifier prompt (
loop-verifier.md) -- task 6. The runner just reads the filename fromloop.jsonand passes it to the harness; it does not parse the prompt itself. - Orchestrator logic that decides phase transitions -- the orchestrator role does that; the runner only spawns the orchestrator and trusts its
status.pycalls.
Approach
Single new file scripts/loop-runner.py. Stdlib-only (no new pip deps). Reuses small helpers (_read_state_loop, _write_state_loop, _loop_dir, _read_loop_config) duplicated inline rather than imported from status.py -- keeps the scripts independent (no risk of one script accidentally being moved/broken affecting the other). The duplicate helpers are <30 lines total.
Tests file tests/test_loop_runner.py uses tmp_path + a _stub_subprocess helper that pattern-matches on argv to return canned outputs.
Verification
python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -v
python3 -m pytest tests/ -q # ensure no regressions