Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

10 KiB
Raw Permalink Blame History

SPEC: add-loop-runner

Implements scripts/loop-runner.py --mode tick (and --mode daemon opt-in). The runner is the per-tick engine that calls the brakes, spawns the three session roles (Implement / Verify / Orchestrate), parses the graded verifier verdict, and updates .state.loop. It is the runtime partner of the brakes layer landed in task add-status-brakes.

Goal

A single Python entry point that any OS scheduler (launchd / cron / schtasks) or human can invoke as:

python3 <framework>/scripts/loop-runner.py --mode tick --loop <name> --project <p>

It must:

  • Be idempotent in the failure case -- a crash mid-tick does not advance iteration_count or corrupt .state.loop.
  • Never invoke an LLM directly. All role sessions are external subprocesses against the user's configured harness, dispatched from loop.json harness.command.
  • Refuse to run when --check-gate returns not-ok, and exit 0 (clean exit; do not crash the scheduler) so the OS unit's retry backoff stays calm.
  • Apply all six brake gates indirectly via --check-gate (no duplicated gate logic in the runner).

Requirements

R1 -- Entry point and CLI shape

  • --mode {tick,daemon} required.
  • --loop NAME required.
  • --project PATH optional (forwarded to status.py).
  • --json optional -- emit machine-readable tick summary as the last line.
  • --interval SECONDS for --mode daemon only (default: read from loop.json schedule.interval_seconds, else 3600).
  • Unknown --mode → exit 2.
  • Unknown loop (no .state.loop) → log SKIP, exit 0 (not 2; the runner never escalates a missing loop to a hard error, because the OS scheduler must keep firing).

R2 -- Tick flow (per technical.md §7)

In order:

  1. Load: read .state.loop and loop.json from <loops>/<name>/. Treat missing .state.loop as untracked SKIP (R1).
  2. Gate: subprocess.run([python, status.py, "--check-gate", NAME, "--project", P, "--json"]). Parse JSON. If ok == false: append SKIP reason=… to .state.log, exit 0.
  3. Find work (v1: only single work_source): current_task = state["current_task"]. If null: SKIP no_current_task. audit / backlog work_sources are stubbed for v1 (return SKIP) and fleshed out in tasks 4 and 6.
  4. Worktree: deferred to task add-blast-radius-scheduler. The runner uses state["worktree_path"] if set else project_root as cwd. If worktree configured but missing, write a worktree_missing warning to .state.log and SKIP (human_intervention halts are owned by --check-gate, not the runner).
  5. Spawn Implement: build harness command from loop.json harness.command with {prompt} = roles.implement.prompt, {cwd} = resolved cwd, {output} = unique artifact path under <loop>/outputs/<tickN>-<role>.json. Invoke via subprocess.run. Capture stdout. Do not block on harness timeout; v1 has no timeout (the harness owns its own timeout policy).
  6. Spawn Verify: same as Implement, with {prompt} = roles.verify.prompt. Add {artifact} substitution token (pointing at Implement's output path). Capture stdout -- this must parse as JSON (verdict).
  7. Parse verdict: accept either raw JSON or ```json fenced blocks or JSON with leading // / # line comments. Strict keys: pass (bool, required), score (float 0.0–1.0, required), reasons (list of strings, optional), next_hint (string, optional). On parse failure → halt as verifier_failed, write HALT verifier_failed:unparseable to .state.log, exit 0.
  8. Append score: push verdict["score"] to state["score_history"], capped at brakes.score_plateau_window (drop oldest beyond window).
  9. Spawn Orchestrate: {prompt} = roles.orchestrate.prompt, plus inject {verdict} (JSON-serialized) and {current_task} and {current_phase} as substitution tokens. The orchestrator's stdout is captured but not parsed in v1 -- the orchestrator is the actor that calls status.py --transition / --approve itself (no auto-approve path).
  10. Update state (the runner's own writes -- never overlap with orchestrator writes):
    • state["iteration_count"] += 1
    • state["last_tick_at"] = iso8601_now
    • state["last_verdict"] = verdict
    • Atomic write via tmp+rename (same helper as status.py -- duplicate the small writer rather than import across scripts).
  11. Tick log: append TICK pass=<bool> score=<f> iter=<N> to .state.log.
  12. Exit 0.

Order of failure-mode Halt writes (all delegated to status.py via _disable_schedule best-effort, but the halt itself is a direct .state.loop write from the runner):

  • Parse failure → verifier_failed (R7 above).

The runner does not check iterations / budget / drift / task-phase gates itself -- --check-gate (R2 step 2) already did. The runner is responsible only for verifier_failed (verdict parse) and for verifier_failed (score plateau) indirectly via the next tick's --check-gate.

R3 -- --mode daemon

  • time.sleep(interval) loop calling cmd_tick().
  • KeyboardInterrupt → exit 0 cleanly with a DAEMON_STOPPED log entry.
  • --max-iterations N (optional) caps daemon loop count. 0 / unset = unbounded.

R4 -- Harness command substitution

loop.json harness.command is a list of strings. The runner walks each element, replacing {prompt}, {cwd}, {output}, {artifact}, {verdict}, {current_task}, {current_phase} with values from the tick context. Missing tokens stay literal (so configurations can opt out of, say, the {output} token by simply not including it).

Default harness.command (when loop.json doesn't specify one) is ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"], matching the user's primary harness (D8 -- never inspect model capability).

R5 -- Context-floor guard (D13)

Before invoking the Implement role, the runner calls vram_detect.py --loop-mode --json. If the JSON loop_mode_eligible == false, the runner halts the loop with human_intervention and writes HALT human_intervention:context_below_floor. Existing shell: a "context too small" loop cannot burn tokens through a harness call that would fail anyway.

This guard is implemented in the runner (not in --check-gate) because --check-gate is per-tick and the available-context value is hardware-state, not loop-state -- we don't want it cached in .state.loop between ticks.

R6 -- Idempotence

  • State writes are atomic (tmp+rename).
  • The Implement / Verify / Orchestrate invocations do not mutate state; only step 10 writes.
  • Verifier parse failure short-circuits before step 10, so a tick that fails to parse its verifier does not increment iteration_count. The harness retry on next tick starts from the same current_task and iteration_count.
  • A KeyboardInterrupt or SIGTERM between steps 5 and 10 leaves .state.loop unchanged. The harness subprocess may be left running (the runner does not own process groups in v1).

R7 -- Tests (tests/test_loop_runner.py)

Required by AGENTS.md. All harness calls are stubbed via monkeypatch.setattr(subprocess, "run", fake_run). No live LLM calls in CI.

  1. test_tick_pass -- fixture loop with a current_task in implement, mock --check-gate returns ok, mock verifier returns {"pass": true, "score": 0.9}. Assert iteration_count == 1, last_verdict["pass"] is True, .state.log has TICK pass=True score=0.9 iter=1.
  2. test_tick_skip_when_halted -- pre-halt .state.loop, mock --check-gate returns not-ok. Assert iteration_count unchanged, .state.log has SKIP reason=halted:….
  3. test_tick_skip_when_untracked -- no .state.loop. Assert exit 0, .state.log has SKIP untracked.
  4. test_tick_skip_no_current_task -- .state.loop has current_task: null. Assert SKIP no_current_task.
  5. test_verifier_parse_failure_halts -- mock verifier returns garbage. Assert loop halted as verifier_failed, last_verdict is null, iteration_count unchanged (R6 idempotence).
  6. test_score_history_capped -- loop with score_plateau_window: 3, run 5 ticks with mock verifier returning scores 0.5, 0.4, 0.4, 0.4, 0.4. Assert score_history length is 3 (the last three).
  7. test_json_output_mode -- --json prints a structured tick summary on the last line.
  8. test_daemon_mode_runs_n_iterations -- --mode daemon --max-iterations 3 runs cmd_tick three times then exits 0.
  9. test_context_floor_refuses -- mock vram_detect.py returns loop_mode_eligible: false. Assert loop halted human_intervention, harness subprocess never invoked.
  10. test_unknown_mode_rejected -- --mode bogus exits 2.
  11. test_unknown_loop_skip_clean_exit -- --loop ghost exits 0 (R1).
  12. test_harness_command_substitution -- fixture loop.json with custom harness.command containing {prompt}, {cwd}, {output}. Assert stub subprocess.run saw the substituted values verbatim.
  13. test_orchestrator_invoked_after_verifier -- assert subprocess invocations happen in order: gate → implement → verify → orchestrate. Capture argv patterns to confirm.

R8 -- Out of scope (other tasks)

  • Live harness adapter -- provided by user as harness.command; no new adapter code.
  • audit work_source -- task 4 (goal-mode / verifier session) and task 6 (self-improvement template).
  • backlog work_source -- task 7 (self-improvement loop) and design/<area>/BACKLOG.md integration.
  • Worktree creation plumbing -- task add-blast-radius-scheduler.
  • Verifier prompt (loop-verifier.md) -- task 6. The runner just reads the filename from loop.json and passes it to the harness; it does not parse the prompt itself.
  • Orchestrator logic that decides phase transitions -- the orchestrator role does that; the runner only spawns the orchestrator and trusts its status.py calls.

Approach

Single new file scripts/loop-runner.py. Stdlib-only (no new pip deps). Reuses small helpers (_read_state_loop, _write_state_loop, _loop_dir, _read_loop_config) duplicated inline rather than imported from status.py -- keeps the scripts independent (no risk of one script accidentally being moved/broken affecting the other). The duplicate helpers are <30 lines total.

Tests file tests/test_loop_runner.py uses tmp_path + a _stub_subprocess helper that pattern-matches on argv to return canned outputs.

Verification

python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -v
python3 -m pytest tests/ -q   # ensure no regressions