Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

4.7 KiB
Raw Permalink Blame History

Implementation: add-loop-runner

Implements scripts/loop-runner.py per SPEC R1–R8.

File added

scripts/loop-runner.py — single entry point for --mode tick and --mode daemon. Stdlib only (no new pip deps).

Layout

  • LOOP_* constants mirroring status.py for the few state-shape facts the runner needs.
  • Small helpers duplicated inline rather than imported across scripts (per technical.md: scripts stay independent; no cross-script imports): _find_project_dir, _loops_dir, _loop_dir, _read_state_loop, _write_state_loop, _read_loop_config, _append_tick_log, _halt_loop.
  • _run_json(args) — invokes a subprocess and parses the last stdout line as JSON. Returns None on subprocess failure, non-zero exit, empty output, or JSON parse failure. Used by both _gate and _context_floor_ok.
  • _substitute(template, mapping) — token substitution for loop.json harness.command strings. Recognized tokens: {prompt}, {cwd}, {output}, {artifact}, {verdict}, {current_task}, {current_phase}.
  • _invoke_harness(harness_cfg, role, prompt_path, cwd, extras) — builds the harness command, substitutes tokens, runs subprocess.run, returns stdout. Default command when harness.command is missing is ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"].
  • parse_verdict(text) — strict graded-JSON parser. Accepts raw JSON, fenced ```json blocks, lines with leading // or # comments stripped. Returns None when missing pass key or total garbage. Otherwise returns {"pass": bool, "score": float, "reasons": list, "next_hint": str?}.
  • _gate(...), _context_floor_ok(), _role_prompt(...), _score_window(...), _loop_max_iterations(...), _outputs_dir(...), _make_completed (test helper used inline).
  • cmd_tick(args) — the tick flow per technical.md §7. Returns a summary dict, never raises (clean-exit on every path).
  • cmd_daemon(args) — time.sleep(interval) loop bounded by --max-iterations. KeyboardInterrupt stops cleanly with a DAEMON_STOPPED log entry.
  • main() — argparse with --mode {tick,daemon}, --loop, --project, --interval, --max-iterations, --json.

R-by-R coverage

Req Code
R1 entrypoint main() argparse, --mode required choices; cmd_tick returns summary with skipped:True and reason:"untracked" for missing .state.loop
R2 tick flow cmd_tick 7-route: load → gate → find_work → cwd → ctx-floor → spawn Implement → spawn Verify → parse verdict → cap score → spawn Orchestrate → atomic write state → tick log
R3 daemon cmd_daemon
R4 harness substitution _substitute, _invoke_harness
R5 context-floor guard _context_floor_ok called before any harness subprocess; halts human_intervention on loop_mode_eligible=False
R6 idempotence state writes only in step 10 (after parse_verdict succeeds and orchestrator ran); pre-step-10 crashes leave .state.loop untouched
R7 tests tests/test_loop_runner.py (18 tests)
R8 out-of-scope none — deferred to tasks 4–7 (audit work_source, backlog, worktree creation, the prompts themselves)

Tests (tests/test_loop_runner.py)

18 tests across 7 classes; all subprocess.run and _run_json calls stubbed via monkeypatch so no live LLM calls hit in CI.

  • TestEntrypoint (2): unknown-loop exits 0; unknown-mode exits 2.
  • TestTickFlow (5): tick-pass advances iteration_count; skip-when-halted; skip-when-untracked; skip-no-current-task; skip-when-gate-subprocess-fails.
  • TestContextFloor (1): refuses below floor; halts human_intervention; implement harness never invoked.
  • TestVerifierParseFailure (5): parse-failure halts and does not advance iteration_count (idempotence); fenced JSON parses; JSON with line comments parses; missing pass key → None; empty text → None.
  • TestScoreHistory (1): 5 ticks with window=3 → final score_history length is 3 and equals [0.4, 0.4, 0.4].
  • TestHarnessSubstitution (1): custom harness.command with --prompt/--cwd/--out/--artifact tokens; verify-role invocation sees the implement role's output path as --artifact <...-implement.json>.
  • TestDaemonMode (1): --max-iterations 3 runs 3 ticks then exits 0; time.sleep no-op via monkeypatch.
  • TestOrchestratorOrdering (1): implement → verify → orchestrate order observed via tagged handlers.
  • TestJsonOutput (1): --json prints structured tick summary as last line; parsed via lr.main() + capsys (since subprocess.run is patched).

Verification

python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -q   # 18 passed
python3 -m pytest tests/ -q                       # 328 passed (was 310 + 18 new)