Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test

- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
This commit is contained in:
Lap Tran
2026-06-26 10:05:18 -04:00
parent fe43b9e1fc
commit bc7daf8590
666 changed files with 15994 additions and 69 deletions
+55
View File
@@ -0,0 +1,55 @@
# Implementation: add-loop-runner
Implements `scripts/loop-runner.py` per SPEC R1–R8.
## File added
`scripts/loop-runner.py` — single entry point for `--mode tick` and `--mode daemon`. Stdlib only (no new pip deps).
## Layout
- `LOOP_*` constants mirroring `status.py` for the few state-shape facts the runner needs.
- Small helpers duplicated inline rather than imported across scripts (per technical.md: scripts stay independent; no cross-script imports): `_find_project_dir`, `_loops_dir`, `_loop_dir`, `_read_state_loop`, `_write_state_loop`, `_read_loop_config`, `_append_tick_log`, `_halt_loop`.
- `_run_json(args)` — invokes a subprocess and parses the last stdout line as JSON. Returns `None` on subprocess failure, non-zero exit, empty output, or JSON parse failure. Used by both `_gate` and `_context_floor_ok`.
- `_substitute(template, mapping)` — token substitution for `loop.json` `harness.command` strings. Recognized tokens: `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`.
- `_invoke_harness(harness_cfg, role, prompt_path, cwd, extras)` — builds the harness command, substitutes tokens, runs `subprocess.run`, returns stdout. Default command when `harness.command` is missing is `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]`.
- `parse_verdict(text)` — strict graded-JSON parser. Accepts raw JSON, fenced ```json blocks, lines with leading `//` or `#` comments stripped. Returns `None` when missing `pass` key or total garbage. Otherwise returns `{"pass": bool, "score": float, "reasons": list, "next_hint": str?}`.
- `_gate(...)`, `_context_floor_ok()`, `_role_prompt(...)`, `_score_window(...)`, `_loop_max_iterations(...)`, `_outputs_dir(...)`, `_make_completed` (test helper used inline).
- `cmd_tick(args)` — the tick flow per technical.md §7. Returns a summary dict, never raises (clean-exit on every path).
- `cmd_daemon(args)` — `time.sleep(interval)` loop bounded by `--max-iterations`. `KeyboardInterrupt` stops cleanly with a `DAEMON_STOPPED` log entry.
- `main()` — argparse with `--mode {tick,daemon}`, `--loop`, `--project`, `--interval`, `--max-iterations`, `--json`.
## R-by-R coverage
| Req | Code |
|-----|------|
| R1 entrypoint | `main()` argparse, `--mode` required choices; `cmd_tick` returns summary with `skipped:True` and `reason:"untracked"` for missing `.state.loop` |
| R2 tick flow | `cmd_tick` 7-route: load → gate → find_work → cwd → ctx-floor → spawn Implement → spawn Verify → parse verdict → cap score → spawn Orchestrate → atomic write state → tick log |
| R3 daemon | `cmd_daemon` |
| R4 harness substitution | `_substitute`, `_invoke_harness` |
| R5 context-floor guard | `_context_floor_ok` called before any harness subprocess; halts `human_intervention` on `loop_mode_eligible=False` |
| R6 idempotence | state writes only in step 10 (after parse_verdict succeeds and orchestrator ran); pre-step-10 crashes leave `.state.loop` untouched |
| R7 tests | `tests/test_loop_runner.py` (18 tests) |
| R8 out-of-scope | none — deferred to tasks 4–7 (audit work_source, backlog, worktree creation, the prompts themselves) |
## Tests (`tests/test_loop_runner.py`)
18 tests across 7 classes; all `subprocess.run` and `_run_json` calls stubbed via monkeypatch so no live LLM calls hit in CI.
- `TestEntrypoint` (2): unknown-loop exits 0; unknown-mode exits 2.
- `TestTickFlow` (5): tick-pass advances iteration_count; skip-when-halted; skip-when-untracked; skip-no-current-task; skip-when-gate-subprocess-fails.
- `TestContextFloor` (1): refuses below floor; halts `human_intervention`; implement harness never invoked.
- `TestVerifierParseFailure` (5): parse-failure halts and **does not** advance iteration_count (idempotence); fenced JSON parses; JSON with line comments parses; missing `pass` key → None; empty text → None.
- `TestScoreHistory` (1): 5 ticks with window=3 → final `score_history` length is 3 and equals `[0.4, 0.4, 0.4]`.
- `TestHarnessSubstitution` (1): custom `harness.command` with `--prompt/--cwd/--out/--artifact` tokens; verify-role invocation sees the implement role's output path as `--artifact <...-implement.json>`.
- `TestDaemonMode` (1): `--max-iterations 3` runs 3 ticks then exits 0; `time.sleep` no-op via monkeypatch.
- `TestOrchestratorOrdering` (1): implement → verify → orchestrate order observed via tagged handlers.
- `TestJsonOutput` (1): `--json` prints structured tick summary as last line; parsed via `lr.main()` + `capsys` (since `subprocess.run` is patched).
## Verification
```
python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -q # 18 passed
python3 -m pytest tests/ -q # 328 passed (was 310 + 18 new)
```