Files
automaton/tasks/complete/add-loop-runner/SPEC.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

120 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPEC: add-loop-runner
Implements `scripts/loop-runner.py --mode tick` (and `--mode daemon` opt-in). The runner is the per-tick engine that calls the brakes, spawns the three session roles (Implement / Verify / Orchestrate), parses the graded verifier verdict, and updates `.state.loop`. It is the runtime partner of the brakes layer landed in task `add-status-brakes`.
## Goal
A single Python entry point that any OS scheduler (`launchd` / `cron` / `schtasks`) or human can invoke as:
```
python3 <framework>/scripts/loop-runner.py --mode tick --loop <name> --project <p>
```
It must:
- Be **idempotent in the failure case** -- a crash mid-tick does not advance `iteration_count` or corrupt `.state.loop`.
- Never invoke an LLM directly. All role sessions are external subprocesses against the user's configured harness, dispatched from `loop.json` `harness.command`.
- Refuse to run when `--check-gate` returns not-ok, and exit 0 (clean exit; do not crash the scheduler) so the OS unit's retry backoff stays calm.
- Apply all six brake gates indirectly via `--check-gate` (no duplicated gate logic in the runner).
## Requirements
### R1 -- Entry point and CLI shape
- `--mode {tick,daemon}` required.
- `--loop NAME` required.
- `--project PATH` optional (forwarded to `status.py`).
- `--json` optional -- emit machine-readable tick summary as the last line.
- `--interval SECONDS` for `--mode daemon` only (default: read from `loop.json` `schedule.interval_seconds`, else 3600).
- Unknown `--mode` → exit 2.
- Unknown loop (no `.state.loop`) → log SKIP, exit 0 (not 2; the runner never escalates a missing loop to a hard error, because the OS scheduler must keep firing).
### R2 -- Tick flow (per `technical.md` §7)
In order:
1. **Load**: read `.state.loop` and `loop.json` from `<loops>/<name>/`. Treat missing `.state.loop` as `untracked` SKIP (R1).
2. **Gate**: `subprocess.run([python, status.py, "--check-gate", NAME, "--project", P, "--json"])`. Parse JSON. If `ok == false`: append `SKIP reason=…` to `.state.log`, exit 0.
3. **Find work** (v1: only `single` work_source): `current_task = state["current_task"]`. If null: SKIP `no_current_task`. `audit` / `backlog` work_sources are stubbed for v1 (return SKIP) and fleshed out in tasks 4 and 6.
4. **Worktree**: deferred to task `add-blast-radius-scheduler`. The runner uses `state["worktree_path"]` if set else `project_root` as cwd. If worktree configured but missing, write a `worktree_missing` warning to `.state.log` and SKIP (`human_intervention` halts are owned by `--check-gate`, not the runner).
5. **Spawn Implement**: build harness command from `loop.json` `harness.command` with `{prompt}` = `roles.implement.prompt`, `{cwd}` = resolved cwd, `{output}` = unique artifact path under `<loop>/outputs/<tickN>-<role>.json`. Invoke via `subprocess.run`. Capture stdout. Do not block on harness timeout; v1 has no timeout (the harness owns its own timeout policy).
6. **Spawn Verify**: same as Implement, with `{prompt}` = `roles.verify.prompt`. Add `{artifact}` substitution token (pointing at Implement's output path). Capture stdout -- **this must parse as JSON** (verdict).
7. **Parse verdict**: accept either raw JSON or ```json fenced blocks or JSON with leading `// / #` line comments. Strict keys: `pass` (bool, required), `score` (float 0.0–1.0, required), `reasons` (list of strings, optional), `next_hint` (string, optional). On parse failure → halt as `verifier_failed`, write `HALT verifier_failed:unparseable` to `.state.log`, exit 0.
8. **Append score**: push `verdict["score"]` to `state["score_history"]`, capped at `brakes.score_plateau_window` (drop oldest beyond window).
9. **Spawn Orchestrate**: `{prompt}` = `roles.orchestrate.prompt`, plus inject `{verdict}` (JSON-serialized) and `{current_task}` and `{current_phase}` as substitution tokens. The orchestrator's stdout is captured but not parsed in v1 -- the orchestrator is the actor that calls `status.py --transition` / `--approve` itself (no auto-approve path).
10. **Update state** (the runner's own writes -- never overlap with orchestrator writes):
- `state["iteration_count"] += 1`
- `state["last_tick_at"] = iso8601_now`
- `state["last_verdict"] = verdict`
- Atomic write via tmp+rename (same helper as status.py -- duplicate the small writer rather than import across scripts).
11. **Tick log**: append `TICK pass=<bool> score=<f> iter=<N>` to `.state.log`.
12. Exit 0.
Order of failure-mode Halt writes (all delegated to status.py via `_disable_schedule` best-effort, but the halt itself is a direct `.state.loop` write from the runner):
- Parse failure → `verifier_failed` (R7 above).
The runner **does not** check iterations / budget / drift / task-phase gates itself -- `--check-gate` (R2 step 2) already did. The runner is responsible only for `verifier_failed` (verdict parse) and for `verifier_failed` (score plateau) indirectly via the next tick's `--check-gate`.
### R3 -- `--mode daemon`
- `time.sleep(interval)` loop calling `cmd_tick()`.
- `KeyboardInterrupt` → exit 0 cleanly with a `DAEMON_STOPPED` log entry.
- `--max-iterations N` (optional) caps daemon loop count. 0 / unset = unbounded.
### R4 -- Harness command substitution
`loop.json` `harness.command` is a list of strings. The runner walks each element, replacing `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}` with values from the tick context. Missing tokens stay literal (so configurations can opt out of, say, the `{output}` token by simply not including it).
Default `harness.command` (when `loop.json` doesn't specify one) is `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]`, matching the user's primary harness (D8 -- never inspect model capability).
### R5 -- Context-floor guard (D13)
Before invoking the Implement role, the runner calls `vram_detect.py --loop-mode --json`. If the JSON `loop_mode_eligible == false`, the runner halts the loop with `human_intervention` and writes `HALT human_intervention:context_below_floor`. Existing shell: a "context too small" loop cannot burn tokens through a harness call that would fail anyway.
This guard is implemented in the runner (not in `--check-gate`) because `--check-gate` is per-tick and the available-context value is hardware-state, not loop-state -- we don't want it cached in `.state.loop` between ticks.
### R6 -- Idempotence
- State writes are atomic (tmp+rename).
- The Implement / Verify / Orchestrate invocations do not mutate state; only step 10 writes.
- Verifier parse failure short-circuits before step 10, so a tick that fails to parse its verifier does not increment `iteration_count`. The harness retry on next tick starts from the same `current_task` and `iteration_count`.
- A `KeyboardInterrupt` or `SIGTERM` between steps 5 and 10 leaves `.state.loop` unchanged. The harness subprocess may be left running (the runner does not own process groups in v1).
### R7 -- Tests (`tests/test_loop_runner.py`)
Required by AGENTS.md. All harness calls are stubbed via `monkeypatch.setattr(subprocess, "run", fake_run)`. No live LLM calls in CI.
1. `test_tick_pass` -- fixture loop with a `current_task` in `implement`, mock `--check-gate` returns ok, mock verifier returns `{"pass": true, "score": 0.9}`. Assert `iteration_count == 1`, `last_verdict["pass"] is True`, `.state.log` has `TICK pass=True score=0.9 iter=1`.
2. `test_tick_skip_when_halted` -- pre-halt `.state.loop`, mock `--check-gate` returns not-ok. Assert `iteration_count` unchanged, `.state.log` has `SKIP reason=halted:…`.
3. `test_tick_skip_when_untracked` -- no `.state.loop`. Assert exit 0, `.state.log` has `SKIP untracked`.
4. `test_tick_skip_no_current_task` -- `.state.loop` has `current_task: null`. Assert SKIP `no_current_task`.
5. `test_verifier_parse_failure_halts` -- mock verifier returns garbage. Assert loop halted as `verifier_failed`, `last_verdict` is null, `iteration_count` **unchanged** (R6 idempotence).
6. `test_score_history_capped` -- loop with `score_plateau_window: 3`, run 5 ticks with mock verifier returning scores 0.5, 0.4, 0.4, 0.4, 0.4. Assert `score_history` length is 3 (the last three).
7. `test_json_output_mode` -- `--json` prints a structured tick summary on the last line.
8. `test_daemon_mode_runs_n_iterations` -- `--mode daemon --max-iterations 3` runs `cmd_tick` three times then exits 0.
9. `test_context_floor_refuses` -- mock `vram_detect.py` returns `loop_mode_eligible: false`. Assert loop halted `human_intervention`, harness subprocess never invoked.
10. `test_unknown_mode_rejected` -- `--mode bogus` exits 2.
11. `test_unknown_loop_skip_clean_exit` -- `--loop ghost` exits 0 (R1).
12. `test_harness_command_substitution` -- fixture loop.json with custom `harness.command` containing `{prompt}`, `{cwd}`, `{output}`. Assert stub `subprocess.run` saw the substituted values verbatim.
13. `test_orchestrator_invoked_after_verifier` -- assert subprocess invocations happen in order: gate → implement → verify → orchestrate. Capture argv patterns to confirm.
### R8 -- Out of scope (other tasks)
- Live harness adapter -- provided by user as `harness.command`; no new adapter code.
- `audit` work_source -- task 4 (goal-mode / verifier session) and task 6 (self-improvement template).
- `backlog` work_source -- task 7 (self-improvement loop) and `design/<area>/BACKLOG.md` integration.
- Worktree creation plumbing -- task `add-blast-radius-scheduler`.
- Verifier prompt (`loop-verifier.md`) -- task 6. The runner just reads the filename from `loop.json` and passes it to the harness; it does not parse the prompt itself.
- Orchestrator logic that decides phase transitions -- the orchestrator role does that; the runner only spawns the orchestrator and trusts its `status.py` calls.
## Approach
Single new file `scripts/loop-runner.py`. Stdlib-only (no new pip deps). Reuses small helpers (`_read_state_loop`, `_write_state_loop`, `_loop_dir`, `_read_loop_config`) duplicated inline rather than imported from status.py -- keeps the scripts independent (no risk of one script accidentally being moved/broken affecting the other). The duplicate helpers are <30 lines total.
Tests file `tests/test_loop_runner.py` uses `tmp_path` + a `_stub_subprocess` helper that pattern-matches on argv to return canned outputs.
## Verification
```
python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -v
python3 -m pytest tests/ -q # ensure no regressions
```