120 lines
10 KiB
Markdown
120 lines
10 KiB
Markdown
# SPEC: add-loop-runner
|
||||
|
|
|
|||
|
|
Implements `scripts/loop-runner.py --mode tick` (and `--mode daemon` opt-in). The runner is the per-tick engine that calls the brakes, spawns the three session roles (Implement / Verify / Orchestrate), parses the graded verifier verdict, and updates `.state.loop`. It is the runtime partner of the brakes layer landed in task `add-status-brakes`.
|
|||
|
|
|
|||
|
|
## Goal
|
|||
|
|
|
|||
|
|
A single Python entry point that any OS scheduler (`launchd` / `cron` / `schtasks`) or human can invoke as:
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
python3 <framework>/scripts/loop-runner.py --mode tick --loop <name> --project <p>
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
It must:
|
|||
|
|
- Be **idempotent in the failure case** -- a crash mid-tick does not advance `iteration_count` or corrupt `.state.loop`.
|
|||
|
|
- Never invoke an LLM directly. All role sessions are external subprocesses against the user's configured harness, dispatched from `loop.json` `harness.command`.
|
|||
|
|
- Refuse to run when `--check-gate` returns not-ok, and exit 0 (clean exit; do not crash the scheduler) so the OS unit's retry backoff stays calm.
|
|||
|
|
- Apply all six brake gates indirectly via `--check-gate` (no duplicated gate logic in the runner).
|
|||
|
|
|
|||
|
|
## Requirements
|
|||
|
|
|
|||
|
|
### R1 -- Entry point and CLI shape
|
|||
|
|
- `--mode {tick,daemon}` required.
|
|||
|
|
- `--loop NAME` required.
|
|||
|
|
- `--project PATH` optional (forwarded to `status.py`).
|
|||
|
|
- `--json` optional -- emit machine-readable tick summary as the last line.
|
|||
|
|
- `--interval SECONDS` for `--mode daemon` only (default: read from `loop.json` `schedule.interval_seconds`, else 3600).
|
|||
|
|
- Unknown `--mode` → exit 2.
|
|||
|
|
- Unknown loop (no `.state.loop`) → log SKIP, exit 0 (not 2; the runner never escalates a missing loop to a hard error, because the OS scheduler must keep firing).
|
|||
|
|
|
|||
|
|
### R2 -- Tick flow (per `technical.md` §7)
|
|||
|
|
|
|||
|
|
In order:
|
|||
|
|
|
|||
|
|
1. **Load**: read `.state.loop` and `loop.json` from `<loops>/<name>/`. Treat missing `.state.loop` as `untracked` SKIP (R1).
|
|||
|
|
2. **Gate**: `subprocess.run([python, status.py, "--check-gate", NAME, "--project", P, "--json"])`. Parse JSON. If `ok == false`: append `SKIP reason=…` to `.state.log`, exit 0.
|
|||
|
|
3. **Find work** (v1: only `single` work_source): `current_task = state["current_task"]`. If null: SKIP `no_current_task`. `audit` / `backlog` work_sources are stubbed for v1 (return SKIP) and fleshed out in tasks 4 and 6.
|
|||
|
|
4. **Worktree**: deferred to task `add-blast-radius-scheduler`. The runner uses `state["worktree_path"]` if set else `project_root` as cwd. If worktree configured but missing, write a `worktree_missing` warning to `.state.log` and SKIP (`human_intervention` halts are owned by `--check-gate`, not the runner).
|
|||
|
|
5. **Spawn Implement**: build harness command from `loop.json` `harness.command` with `{prompt}` = `roles.implement.prompt`, `{cwd}` = resolved cwd, `{output}` = unique artifact path under `<loop>/outputs/<tickN>-<role>.json`. Invoke via `subprocess.run`. Capture stdout. Do not block on harness timeout; v1 has no timeout (the harness owns its own timeout policy).
|
|||
|
|
6. **Spawn Verify**: same as Implement, with `{prompt}` = `roles.verify.prompt`. Add `{artifact}` substitution token (pointing at Implement's output path). Capture stdout -- **this must parse as JSON** (verdict).
|
|||
|
|
7. **Parse verdict**: accept either raw JSON or ```json fenced blocks or JSON with leading `// / #` line comments. Strict keys: `pass` (bool, required), `score` (float 0.0–1.0, required), `reasons` (list of strings, optional), `next_hint` (string, optional). On parse failure → halt as `verifier_failed`, write `HALT verifier_failed:unparseable` to `.state.log`, exit 0.
|
|||
|
|
8. **Append score**: push `verdict["score"]` to `state["score_history"]`, capped at `brakes.score_plateau_window` (drop oldest beyond window).
|
|||
|
|
9. **Spawn Orchestrate**: `{prompt}` = `roles.orchestrate.prompt`, plus inject `{verdict}` (JSON-serialized) and `{current_task}` and `{current_phase}` as substitution tokens. The orchestrator's stdout is captured but not parsed in v1 -- the orchestrator is the actor that calls `status.py --transition` / `--approve` itself (no auto-approve path).
|
|||
|
|
10. **Update state** (the runner's own writes -- never overlap with orchestrator writes):
|
|||
|
|
- `state["iteration_count"] += 1`
|
|||
|
|
- `state["last_tick_at"] = iso8601_now`
|
|||
|
|
- `state["last_verdict"] = verdict`
|
|||
|
|
- Atomic write via tmp+rename (same helper as status.py -- duplicate the small writer rather than import across scripts).
|
|||
|
|
11. **Tick log**: append `TICK pass=<bool> score=<f> iter=<N>` to `.state.log`.
|
|||
|
|
12. Exit 0.
|
|||
|
|
|
|||
|
|
Order of failure-mode Halt writes (all delegated to status.py via `_disable_schedule` best-effort, but the halt itself is a direct `.state.loop` write from the runner):
|
|||
|
|
- Parse failure → `verifier_failed` (R7 above).
|
|||
|
|
|
|||
|
|
The runner **does not** check iterations / budget / drift / task-phase gates itself -- `--check-gate` (R2 step 2) already did. The runner is responsible only for `verifier_failed` (verdict parse) and for `verifier_failed` (score plateau) indirectly via the next tick's `--check-gate`.
|
|||
|
|
|
|||
|
|
### R3 -- `--mode daemon`
|
|||
|
|
|
|||
|
|
- `time.sleep(interval)` loop calling `cmd_tick()`.
|
|||
|
|
- `KeyboardInterrupt` → exit 0 cleanly with a `DAEMON_STOPPED` log entry.
|
|||
|
|
- `--max-iterations N` (optional) caps daemon loop count. 0 / unset = unbounded.
|
|||
|
|
|
|||
|
|
### R4 -- Harness command substitution
|
|||
|
|
|
|||
|
|
`loop.json` `harness.command` is a list of strings. The runner walks each element, replacing `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}` with values from the tick context. Missing tokens stay literal (so configurations can opt out of, say, the `{output}` token by simply not including it).
|
|||
|
|
|
|||
|
|
Default `harness.command` (when `loop.json` doesn't specify one) is `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]`, matching the user's primary harness (D8 -- never inspect model capability).
|
|||
|
|
|
|||
|
|
### R5 -- Context-floor guard (D13)
|
|||
|
|
|
|||
|
|
Before invoking the Implement role, the runner calls `vram_detect.py --loop-mode --json`. If the JSON `loop_mode_eligible == false`, the runner halts the loop with `human_intervention` and writes `HALT human_intervention:context_below_floor`. Existing shell: a "context too small" loop cannot burn tokens through a harness call that would fail anyway.
|
|||
|
|
|
|||
|
|
This guard is implemented in the runner (not in `--check-gate`) because `--check-gate` is per-tick and the available-context value is hardware-state, not loop-state -- we don't want it cached in `.state.loop` between ticks.
|
|||
|
|
|
|||
|
|
### R6 -- Idempotence
|
|||
|
|
|
|||
|
|
- State writes are atomic (tmp+rename).
|
|||
|
|
- The Implement / Verify / Orchestrate invocations do not mutate state; only step 10 writes.
|
|||
|
|
- Verifier parse failure short-circuits before step 10, so a tick that fails to parse its verifier does not increment `iteration_count`. The harness retry on next tick starts from the same `current_task` and `iteration_count`.
|
|||
|
|
- A `KeyboardInterrupt` or `SIGTERM` between steps 5 and 10 leaves `.state.loop` unchanged. The harness subprocess may be left running (the runner does not own process groups in v1).
|
|||
|
|
|
|||
|
|
### R7 -- Tests (`tests/test_loop_runner.py`)
|
|||
|
|
|
|||
|
|
Required by AGENTS.md. All harness calls are stubbed via `monkeypatch.setattr(subprocess, "run", fake_run)`. No live LLM calls in CI.
|
|||
|
|
|
|||
|
|
1. `test_tick_pass` -- fixture loop with a `current_task` in `implement`, mock `--check-gate` returns ok, mock verifier returns `{"pass": true, "score": 0.9}`. Assert `iteration_count == 1`, `last_verdict["pass"] is True`, `.state.log` has `TICK pass=True score=0.9 iter=1`.
|
|||
|
|
2. `test_tick_skip_when_halted` -- pre-halt `.state.loop`, mock `--check-gate` returns not-ok. Assert `iteration_count` unchanged, `.state.log` has `SKIP reason=halted:…`.
|
|||
|
|
3. `test_tick_skip_when_untracked` -- no `.state.loop`. Assert exit 0, `.state.log` has `SKIP untracked`.
|
|||
|
|
4. `test_tick_skip_no_current_task` -- `.state.loop` has `current_task: null`. Assert SKIP `no_current_task`.
|
|||
|
|
5. `test_verifier_parse_failure_halts` -- mock verifier returns garbage. Assert loop halted as `verifier_failed`, `last_verdict` is null, `iteration_count` **unchanged** (R6 idempotence).
|
|||
|
|
6. `test_score_history_capped` -- loop with `score_plateau_window: 3`, run 5 ticks with mock verifier returning scores 0.5, 0.4, 0.4, 0.4, 0.4. Assert `score_history` length is 3 (the last three).
|
|||
|
|
7. `test_json_output_mode` -- `--json` prints a structured tick summary on the last line.
|
|||
|
|
8. `test_daemon_mode_runs_n_iterations` -- `--mode daemon --max-iterations 3` runs `cmd_tick` three times then exits 0.
|
|||
|
|
9. `test_context_floor_refuses` -- mock `vram_detect.py` returns `loop_mode_eligible: false`. Assert loop halted `human_intervention`, harness subprocess never invoked.
|
|||
|
|
10. `test_unknown_mode_rejected` -- `--mode bogus` exits 2.
|
|||
|
|
11. `test_unknown_loop_skip_clean_exit` -- `--loop ghost` exits 0 (R1).
|
|||
|
|
12. `test_harness_command_substitution` -- fixture loop.json with custom `harness.command` containing `{prompt}`, `{cwd}`, `{output}`. Assert stub `subprocess.run` saw the substituted values verbatim.
|
|||
|
|
13. `test_orchestrator_invoked_after_verifier` -- assert subprocess invocations happen in order: gate → implement → verify → orchestrate. Capture argv patterns to confirm.
|
|||
|
|
|
|||
|
|
### R8 -- Out of scope (other tasks)
|
|||
|
|
|
|||
|
|
- Live harness adapter -- provided by user as `harness.command`; no new adapter code.
|
|||
|
|
- `audit` work_source -- task 4 (goal-mode / verifier session) and task 6 (self-improvement template).
|
|||
|
|
- `backlog` work_source -- task 7 (self-improvement loop) and `design/<area>/BACKLOG.md` integration.
|
|||
|
|
- Worktree creation plumbing -- task `add-blast-radius-scheduler`.
|
|||
|
|
- Verifier prompt (`loop-verifier.md`) -- task 6. The runner just reads the filename from `loop.json` and passes it to the harness; it does not parse the prompt itself.
|
|||
|
|
- Orchestrator logic that decides phase transitions -- the orchestrator role does that; the runner only spawns the orchestrator and trusts its `status.py` calls.
|
|||
|
|
|
|||
|
|
## Approach
|
|||
|
|
|
|||
|
|
Single new file `scripts/loop-runner.py`. Stdlib-only (no new pip deps). Reuses small helpers (`_read_state_loop`, `_write_state_loop`, `_loop_dir`, `_read_loop_config`) duplicated inline rather than imported from status.py -- keeps the scripts independent (no risk of one script accidentally being moved/broken affecting the other). The duplicate helpers are <30 lines total.
|
|||
|
|
|
|||
|
|
Tests file `tests/test_loop_runner.py` uses `tmp_path` + a `_stub_subprocess` helper that pattern-matches on argv to return canned outputs.
|
|||
|
|
|
|||
|
|
## Verification
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
python3 -m py_compile scripts/loop-runner.py
|
|||
|
|
python3 -m pytest tests/test_loop_runner.py -v
|
|||
|
|
python3 -m pytest tests/ -q # ensure no regressions
|
|||
|
|
```
|