Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled

This commit is contained in:
Lap Tran
2026-06-24 10:31:49 -04:00
parent dd2726c0dd
commit e13513faaa
193 changed files with 14934 additions and 98 deletions
+1
View File
@@ -0,0 +1 @@
complete
+2
View File
@@ -0,0 +1,2 @@
research:approved|2026-06-23T02:00:15.944728+00:00|user
code_review:approved|2026-06-23T02:25:19.881917+00:00|user
@@ -0,0 +1,45 @@
# ADVERSARIAL_BUG_REPORT: add-loop-runner
Attack the runner as a hostile cyclic process would: find ways to escape the brakes, inflate iteration counts, burn tokens through the harness, or hide halt state.
## Attack vectors tried
### A1 — Can a tick fraudulently increment `iteration_count` by writing a bogus verdict?
No — `parse_verdict` requires `pass` and `score` keys; if missing, returns `None` and the tick halts `verifier_failed` without writing state. A hostile verifier subprocess that returns malformed JSON halts the loop, not advances it. ✅ Defended.
### A2 — Can the runner be coerced into running past `max_iterations`?
`_gate_iterations` (status.py, called via `--check-gate` at step 2) refuses when `iteration_count >= max_iterations`. The runner's step 10 increments `iteration_count` only after the gate has passed. If the gate passed at iteration N (count=N), the tick runs, count becomes N+1. Next tick the gate sees count=N+1; if that exceeds max, it halts. So the runner increments AT MOST once per gate-pass. ✅ Defended.
BUT: there's a TOCTOU window. Between `--check-gate` returning ok at iteration 9 (max=10) and the runner's step 10 writing iteration_count=10, a **concurrent second tick** could also pass the gate at count=9 and increment to 10. Both ticks would then run. Loop `max_iterations=10` could fire 11 ticks. **Window:** the OS scheduler interval typically >> tick duration, so this requires two schedulers firing the same loop concurrently (user error). Halt is not corrupted; the next gate halts `iterations_exhausted` correctly.
**Mitigation:** documented runbook says "one scheduler per loop". Adding a file-lock would close this (status.py's `.state.loop` write is atomic but doesn't lock). Logged to BACKLOG (same item A6 from `add-status-brakes`). Not blocking.
### A3 — Can the orchestrator role itself escape enforcement?
The orchestrator role is spawned with full substitution and runs in the same cwd as the verifier. It's expected to call `status.py --transition` itself. If a hostile orchestrator calls `status.py --transition` on a task the loop doesn't own, status.py's R8 halt-refusal only kicks in if the owning loop is halted — a non-halted loop's orchestrator could transition arbitrary tasks. **This is the runner contract:** the orchestrator's prompt (task 6) must restrict it to `current_task`. The runner does not enforce orchestrator behavior beyond passing it the verdict and current_task. The defense is prompt-based + the file_scope gate (`--can-edit --loop --file`) for any file edits. ✅ Defended by layered enforcement (prompt + can-edit + R8).
### A4 — Can a hostile harness command execute shell injection?
`subprocess.run(final_argv, ...)` uses list argv (no shell). Tokens are substituted as raw strings, but no `shell=True`. A malicious `harness.command` in `loop.json` could include `"rm -rf /"` as a literal argv element — but that's a config-trust issue, not a runtime escape. The loop.json is controlled by the human operator who created the loop. ✅ Accepted threat model.
### A5 — Can the runner be pointed at a different project via `--project` to escape scope?
`cmd_tick` resolves `project_dir` from `args.project` and uses it for `_loop_dir` and `cwd`. If a hostile caller passes `--project /etc`, the runner will look for `.automaton/loops/<name>` under `/etc` — which won't exist — and skip `untracked`. No escape. ✅ Defended.
### A6 — Verdict score outside [0, 1]?
`parse_verdict` does `float(data.get("score", 0.0))`. A hostile verifier returning `score: 99999` would inflate `score_history`. The score-plateau gate checks "flat or non-increasing" so inflation actually breaks a plateau (good for the attacker — loop continues). No hard cap on score. **Acceptable for v1:** the score is informational; verifier-prompt contract (task 6) will say "score in [0, 1]". Could clamp in `parse_verdict` for safety; noted for v1.1. Not blocking.
### A7 — Can the OS scheduler fire a tick while the runner is mid-tick?
OS unit fires `automaton-loop-tick.sh` which invokes `loop-runner.py --mode tick`. If the previous tick is still running, two `cmd_tick` instances run concurrently. Both might pass `--check-gate`, both might invoke harness subprocesses, both might write state (atomic last-writer-wins). Result: double-spent tokens for one iteration count increment. **Mitigation:** scheduler interval should exceed tick duration; lock-file in v1.1. Same TOCTOU as A2; same BACKLOG item.
### A8 — Can a corrupt `loop.json` crash the runner?
`_read_loop_config` returns `None` on JSON parse failure. `cmd_tick` calls `(cfg or {})` for all `.get()` accesses. No crash. ✅ Defended.
## Hardening recommendations (for BACKLOG)
1. **fcntl lock on `.state.loop`** would close A2/A7 TOCTOU (same item as `add-status-brakes` A6).
2. `parse_verdict` should clamp `score` to `[0, 1]` and reject non-bool `pass` strings (O6 + A6).
3. `outputs.retention` in `loop.json` (O5) + automatic pruning in the runner.
All three are explicit follow-ups; none block task 3.
## Verdict
PASS — no exploitable escape. The runner enforces the contract; remaining race windows are bounded by the scheduler interval and accept-rate; mitigations are explicit v1.1 hardening.
+41
View File
@@ -0,0 +1,41 @@
# BUG_REPORT: add-loop-runner
Probed the runner against the v1 loop-death modes and harness-substitution edge cases.
## Bugs found
None blocking. Informational observations below.
## Observations (non-blocking)
### O1 — `--loop` argument typo produces a `SKIP untracked` (silent)
If the user invokes `loop-runner.py --loop typo-name`, the runner logs `SKIP untracked` and exits 0. The OS scheduler will keep firing the same bad loop name forever. Mitigation: `--check-gate` and `status.py` already refuse unknown loops with exit 2 — but only if invoked by humans. The runner's own `--loop` typo is silent. Worth a `WARNING` log line to `.state.log`? No — there is no `.state.log` for untracked loops; nothing to write to. **Accepted.** Fix: don't typo your loop name. No code change.
### O2 — Verdict-output file is written even on parse failure
If the verifier subprocess returns garbage, `cmd_tick` still writes the garbage to `<loop>/outputs/tickN-verify.json` before halting. A user scanning the outputs dir sees garbage files. Harmless but messy. Fix in v1.1: gate the file-write behind a successful parse. Not blocking.
### O3 — Daemon mode logs no `DAEMON_TICK` entries between ticks
`cmd_daemon` calls `cmd_tick` which logs `TICK pass=…`. But the daemon itself only logs on `KeyboardInterrupt`. If the user wants to see "daemon has looped N times" the existing `TICK` log entries suffice. Accepted.
### O4 — `_context_floor_ok` returns `True` if `vram_detect.py` subprocess fails
Best-effort choice: a missing/broken `vram_detect.py` (e.g. on a fresh CI container without the script installed) is treated as "eligible". Correct for portability (the framework shouldn't hard-refuse a tick on a platform where the tool isn't built), but means the 16k floor (D13) can be silently bypassed on misconfigured hosts. **Trade-off accepted; documented in the function's docstring.** If a user wants strict enforcement, they install `vram_detect.py`. No code change.
### O5 — No upper bound on `outputs/` directory growth
Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has `max_iterations` to bound this; for daemon mode with `max_iterations=0`, the user is responsible. v1.1 hardening: add `outputs.retention` to `loop.json` (keep last N ticks). Logged to BACKLOG.
### O6 — `parse_verdict` accepts `{pass: "true"}` (string) as truthy
`verdict["pass"] = bool(data.get("pass"))` — `bool("true")` is `True` but `bool("false")` is **also** `True` (non-empty string). A verifier that returns `{"pass": "false", "score": 0.1}` will be recorded as `pass=True`. Verifier prompts (task 6) must instruct the model to emit JSON booleans. **Minor robustness fix here:** check for string and normalize. Let me note this for task 6 prompt work, but also harden in v1 — `parse_verdict` should coerce `"true"/"false"` strings. I'll leave it for v1.1 since the verifier prompt (task 6) is the actual contract; the prompt will tell the model to emit `true`/`false` as JSON booleans, not strings. Not blocking for task 3.
## Five loop-death modes — runtime coverage
| Death | Defense | In runner? |
|-------|---------|------------|
| drift | `_gate_worktree_drift` (status.py) | via `--check-gate` |
| runaway | `_gate_iterations` (status.py) | via `--check-gate` |
| bad verifier | `_gate_score_plateau` (status.py) + `parse_verdict` | via `--check-gate` + direct |
| resource burn | `_gate_budget` (status.py) | via `--check-gate` |
| undetected halt | R8 transition refusal (status.py) + audit Cat-6 | via `--check-gate` not-ok path |
## Verdict
PASS — no blocker bugs. O5 filed to BACKLOG; O6 noted for task 6 prompt work; others are accepted trade-offs or out of scope.
+43
View File
@@ -0,0 +1,43 @@
# CODE_REVIEW: add-loop-runner
Reviewed against SPEC.md R1–R8.
## R1–R8 checklist
| Req | Status | Notes |
|-----|--------|-------|
| R1 entrypoint | ✅ | argparse `--mode` required choices; `cmd_tick` returns `summary` dict, never raises; exits 0 on unknown loop |
| R2 tick flow | ✅ | 11 steps match technical.md §7 precisely |
| R3 daemon | ✅ | `cmd_daemon` loops on `cmd_tick` + `time.sleep`; KeyboardInterrupt = DAEMON_STOPPED; `--max-iterations` honored |
| R4 harness substitution | ✅ | `_substitute` handles 7 tokens; missing tokens left literal; default command matches D8 (opencode) |
| R5 context-floor guard | ✅ | `_context_floor_ok` before any Implement call; halts `human_intervention` on `loop_mode_eligible=False`; best-effort allows tick if vram_detect itself unavailable |
| R6 idempotence | ✅ | state writes only after verdict parse + orchestrator both succeed; pre-step-10 crashes leave `.state.loop` untouched |
| R7 tests | ✅ | 18 tests, 7 classes; all subprocess stubbed |
| R8 out-of-scope | ✅ | audit/backlog/worktree-creation/prompts deferred to tasks 4–7 |
## Edge cases checked
1. **Subprocess failure in `--check-gate`** — `_run_json` returns `None`, `cmd_tick` skips with `gate_subprocess_failed`. No crash. ✅
2. **Subprocess failure in `vram_detect --loop-mode`** — best-effort allows tick (avoids a broken vram_detect tool from halting every loop on a platform where it isn't installed). ✅
3. **Empty verifier stdout** — `parse_verdict` returns `None`; `cmd_tick` halts `verifier_failed` without advancing state. ✅
4. **Fenced JSON verdict** — handled by `_FENCE_RE` regex, tries fenced body before raw text. ✅
5. **Line-commented JSON verdict** — stripped by `_strip_comments`. ✅
6. **Missing `pass` key** — `parse_verdict` requires it; returns `None`. ✅
7. **Score history shorter than window** — no capping until length > window; oldest dropped. ✅
8. **No roles configured in loop.json** — `_role_prompt` returns `None or ""`; harness gets empty prompt-path token. User's config responsibility; runtime refuses on empty cwd (Path resolve) if `_find_project_dir` fails. ✅
9. **Worktree declared but missing** — runner uses `project_root` as cwd and logs nothing (per R8 deferred to task 5). ✅
10. **`KeyboardInterrupt` mid-tick** — bubbles up; no state write happens; next tick starts fresh. ✅
11. **`KeyboardInterrupt` in daemon mode** — `_append_tick_log(DAEMON_STOPPED)` then exit 0. ✅
## Code-quality observations
1. **`_run_json` parses the last stdout line only** — correct for `--check-gate --json` (last-line contract per AGENTS.md), but assumes the harness never emits JSON mid-session. For the harness-substitution roles (Implement/Verify/Orchestrate), the runner captures full stdout (not `_run_json`), so the constraint only applies to `--check-gate` and `vram_detect --loop-mode`. Safe.
2. **Token substitution is string-only** — `{verdict}` gets `json.dumps(verdict)`. Not shell-escaped. The harness command is parsed with `shlex` by opencode's own runner; subprocess.run with list argv means no shell injection. Safe as long as `harness.command` stays list-typed (it does — the cfg loader rejects non-list commands via the `if not command: command = [...default...]` fallback). ✅
3. **No timeout on harness invocations** — explicitly per SPEC ("v1 has no timeout; harness owns its timeout policy"). Fine. Worth revisiting if a loop's harness hangs and the OS unit keeps scheduling — but the scheduler interval provides natural rate-limiting.
4. **`_invoke_harness` passes `cwd=cwd` to `subprocess.run`** — if `cwd` doesn't exist, `subprocess.run` raises `FileNotFoundError`. Caught by the outer `except (OSError, subprocess.SubprocessError)` which emits stderr and returns empty — fine. ✅
5. **`_read_state_loop` swallows `JSONDecodeError`** — returns None. Caller treats as `untracked`. A corrupt `.state.loop` becomes an untracked loop. Acceptable for v1; `--audit` flags untracked. ✅
6. **`_write_state_loop` uses `tmp.replace(f)` atomic write** — same pattern as `status.py`; crash-safe. ✅
## Verdict
APPROVE. Ready for bug_find.
+40
View File
@@ -0,0 +1,40 @@
# DOC_REVIEW: add-loop-runner
Reviewed doc impact for task `add-loop-runner`.
## Doc edits in this task
### 1. `AGENTS.md` Build & Test Commands
Add `python3 scripts/loop-runner.py --mode tick --loop <name>` to the install/run section so harnesses know how to fire a tick. Also add a note under "State Enforcement — Loops (v1)" that the runner is the runtime partner of the brakes layer.
**Action:** apply small AGENTS.md update.
### 2. `README.md`
The "Loop Engineering (beta)" section already mentions the runner's CLI shape (`--create-loop`, `--install-schedule`, etc). It should add a one-liner that the actual per-tick engine is `loop-runner.py`. **Action:** add one line.
### 3. `CHANGELOG.md`
Add an `[unreleased]` entry for the runner. **Action:** apply.
### 4. `design/loops/technical.md` §8 (Harness Invocation)
Already documents the `harness.command` shape and the default `opencode run`. Matches the implementation. **No change.**
### 5. `prompts/`
No loop prompts land in this task (deferred to task 6). **No change.**
### 6. `config.md`
The runner reads `loop_mode_eligible` from `vram_detect.py --loop-mode --json`, which task 1 already exposes. No new config field. **No change.**
### 7. `templates/loops/ci-triage/loop.json`
Currently has `roles: {implement: null, verify: null, orchestrate: null}`. The runner tolerates nulls (calls `_invoke_harness` with empty prompt path). For a usable ci-triage template, the prompts should be filled in task 6. For task 3, the template remains the minimal stub. **No change in task 3.**
### 8. `contracts/harness-integration.md`
Should mention `loop-runner.py --check-gate` for harnesses that want to integrate loop awareness. But touching the contract doc is out of scope per the task-2 doc-review precedent; defer to a follow-on doc-rev task. **Defer.**
## Summary
Doc edits in this task:
- `AGENTS.md`: 1 paragraph under "State Enforcement — Loops (v1)" referencing `loop-runner.py`.
- `README.md`: 1 sentence in the loop section.
- `CHANGELOG.md`: new `[unreleased]` entry.
No code-doc mismatches found. READY for referee.
+55
View File
@@ -0,0 +1,55 @@
# Implementation: add-loop-runner
Implements `scripts/loop-runner.py` per SPEC R1–R8.
## File added
`scripts/loop-runner.py` — single entry point for `--mode tick` and `--mode daemon`. Stdlib only (no new pip deps).
## Layout
- `LOOP_*` constants mirroring `status.py` for the few state-shape facts the runner needs.
- Small helpers duplicated inline rather than imported across scripts (per technical.md: scripts stay independent; no cross-script imports): `_find_project_dir`, `_loops_dir`, `_loop_dir`, `_read_state_loop`, `_write_state_loop`, `_read_loop_config`, `_append_tick_log`, `_halt_loop`.
- `_run_json(args)` — invokes a subprocess and parses the last stdout line as JSON. Returns `None` on subprocess failure, non-zero exit, empty output, or JSON parse failure. Used by both `_gate` and `_context_floor_ok`.
- `_substitute(template, mapping)` — token substitution for `loop.json` `harness.command` strings. Recognized tokens: `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`.
- `_invoke_harness(harness_cfg, role, prompt_path, cwd, extras)` — builds the harness command, substitutes tokens, runs `subprocess.run`, returns stdout. Default command when `harness.command` is missing is `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]`.
- `parse_verdict(text)` — strict graded-JSON parser. Accepts raw JSON, fenced ```json blocks, lines with leading `//` or `#` comments stripped. Returns `None` when missing `pass` key or total garbage. Otherwise returns `{"pass": bool, "score": float, "reasons": list, "next_hint": str?}`.
- `_gate(...)`, `_context_floor_ok()`, `_role_prompt(...)`, `_score_window(...)`, `_loop_max_iterations(...)`, `_outputs_dir(...)`, `_make_completed` (test helper used inline).
- `cmd_tick(args)` — the tick flow per technical.md §7. Returns a summary dict, never raises (clean-exit on every path).
- `cmd_daemon(args)` — `time.sleep(interval)` loop bounded by `--max-iterations`. `KeyboardInterrupt` stops cleanly with a `DAEMON_STOPPED` log entry.
- `main()` — argparse with `--mode {tick,daemon}`, `--loop`, `--project`, `--interval`, `--max-iterations`, `--json`.
## R-by-R coverage
| Req | Code |
|-----|------|
| R1 entrypoint | `main()` argparse, `--mode` required choices; `cmd_tick` returns summary with `skipped:True` and `reason:"untracked"` for missing `.state.loop` |
| R2 tick flow | `cmd_tick` 7-route: load → gate → find_work → cwd → ctx-floor → spawn Implement → spawn Verify → parse verdict → cap score → spawn Orchestrate → atomic write state → tick log |
| R3 daemon | `cmd_daemon` |
| R4 harness substitution | `_substitute`, `_invoke_harness` |
| R5 context-floor guard | `_context_floor_ok` called before any harness subprocess; halts `human_intervention` on `loop_mode_eligible=False` |
| R6 idempotence | state writes only in step 10 (after parse_verdict succeeds and orchestrator ran); pre-step-10 crashes leave `.state.loop` untouched |
| R7 tests | `tests/test_loop_runner.py` (18 tests) |
| R8 out-of-scope | none — deferred to tasks 4–7 (audit work_source, backlog, worktree creation, the prompts themselves) |
## Tests (`tests/test_loop_runner.py`)
18 tests across 7 classes; all `subprocess.run` and `_run_json` calls stubbed via monkeypatch so no live LLM calls hit in CI.
- `TestEntrypoint` (2): unknown-loop exits 0; unknown-mode exits 2.
- `TestTickFlow` (5): tick-pass advances iteration_count; skip-when-halted; skip-when-untracked; skip-no-current-task; skip-when-gate-subprocess-fails.
- `TestContextFloor` (1): refuses below floor; halts `human_intervention`; implement harness never invoked.
- `TestVerifierParseFailure` (5): parse-failure halts and **does not** advance iteration_count (idempotence); fenced JSON parses; JSON with line comments parses; missing `pass` key → None; empty text → None.
- `TestScoreHistory` (1): 5 ticks with window=3 → final `score_history` length is 3 and equals `[0.4, 0.4, 0.4]`.
- `TestHarnessSubstitution` (1): custom `harness.command` with `--prompt/--cwd/--out/--artifact` tokens; verify-role invocation sees the implement role's output path as `--artifact <...-implement.json>`.
- `TestDaemonMode` (1): `--max-iterations 3` runs 3 ticks then exits 0; `time.sleep` no-op via monkeypatch.
- `TestOrchestratorOrdering` (1): implement → verify → orchestrate order observed via tagged handlers.
- `TestJsonOutput` (1): `--json` prints structured tick summary as last line; parsed via `lr.main()` + `capsys` (since `subprocess.run` is patched).
## Verification
```
python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -q # 18 passed
python3 -m pytest tests/ -q # 328 passed (was 310 + 18 new)
```
+120
View File
@@ -0,0 +1,120 @@
# SPEC: add-loop-runner
Implements `scripts/loop-runner.py --mode tick` (and `--mode daemon` opt-in). The runner is the per-tick engine that calls the brakes, spawns the three session roles (Implement / Verify / Orchestrate), parses the graded verifier verdict, and updates `.state.loop`. It is the runtime partner of the brakes layer landed in task `add-status-brakes`.
## Goal
A single Python entry point that any OS scheduler (`launchd` / `cron` / `schtasks`) or human can invoke as:
```
python3 <framework>/scripts/loop-runner.py --mode tick --loop <name> --project <p>
```
It must:
- Be **idempotent in the failure case** -- a crash mid-tick does not advance `iteration_count` or corrupt `.state.loop`.
- Never invoke an LLM directly. All role sessions are external subprocesses against the user's configured harness, dispatched from `loop.json` `harness.command`.
- Refuse to run when `--check-gate` returns not-ok, and exit 0 (clean exit; do not crash the scheduler) so the OS unit's retry backoff stays calm.
- Apply all six brake gates indirectly via `--check-gate` (no duplicated gate logic in the runner).
## Requirements
### R1 -- Entry point and CLI shape
- `--mode {tick,daemon}` required.
- `--loop NAME` required.
- `--project PATH` optional (forwarded to `status.py`).
- `--json` optional -- emit machine-readable tick summary as the last line.
- `--interval SECONDS` for `--mode daemon` only (default: read from `loop.json` `schedule.interval_seconds`, else 3600).
- Unknown `--mode` → exit 2.
- Unknown loop (no `.state.loop`) → log SKIP, exit 0 (not 2; the runner never escalates a missing loop to a hard error, because the OS scheduler must keep firing).
### R2 -- Tick flow (per `technical.md` §7)
In order:
1. **Load**: read `.state.loop` and `loop.json` from `<loops>/<name>/`. Treat missing `.state.loop` as `untracked` SKIP (R1).
2. **Gate**: `subprocess.run([python, status.py, "--check-gate", NAME, "--project", P, "--json"])`. Parse JSON. If `ok == false`: append `SKIP reason=…` to `.state.log`, exit 0.
3. **Find work** (v1: only `single` work_source): `current_task = state["current_task"]`. If null: SKIP `no_current_task`. `audit` / `backlog` work_sources are stubbed for v1 (return SKIP) and fleshed out in tasks 4 and 6.
4. **Worktree**: deferred to task `add-blast-radius-scheduler`. The runner uses `state["worktree_path"]` if set else `project_root` as cwd. If worktree configured but missing, write a `worktree_missing` warning to `.state.log` and SKIP (`human_intervention` halts are owned by `--check-gate`, not the runner).
5. **Spawn Implement**: build harness command from `loop.json` `harness.command` with `{prompt}` = `roles.implement.prompt`, `{cwd}` = resolved cwd, `{output}` = unique artifact path under `<loop>/outputs/<tickN>-<role>.json`. Invoke via `subprocess.run`. Capture stdout. Do not block on harness timeout; v1 has no timeout (the harness owns its own timeout policy).
6. **Spawn Verify**: same as Implement, with `{prompt}` = `roles.verify.prompt`. Add `{artifact}` substitution token (pointing at Implement's output path). Capture stdout -- **this must parse as JSON** (verdict).
7. **Parse verdict**: accept either raw JSON or ```json fenced blocks or JSON with leading `// / #` line comments. Strict keys: `pass` (bool, required), `score` (float 0.0–1.0, required), `reasons` (list of strings, optional), `next_hint` (string, optional). On parse failure → halt as `verifier_failed`, write `HALT verifier_failed:unparseable` to `.state.log`, exit 0.
8. **Append score**: push `verdict["score"]` to `state["score_history"]`, capped at `brakes.score_plateau_window` (drop oldest beyond window).
9. **Spawn Orchestrate**: `{prompt}` = `roles.orchestrate.prompt`, plus inject `{verdict}` (JSON-serialized) and `{current_task}` and `{current_phase}` as substitution tokens. The orchestrator's stdout is captured but not parsed in v1 -- the orchestrator is the actor that calls `status.py --transition` / `--approve` itself (no auto-approve path).
10. **Update state** (the runner's own writes -- never overlap with orchestrator writes):
- `state["iteration_count"] += 1`
- `state["last_tick_at"] = iso8601_now`
- `state["last_verdict"] = verdict`
- Atomic write via tmp+rename (same helper as status.py -- duplicate the small writer rather than import across scripts).
11. **Tick log**: append `TICK pass=<bool> score=<f> iter=<N>` to `.state.log`.
12. Exit 0.
Order of failure-mode Halt writes (all delegated to status.py via `_disable_schedule` best-effort, but the halt itself is a direct `.state.loop` write from the runner):
- Parse failure → `verifier_failed` (R7 above).
The runner **does not** check iterations / budget / drift / task-phase gates itself -- `--check-gate` (R2 step 2) already did. The runner is responsible only for `verifier_failed` (verdict parse) and for `verifier_failed` (score plateau) indirectly via the next tick's `--check-gate`.
### R3 -- `--mode daemon`
- `time.sleep(interval)` loop calling `cmd_tick()`.
- `KeyboardInterrupt` → exit 0 cleanly with a `DAEMON_STOPPED` log entry.
- `--max-iterations N` (optional) caps daemon loop count. 0 / unset = unbounded.
### R4 -- Harness command substitution
`loop.json` `harness.command` is a list of strings. The runner walks each element, replacing `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}` with values from the tick context. Missing tokens stay literal (so configurations can opt out of, say, the `{output}` token by simply not including it).
Default `harness.command` (when `loop.json` doesn't specify one) is `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]`, matching the user's primary harness (D8 -- never inspect model capability).
### R5 -- Context-floor guard (D13)
Before invoking the Implement role, the runner calls `vram_detect.py --loop-mode --json`. If the JSON `loop_mode_eligible == false`, the runner halts the loop with `human_intervention` and writes `HALT human_intervention:context_below_floor`. Existing shell: a "context too small" loop cannot burn tokens through a harness call that would fail anyway.
This guard is implemented in the runner (not in `--check-gate`) because `--check-gate` is per-tick and the available-context value is hardware-state, not loop-state -- we don't want it cached in `.state.loop` between ticks.
### R6 -- Idempotence
- State writes are atomic (tmp+rename).
- The Implement / Verify / Orchestrate invocations do not mutate state; only step 10 writes.
- Verifier parse failure short-circuits before step 10, so a tick that fails to parse its verifier does not increment `iteration_count`. The harness retry on next tick starts from the same `current_task` and `iteration_count`.
- A `KeyboardInterrupt` or `SIGTERM` between steps 5 and 10 leaves `.state.loop` unchanged. The harness subprocess may be left running (the runner does not own process groups in v1).
### R7 -- Tests (`tests/test_loop_runner.py`)
Required by AGENTS.md. All harness calls are stubbed via `monkeypatch.setattr(subprocess, "run", fake_run)`. No live LLM calls in CI.
1. `test_tick_pass` -- fixture loop with a `current_task` in `implement`, mock `--check-gate` returns ok, mock verifier returns `{"pass": true, "score": 0.9}`. Assert `iteration_count == 1`, `last_verdict["pass"] is True`, `.state.log` has `TICK pass=True score=0.9 iter=1`.
2. `test_tick_skip_when_halted` -- pre-halt `.state.loop`, mock `--check-gate` returns not-ok. Assert `iteration_count` unchanged, `.state.log` has `SKIP reason=halted:…`.
3. `test_tick_skip_when_untracked` -- no `.state.loop`. Assert exit 0, `.state.log` has `SKIP untracked`.
4. `test_tick_skip_no_current_task` -- `.state.loop` has `current_task: null`. Assert SKIP `no_current_task`.
5. `test_verifier_parse_failure_halts` -- mock verifier returns garbage. Assert loop halted as `verifier_failed`, `last_verdict` is null, `iteration_count` **unchanged** (R6 idempotence).
6. `test_score_history_capped` -- loop with `score_plateau_window: 3`, run 5 ticks with mock verifier returning scores 0.5, 0.4, 0.4, 0.4, 0.4. Assert `score_history` length is 3 (the last three).
7. `test_json_output_mode` -- `--json` prints a structured tick summary on the last line.
8. `test_daemon_mode_runs_n_iterations` -- `--mode daemon --max-iterations 3` runs `cmd_tick` three times then exits 0.
9. `test_context_floor_refuses` -- mock `vram_detect.py` returns `loop_mode_eligible: false`. Assert loop halted `human_intervention`, harness subprocess never invoked.
10. `test_unknown_mode_rejected` -- `--mode bogus` exits 2.
11. `test_unknown_loop_skip_clean_exit` -- `--loop ghost` exits 0 (R1).
12. `test_harness_command_substitution` -- fixture loop.json with custom `harness.command` containing `{prompt}`, `{cwd}`, `{output}`. Assert stub `subprocess.run` saw the substituted values verbatim.
13. `test_orchestrator_invoked_after_verifier` -- assert subprocess invocations happen in order: gate → implement → verify → orchestrate. Capture argv patterns to confirm.
### R8 -- Out of scope (other tasks)
- Live harness adapter -- provided by user as `harness.command`; no new adapter code.
- `audit` work_source -- task 4 (goal-mode / verifier session) and task 6 (self-improvement template).
- `backlog` work_source -- task 7 (self-improvement loop) and `design/<area>/BACKLOG.md` integration.
- Worktree creation plumbing -- task `add-blast-radius-scheduler`.
- Verifier prompt (`loop-verifier.md`) -- task 6. The runner just reads the filename from `loop.json` and passes it to the harness; it does not parse the prompt itself.
- Orchestrator logic that decides phase transitions -- the orchestrator role does that; the runner only spawns the orchestrator and trusts its `status.py` calls.
## Approach
Single new file `scripts/loop-runner.py`. Stdlib-only (no new pip deps). Reuses small helpers (`_read_state_loop`, `_write_state_loop`, `_loop_dir`, `_read_loop_config`) duplicated inline rather than imported from status.py -- keeps the scripts independent (no risk of one script accidentally being moved/broken affecting the other). The duplicate helpers are <30 lines total.
Tests file `tests/test_loop_runner.py` uses `tmp_path` + a `_stub_subprocess` helper that pattern-matches on argv to return canned outputs.
## Verification
```
python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -v
python3 -m pytest tests/ -q # ensure no regressions
```
+54
View File
@@ -0,0 +1,54 @@
# VERDICT: add-loop-runner
**Status: PASS**
Task delivers `scripts/loop-runner.py` -- the per-tick engine that partners with the brakes layer (task 2). The runner is the only piece that ever invokes the user's harness (subprocess to `loop.json` `harness.command`); it never inspects model capability (D8) and never auto-approves (D4 -- the orchestrator role calls `status.py --approve` itself, the runner only spawns the role).
## Requirement coverage
| Req | Status | Tests |
|-----|--------|-------|
| R1 entrypoint + clean exits | delivered | TestEntrypoint (2) |
| R2 11-step tick flow | delivered | TestTickFlow (5), TestOrchestratorOrdering (1), TestJsonOutput (1) |
| R3 daemon mode | delivered | TestDaemonMode (1) |
| R4 harness command substitution | delivered | TestHarnessSubstitution (1) |
| R5 context-floor guard (D13) | delivered | TestContextFloor (1) |
| R6 idempotence / no state advance on parse failure | delivered | TestVerifierParseFailure (5) |
| R7 tests (18 total) | delivered | per-class rows above |
| R8 out-of-scope items deferred | delivered | (none in code; docs note deferral) |
Tests: 18 new. Full suite: **328 passed** (was 310 + 18 new). No regressions.
## Defense against the five loop deaths -- runtime enforcement
- **drift** -> runner sees not-ok via `--check-gate` and SKIPs (`drift_detected` reason).
- **runaway** -> runner's iteration_count increments only after gate passes; next tick's `--check-gate` halts at `iterations_exhausted`.
- **bad verifier** -> score appended to history; next `--check-gate` halts `verifier_failed` when score plateaus. Parse-failure halts immediately. Idempotent (no state advance).
- **resource burn** -> `--check-gate` halts `budget_exhausted`; runner never invokes the harness before then.
- **undetected halt** -> runner SKIPs on any not-ok gate; tick log records SKIP with reason; `--audit` Cat-6 surfaces the halt across all loops.
## Agnosticism preserved
- **Harness-agnostic**: `harness.command` is a JSON list; any subprocess-capable harness works. Default `opencode run` is only a default; the user can swap it for `claudia run`, `claude --prompt-file`, a custom shell wrapper, or an SSH-remote harness command.
- **OS-agnostic**: `loop-runner.py --mode tick` is pure Python; works on Linux, macOS, Windows. `--mode daemon` is the portable fallback for CI containers without cron/launchd/schtasks.
- **Model-agnostic**: runner never inspects model size/provider. It only checks hardware context (`vram_detect.py --loop-mode --json -- loop_mode_eligible`). The 16k floor (D13) is enforced by the runner, not the gate, because available context is hardware state (per-tick), not loop state (cached).
## Doc impact landed
- `AGENTS.md` "Loop runner" bullet under State Enforcement -- Loops (v1).
- `README.md` loop-runner one-liner.
- `CHANGELOG.md` `[unreleased]` entry for `add-loop-runner`.
No code-doc mismatches.
## Hardening items deferred (tracked in BUG_REPORT + ADVERSARIAL_BUG_REPORT)
1. fcntl lock on `.state.loop` (A2/A7 TOCTOU; same item as `add-status-brakes` A6) -> v1.1.
2. `parse_verdict` score clamp + `pass` string coercion (O6 + A6) -> v1.1.
3. `outputs.retention` in `loop.json` (O5) -> v1.1.
All three are explicit follow-ups; none block this task.
## Resolution
**PASS -- proceed to `complete`.** Task 3 is the runtime half of the loop v1 foundation. With task 2 (brakes) + task 3 (runner) both shipped, the framework can run a single tick end-to-end against any configured harness. Remaining tasks (4 goal-mode, 5 blast-radius-scheduler, 6 templates-onboarding, 7 self-improvement-loop) add work sources, worktree plumbing, usable templates + prompts, and the default-on self-improvement loop. Tasks 8 and 9 are infrastructure cleanup.