Companion to `functional.md`. This file is the implementation contract: every line here is what the bootstrap tasks implement. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update.
## 1. File Map (what v1 adds)
```
~/.automaton/
├── scripts/
│ ├── loop-runner.py # NEW — entry point, --mode tick
│ └── status.py # EXTENDED — new flags (see §3)
├── prompts/
│ ├── loop-implement.md # NEW — minimal implement-role prompt
`.automaton/loops/` is **per-project** — under the project's `.automaton/`, not `~/.automaton/loops/`. For framework self-hosting the project is `~/.automaton/` itself, so loops live at `~/.automaton/.automaton/loops/`. The exception is the self-improvement loop which is the framework's own — it lives at `~/.automaton/.automaton/loops/self-improvement/` when running on the framework repo.
## 2. `.state.loop` Schema
Per-loop state is stored at `<loop_dir>/.state.loop` as JSON. Single source of truth for loop runtime state. Loops without `.state.loop` are UNTRACKED — mirror of the v2.0 task `.state` rule.
-`iteration_count` increments on every tick that the loop **actually runs work**. A tick that finds no work does not increment (and does not produce a verdict).
-`resumed_count` increments when a human issues `--approve --loop <name>` after a halt. This is distinct so consumers can distinguish "ran out of iterations" from "was resumed."
-`score_history` is a capped list (last N entries, where N = `score_plateau_window` from `loop.json`). Older entries evicted FIFO.
-`schema_version` lets future v1.1+ code migrate without guessing.
## 3. `status.py` New Flags
All loop-aware commands route through `status.py` — no second enforcement surface.
```
status.py --create-loop <name> --from-template <t> Create loop dir + loop.json from template
status.py --install-schedule <name> [--interval S] Generate OS-native unit + automaton-loop-tick.sh
status.py --pause-loop <name> Set status=paused, disable schedule
status.py --resume-loop <name> Set status=running (after manual pause; DOES NOT clear halt state)
status.py --approve --loop <name> Clears halt state, increments resumed_count (D4)
status.py --can-continue <name> Check gate: returns OK / HALTED / PAUSED / COMPLETE in JSON
status.py --version Print framework version (from config.md)
```
`--approve --loop <name>` is the **only** way to clear a halt. `--resume-loop` only clears `paused` (user-initiated), never a halt.
## 4. Brake Gate Checks (`--check-gate`)
Called at the top of every tick by `loop-runner.py`. Returns JSON:
```json
{
"ok":false,
"reason":"halted:verifier_failed",
"halt_reason":"verifier_failed",
"remaining_iterations":0,
"remaining_budget_usd":null,
"task_phase":"implement",
"task_in_halt_loop":true,
"out_of_scope_files":[]
}
```
Gate checks, in order:
1.**Loop status** — must be `running`. Anything else halts the tick immediately.
2.**Iteration count** — `iteration_count < max_iterations` from `loop.json`.
3.**Budget** — if `max_budget_usd` is set (informational, remote-only), check the harness's reported cost (best-effort: read from a `cost.json` the harness writes; absence is non-fatal). Below 16k context = hard refuse (D13).
4.**Task phase** — if `current_task` is set, that task's `.state` must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to `human_intervention` by some other path), halt as `human_intervention`.
5.**Worktree drift** — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via `git diff --name-only main...HEAD` restricted to `file_scope`), halt as `drift_detected`.
6.**Score plateau** — last N entries in `score_history` are flat or monotonically decreasing (where N = `score_plateau_window`). Trip → halt as `verifier_failed`.
7.**Model divergence** — in multi-LLM mode (2+ models in `models.json`), checks that the loop's implement and verify roles use different models. If they share the same model, halt as `human_intervention` (this prevents same-model verification / rubber-stamping within a loop tick). Single-LLM mode is exempt. Model is resolved from `roles[<role>].model` if set, otherwise the manifest default.
- 0.0 = no useful progress, or artifact is empty/missing
```
Tier budget (from `loop.json`): **4k minimum, 16k floor** (D13). Below 16k the runner refuses before invoking the verifier — no tiny-context verifier ever runs.
`Implement:` prompt is similar but emits the artifact to a path, not JSON. `Orchestrate:` prompt loads the verdict JSON and calls exactly one `status.py` operation (no edits).
## 6. Scheduler Unit Generation (`--install-schedule`)
Platform detection via `platform.system()`:
- **Darwin** → `~/Library/LaunchAgents/com.automaton.loop.<name>.plist` with `StartInterval` = `interval_seconds`. The plist's `ProgramArguments` calls `automaton-loop-tick.sh` (generated in the loop dir, see below). `--pause-loop` renames the plist to `.disabled` (Launch Agents don't honor a disabled bit portably). `--resume-loop` renames it back and `launchctl load`s it.
- **Linux** → `crontab -l` is read, lines for this loop removed, new line added (`*/N minutes * * * * <automaton-loop-tick.sh>`), `crontab -` written back. `--pause-loop` removes the line; `--resume-loop` re-adds it.
`automaton-loop-tick.sh` (generated in the loop dir, chmod +x) is 3 lines. The name is self-documenting: when the scheduler unit (plist ProgramArguments / cron line / schtasks `/tr`) references the file path, the filename alone conveys "this is automaton's loop-tick entry point" — no need to parse the script to know its role.
This keeps the OS-specific unit trivial. All logic (gate check, role invocation, verdict parse, brakes) lives in Python.
`--daemon` opt-in: `loop-runner.py --mode daemon --loop <name>` runs a `time.sleep(interval)` loop calling `--mode tick` per iteration. For CI / shared servers without cron. v1 supports it; default is native.
on parse failure → halt as verifier_failed, no retry in v1
on parse success, coerce defensively (see §6b below):
* `pass` accepts bool or `"true"`/`"false"` strings (case-insensitive,
whitespace-stripped). Other strings fall through to `bool(...)`.
* `score` is clamped to `[0, 1]`. NaN / ±Infinity / non-numeric
types default to `0.5` (neutral midpoint).
8. append score to score_history (cap = score_plateau_window)
9. spawn Orchestrate: session with loop-orchestrate.md
inputs: verdict, current_task, current_phase
executes exactly one status.py call: transition, approve (not auto; orchestrator refuses auto-approve), or escalate to human_intervention
9.5 release on terminal phase (v1.1):
re-read task .state; if phase is complete or human_intervention:
state.current_task = None (released for other loops)
10. write verdict to .state.log, increment iteration_count, update last_tick_at
11. if verdict.pass == true:
transition task to next phase (orchestrator decides which)
if task reached complete: status=complete in .state.loop
```
`--mode tick` is **idempotent in the failure case**: a crash mid-tick does not advance iteration_count and does not corrupt `.state.loop` (atomic write via tmp file, same pattern as `_write_state`).
The runner holds a cross-process file lock (`_loop_lock`) over the entire tick critical section — from the `_gate` subprocess call through the step-10 state write. The lock file is `<loop_path>/.state.lock` (per-loop granularity). POSIX uses `fcntl.flock(LOCK_EX)`; Windows uses `msvcrt.locking(LK_LOCK, 1)`. Blocking acquire, no timeout in v1.1 (operators notice a wedged tick via `--loop-list` stale `last_tick_at`).
The same `_loop_lock` wraps the read-modify-write blocks in `cmd_pause_loop`, `cmd_resume_loop`, `cmd_approve_loop`, and `cmd_check_gate` (in `status.py`). `--create-loop` is intentionally unwrapped — there is no prior state to race against.
**Subprocess-deadlock avoidance (D-L6)**: the runner's `_gate` call spawns `status.py --check-gate` as a subprocess. If status.py's `cmd_check_gate` also acquired `_loop_lock`, it would deadlock waiting on the parent runner's held flock. To avoid this, the runner passes `$AUTOMATON_NO_LOOP_LOCK=1` in that subprocess's env ONLY (scoped to the `_gate` subprocess; harness subprocesses do NOT inherit it). status.py's `_loop_lock` checks the env var; if set, it yields without flocking (trusting the caller's outer lock). Standalone CLI users don't set the env var, so `--check-gate` invoked manually locks normally and serializes against `--pause-loop` etc.
**Same bypass for claim subprocess**: `--claim-loop-task` (step 3.5) is also spawned from inside the runner's `_loop_lock`. The runner passes `$AUTOMATON_NO_LOOP_LOCK=1` in the claim subprocess env for the same reason — `cmd_claim_loop_task` acquires `_loop_lock` in status.py, but the runner already holds it. The bypass env is set only for this subprocess; standalone `--claim-loop-task` invocations (e.g. from `--can-continue` or future operator tooling) lock normally.
**Forbidding re-entry**: `_loop_lock` is NOT re-entrant across processes. Audit callsites to ensure no nested `_loop_lock` within the same `with` block. All v1.1 callsites are flat — no nested locks.
**Harness contract implication**: harnesses invoked via `harness.command` should NOT call loop-control commands (`--pause-loop`, `--resume-loop`, `--approve --loop`, `--check-gate`) from inside a tick — that would deadlock waiting on the parent runner's lock. Task commands (`--transition`, `--approve --task`, `--can-edit`, `--scope-check`, `--task`, `--claim`) do NOT touch `.state.lock` and are safe. Future work: add this to `contracts/harness-integration.md`.
To bound `outputs/` directory growth (O5 from `add-loop-runner/BUG_REPORT.md`), the runner runs GC after step 10 inside `_loop_lock`, retaining only the last N tick groups (default 20). Configured via `loop.json`:
```json
"outputs":{
"retention":20
}
```
-`retention` = 0 disables GC (unlimited, v1 behavior). Negative coerces to 0 with WARNING.
- GC iterates `outputs/`, parses `tick{N}-` prefix via `^tick(\d+)-` regex, computes `cutoff = max_seen - retention + 1`, deletes files with tick index < cutoff.
- Non-tick files (no `tick{N}-` prefix) are preserved.
- GC failure (permissions, file-not-found mid-iteration) is logged as WARNING and swallowed — never crashes the tick.
## 8. Harness Invocation
`loop-runner.py` invokes the user's harness via a single configured command in `loop.json`:
v1.1's default `harness.command` is `opencode run` -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the `command` list (single argv element per token, no shell expansion):
-`{model}` -- the model assigned to the role (from `roles[<role>].model` or manifest default). Passed via `extras["model"]`. If the role has no model assignment, the token is left unsubstituted.
-`{prompt}` -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path.
-`{prompt_content}` -- the resolved prompt file's text content as a single argv element. Safe under `subprocess.run` list mode; no shell quoting needed. Used by the default command since `opencode run` takes the message as a positional argument and has no `--prompt-file` flag.
-`{cwd}` -- the working directory the harness should run in (the loop's project root or worktree).
-`{output}`, `{artifact}` -- role-specific extras (the implement output path handed to verify).
-`{verdict}`, `{current_task}`, `{current_phase}`, etc. -- other runtime extras; see `_resolve_prompt` below.
-`{model}` -- the model assigned to the role being invoked (from `roles[<role>].model` in `loop.json`, or the manifest default). The runner passes it via the `extras["model"]` key. If the role has no explicit model, `{model}` is left as-is (no substitution). This allows per-role model pinning without hardcoding the model name in `harness.command`.
The default command does NOT hardcode a `--model` flag; the spawned `opencode run` inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set `harness.command` in their `loop.json`:
The runner core has zero knowledge of which harness is invoked (D8: framework never inspects model/provider). Concrete adapters for non-opencode harnesses (Pi Dev, aider, Cursor, Cline, Copilot) are out of scope for v1 -- users override `harness.command` to match their harness's CLI shape. Examples:
Before building the harness command, the runner calls `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)`:
1.**File search**: checks `<loop_path>/<prompt_ref>` first (loop-local override), then `~/.automaton/prompts/<prompt_ref>` (framework default). If neither exists, returns the raw `prompt_ref` string (backward compat -- the harness receives the raw ref).
2.**Content substitution**: reads the prompt file and substitutes content-level tokens in the prompt text:
-`{task_brief}`, `{acceptance_criteria}`, `{next_hint}` -- from the work source and `.state.loop`
-`{current_task}`, `{current_phase}` -- from `.state.loop`
-`{verdict}` -- the JSON verdict from the verifier (orchestrate role only)
-`{artifact_content}` -- reads the file at `extras["artifact"]` (the implement output path) and substitutes its full text; empty string if the file is missing (verify role only)
3.**Temp file write**: writes the substituted content to `<loop_path>/outputs/tickN-<role>-prompt.md`.
4.**Return**: the temp file path replaces `{prompt}` in the harness command template.
This means the harness receives a fully-resolved prompt file with all context baked in -- no runtime token substitution needed inside the harness. The loop-local override path (`<loop_path>/<prompt_ref>`) allows per-loop prompt customization without modifying the framework prompts.
## 9. Self-Improvement Loop Template
`templates/loops/self-improvement/loop.json`:
```json
{
"name":"self-improvement",
"description":"Ticks against status.py --audit on the framework's own repo",
Each role in `roles` accepts an optional `"model"` field to pin a specific model for that role (e.g. `"implement": {"prompt": "loop-implement.md", "tier": 16000, "model": "model-a"}`). When set, the runner passes `model=<value>` in the harness extras for that role, enabling `{model}` substitution in `harness.command`. This is how multi-LLM loops prevent same-model verification — see `CONFLICT_MATRIX` in `status.py`.
Installs default-on at `install.sh` time: `status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` then `status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"`. Both commands use `|| true` so the framework works even if loop creation fails. `update.sh` bootstraps the loop idempotently for existing users (checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`). Disabling: `status.py --pause-loop self-improvement --project ~/.automaton/`.
## 10. Tests (`tests/test_loops.py`)
Required by AGENTS.md ("tests required for any new Python code"). Covers:
1.`test_loop_tick_pass` — fixture task in `implement`, mock verifier returns PASS → loop ticks, task transitions to `code_review`, loop status stays `running`.
2.`test_loop_halt_iterations` — set `max_iterations: 1`, mock verifier returns non-PASS → halts as `iterations_exhausted`.
3.`test_loop_halt_verifier_failed` — three consecutive flat scores → halts as `verifier_failed`.
4.`test_loop_halt_drift_detected` — fixture worktree edited outside `file_scope` → next `--check-gate` halts as `drift_detected`.
tier_kb = role's tier from loop.json (capped to available_kb)
floor_kb = 16000 # D13
if tier_kb < floor_kb: refuse("context window below 16k floor")
```
Double-headroom bug (vram_detect.py:642+:654): the existing `vram_detect.py` applies headroom twice. Fix in task `fix-context-sizing`: remove the inner application, keep only the outer one. The result is what the runner reads.