# Loop Engineering — Technical Design Companion to `functional.md`. This file is the implementation contract: every line here is what the bootstrap tasks implement. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update. ## 1. File Map (what v1 adds) ``` ~/.automaton/ ├── scripts/ │ ├── loop-runner.py # NEW — entry point, --mode tick │ └── status.py # EXTENDED — new flags (see §3) ├── prompts/ │ ├── loop-implement.md # NEW — minimal implement-role prompt │ ├── loop-verifier.md # NEW — minimal verifier-role prompt, emits JSON │ └── loop-orchestrate.md # NEW — orchestrator-role prompt, calls status.py ├── templates/loops/ │ ├── ci-triage/ # NEW — example loop template │ │ └── loop.json │ └── self-improvement/ # NEW — default-on loop template │ └── loop.json ├── .automaton/loops// # NEW (per-loop state, created by --create-loop) │ ├── loop.json # copied from template │ ├── .state.loop # NEW file — loop state (see §2) │ ├── .state.log # tick log, append-only │ ├── automaton-loop-tick.sh # generated by --install-schedule (self-documenting name) │ └── worktree/ # git worktree (unless --no-worktree) └── tests/ └── test_loops.py # NEW — end-to-end coverage ``` `.automaton/loops/` is **per-project** — under the project's `.automaton/`, not `~/.automaton/loops/`. For framework self-hosting the project is `~/.automaton/` itself, so loops live at `~/.automaton/.automaton/loops/`. The exception is the self-improvement loop which is the framework's own — it lives at `~/.automaton/.automaton/loops/self-improvement/` when running on the framework repo. ## 2. `.state.loop` Schema Per-loop state is stored at `/.state.loop` as JSON. Single source of truth for loop runtime state. Loops without `.state.loop` are UNTRACKED — mirror of the v2.0 task `.state` rule. ```json { "schema_version": 1, "name": "self-improvement", "status": "running | halted | paused | complete", "halt_reason": "iterations_exhausted | budget_exhausted | verifier_failed | drift_detected | human_intervention | null", "iteration_count": 0, "resumed_count": 0, "last_tick_at": "1970-01-01T00:00:00Z", "last_verdict": null, "score_history": [], "current_task": null, "worktree_branch": null, "worktree_path": null } ``` - `iteration_count` increments on every tick that the loop **actually runs work**. A tick that finds no work does not increment (and does not produce a verdict). - `resumed_count` increments when a human issues `--approve --loop ` after a halt. This is distinct so consumers can distinguish "ran out of iterations" from "was resumed." - `score_history` is a capped list (last N entries, where N = `score_plateau_window` from `loop.json`). Older entries evicted FIFO. - `schema_version` lets future v1.1+ code migrate without guessing. ## 3. `status.py` New Flags All loop-aware commands route through `status.py` — no second enforcement surface. ``` status.py --create-loop --from-template Create loop dir + loop.json from template status.py --install-schedule [--interval S] Generate OS-native unit + automaton-loop-tick.sh status.py --pause-loop Set status=paused, disable schedule status.py --resume-loop Set status=running (after manual pause; DOES NOT clear halt state) status.py --approve --loop Clears halt state, increments resumed_count (D4) status.py --can-continue Check gate: returns OK / HALTED / PAUSED / COMPLETE in JSON status.py --check-gate [--task ] Pre-tick gate: brakes + scope + budget (see §4) status.py --can-edit --project

[--task ] [--file ] [--loop ] [--loop-worktree] Existing semantics preserved; --loop adds worktree scope status.py --transition --task UNCHANGED, but refuses if a halted loop owns the task status.py --audit UNCHANGED output, + loops section listing statuses status.py --loop-list List loops and their status status.py --version Print framework version (from config.md) ``` `--approve --loop ` is the **only** way to clear a halt. `--resume-loop` only clears `paused` (user-initiated), never a halt. ## 4. Brake Gate Checks (`--check-gate`) Called at the top of every tick by `loop-runner.py`. Returns JSON: ```json { "ok": false, "reason": "halted:verifier_failed", "halt_reason": "verifier_failed", "remaining_iterations": 0, "remaining_budget_usd": null, "task_phase": "implement", "task_in_halt_loop": true, "out_of_scope_files": [] } ``` Gate checks, in order: 1. **Loop status** — must be `running`. Anything else halts the tick immediately. 2. **Iteration count** — `iteration_count < max_iterations` from `loop.json`. 3. **Budget** — if `max_budget_usd` is set (informational, remote-only), check the harness's reported cost (best-effort: read from a `cost.json` the harness writes; absence is non-fatal). Below 16k context = hard refuse (D13). 4. **Task phase** — if `current_task` is set, that task's `.state` must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to `human_intervention` by some other path), halt as `human_intervention`. 5. **Worktree drift** — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via `git diff --name-only main...HEAD` restricted to `file_scope`), halt as `drift_detected`. 6. **Score plateau** — last N entries in `score_history` are flat or monotonically decreasing (where N = `score_plateau_window`). Trip → halt as `verifier_failed`. All halts atomically set `status=halted`, `halt_reason=`, write to `.state.log`, and call `--pause-loop`'s schedule-disable step (see §6). ## 5. Verifier Prompt Shape (`loop-verifier.md`) Intentionally minimal — a fresh-context LLM should be able to grade a single artifact without loading the framework's 11-file context. ``` # Loop Verifier — {loop_name} You are grading one artifact for loop `{loop_name}`. ## Task context (read-only) {task_brief} # from .automaton/tasks//RESEARCH.md or DESIGN.md, capped at 4k tokens ## Artifact under review {artifact_content} # the file(s) the Implement: role just wrote ## Last tick's hint {next_hint} # capped at 1k tokens, may be empty ## What to check {acceptance_criteria} # from loop.json, capped at 2k tokens ## Output (strict JSON, no prose) { "pass": , "score": <0.0-1.0>, "reasons": ["..."], "next_hint": "..." } Score rubric: - 1.0 = acceptance_criteria fully satisfied, no defects - 0.7 = functionally complete, minor defects not in criteria - 0.4 = partial progress, criteria partially addressed - 0.0 = no useful progress, or artifact is empty/missing ``` Tier budget (from `loop.json`): **4k minimum, 16k floor** (D13). Below 16k the runner refuses before invoking the verifier — no tiny-context verifier ever runs. `Implement:` prompt is similar but emits the artifact to a path, not JSON. `Orchestrate:` prompt loads the verdict JSON and calls exactly one `status.py` operation (no edits). ## 6. Scheduler Unit Generation (`--install-schedule`) Platform detection via `platform.system()`: - **Darwin** → `~/Library/LaunchAgents/com.automaton.loop..plist` with `StartInterval` = `interval_seconds`. The plist's `ProgramArguments` calls `automaton-loop-tick.sh` (generated in the loop dir, see below). `--pause-loop` renames the plist to `.disabled` (Launch Agents don't honor a disabled bit portably). `--resume-loop` renames it back and `launchctl load`s it. - **Linux** → `crontab -l` is read, lines for this loop removed, new line added (`*/N minutes * * * * `), `crontab -` written back. `--pause-loop` removes the line; `--resume-loop` re-adds it. - **Windows** → `schtasks /create /tn "AutomatonLoop_" /tr "" /sc minute /mo /f`. `--pause-loop` calls `schtasks /change /tn ... /disable`; `--resume-loop` calls `/enable`. `automaton-loop-tick.sh` (generated in the loop dir, chmod +x) is 3 lines. The name is self-documenting: when the scheduler unit (plist ProgramArguments / cron line / schtasks `/tr`) references the file path, the filename alone conveys "this is automaton's loop-tick entry point" — no need to parse the script to know its role. ```bash #!/usr/bin/env bash cd "" python3 "/scripts/loop-runner.py" --mode tick --loop "" ``` This keeps the OS-specific unit trivial. All logic (gate check, role invocation, verdict parse, brakes) lives in Python. `--daemon` opt-in: `loop-runner.py --mode daemon --loop ` runs a `time.sleep(interval)` loop calling `--mode tick` per iteration. For CI / shared servers without cron. v1 supports it; default is native. ## 7. `loop-runner.py --mode tick` Flow ``` 1. parse --loop → load .state.loop + loop.json 2. status.py --check-gate --json → gate if !ok: log halt, exit 0 (clean exit; do not crash the scheduler) 3. find_work(work_source): audit → status.py --audit --json, pick highest-severity unresolved backlog → read design//BACKLOG.md, pick top not-done item single → use current_task from .state.loop 3.5 claim task (cross-loop ownership check — v1.1): if candidate != state.current_task: status.py --claim-loop-task --task --project

if non-zero exit → SKIP "task_claimed_by_other_loop" (transient; next tick retries) 4. ensure worktree exists (if worktree=true): if worktree_path is null or path missing: git worktree add .automaton/loops//worktree -b loop/ (retry without -b if branch already exists) record worktree_path + worktree_branch in .state.loop if not a git repo or git unavailable: fall back to project root (WARNING) 5. spawn Implement: session with loop-implement.md, cwd = worktree_path (or project root if --no-worktree) captures artifact path 6. spawn Verify: session with loop-verifier.md, reads artifact, emits verdict JSON 7. parse verdict (strict JSON, accept comments / ```json fences) on parse failure → halt as verifier_failed, no retry in v1 on parse success, coerce defensively (see §6b below): * `pass` accepts bool or `"true"`/`"false"` strings (case-insensitive, whitespace-stripped). Other strings fall through to `bool(...)`. * `score` is clamped to `[0, 1]`. NaN / ±Infinity / non-numeric types default to `0.5` (neutral midpoint). 8. append score to score_history (cap = score_plateau_window) 9. spawn Orchestrate: session with loop-orchestrate.md inputs: verdict, current_task, current_phase executes exactly one status.py call: transition, approve (not auto; orchestrator refuses auto-approve), or escalate to human_intervention 9.5 release on terminal phase (v1.1): re-read task .state; if phase is complete or human_intervention: state.current_task = None (released for other loops) 10. write verdict to .state.log, increment iteration_count, update last_tick_at 11. if verdict.pass == true: transition task to next phase (orchestrator decides which) if task reached complete: status=complete in .state.loop ``` `--mode tick` is **idempotent in the failure case**: a crash mid-tick does not advance iteration_count and does not corrupt `.state.loop` (atomic write via tmp file, same pattern as `_write_state`). ### Lock serialization (v1.1 — `add-state-loop-lock`) The runner holds a cross-process file lock (`_loop_lock`) over the entire tick critical section — from the `_gate` subprocess call through the step-10 state write. The lock file is `/.state.lock` (per-loop granularity). POSIX uses `fcntl.flock(LOCK_EX)`; Windows uses `msvcrt.locking(LK_LOCK, 1)`. Blocking acquire, no timeout in v1.1 (operators notice a wedged tick via `--loop-list` stale `last_tick_at`). The same `_loop_lock` wraps the read-modify-write blocks in `cmd_pause_loop`, `cmd_resume_loop`, `cmd_approve_loop`, and `cmd_check_gate` (in `status.py`). `--create-loop` is intentionally unwrapped — there is no prior state to race against. **Subprocess-deadlock avoidance (D-L6)**: the runner's `_gate` call spawns `status.py --check-gate` as a subprocess. If status.py's `cmd_check_gate` also acquired `_loop_lock`, it would deadlock waiting on the parent runner's held flock. To avoid this, the runner passes `$AUTOMATON_NO_LOOP_LOCK=1` in that subprocess's env ONLY (scoped to the `_gate` subprocess; harness subprocesses do NOT inherit it). status.py's `_loop_lock` checks the env var; if set, it yields without flocking (trusting the caller's outer lock). Standalone CLI users don't set the env var, so `--check-gate` invoked manually locks normally and serializes against `--pause-loop` etc. **Same bypass for claim subprocess**: `--claim-loop-task` (step 3.5) is also spawned from inside the runner's `_loop_lock`. The runner passes `$AUTOMATON_NO_LOOP_LOCK=1` in the claim subprocess env for the same reason — `cmd_claim_loop_task` acquires `_loop_lock` in status.py, but the runner already holds it. The bypass env is set only for this subprocess; standalone `--claim-loop-task` invocations (e.g. from `--can-continue` or future operator tooling) lock normally. **Forbidding re-entry**: `_loop_lock` is NOT re-entrant across processes. Audit callsites to ensure no nested `_loop_lock` within the same `with` block. All v1.1 callsites are flat — no nested locks. **Harness contract implication**: harnesses invoked via `harness.command` should NOT call loop-control commands (`--pause-loop`, `--resume-loop`, `--approve --loop`, `--check-gate`) from inside a tick — that would deadlock waiting on the parent runner's lock. Task commands (`--transition`, `--approve --task`, `--can-edit`, `--scope-check`, `--task`, `--claim`) do NOT touch `.state.lock` and are safe. Future work: add this to `contracts/harness-integration.md`. ### Outputs retention (v1.1 — `add-outputs-retention`) To bound `outputs/` directory growth (O5 from `add-loop-runner/BUG_REPORT.md`), the runner runs GC after step 10 inside `_loop_lock`, retaining only the last N tick groups (default 20). Configured via `loop.json`: ```json "outputs": { "retention": 20 } ``` - `retention` = 0 disables GC (unlimited, v1 behavior). Negative coerces to 0 with WARNING. - GC iterates `outputs/`, parses `tick{N}-` prefix via `^tick(\d+)-` regex, computes `cutoff = max_seen - retention + 1`, deletes files with tick index < cutoff. - Non-tick files (no `tick{N}-` prefix) are preserved. - GC failure (permissions, file-not-found mid-iteration) is logged as WARNING and swallowed — never crashes the tick. ## 8. Harness Invocation `loop-runner.py` invokes the user's harness via a single configured command in `loop.json`: ```json "harness": { "command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"], "prompt_var": "{prompt}", "cwd_var": "{cwd}", "output_var": "{output}" } ``` v1.1's default `harness.command` is `opencode run` -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the `command` list (single argv element per token, no shell expansion): - `{prompt}` -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path. - `{prompt_content}` -- the resolved prompt file's text content as a single argv element. Safe under `subprocess.run` list mode; no shell quoting needed. Used by the default command since `opencode run` takes the message as a positional argument and has no `--prompt-file` flag. - `{cwd}` -- the working directory the harness should run in (the loop's project root or worktree). - `{output}`, `{artifact}` -- role-specific extras (the implement output path handed to verify). - `{verdict}`, `{current_task}`, `{current_phase}`, etc. -- other runtime extras; see `_resolve_prompt` below. The default command does NOT hardcode a `--model` flag; the spawned `opencode run` inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set `harness.command` in their `loop.json`: ```json "harness": {"command": ["opencode", "run", "--model", "local-mlx/...", "--dir", "{cwd}", "{prompt_content}"]} ``` The runner core has zero knowledge of which harness is invoked (D8: framework never inspects model/provider). Concrete adapters for non-opencode harnesses (Pi Dev, aider, Cursor, Cline, Copilot) are out of scope for v1 -- users override `harness.command` to match their harness's CLI shape. Examples: ```json // Pi Dev "harness": {"command": ["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]} // aider "harness": {"command": ["aider", "--message", "{prompt_content}", "--yes"]} // Generic shell wrapper (any tool that reads prompt from stdin) "harness": {"command": ["sh", "-c", "cat {prompt} | my-tool --cwd {cwd}"]} ``` No new harness adapter is written in v1.1. ### Prompt Resolution and Token Substitution Before building the harness command, the runner calls `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)`: 1. **File search**: checks `/` first (loop-local override), then `~/.automaton/prompts/` (framework default). If neither exists, returns the raw `prompt_ref` string (backward compat -- the harness receives the raw ref). 2. **Content substitution**: reads the prompt file and substitutes content-level tokens in the prompt text: - `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` -- from the work source and `.state.loop` - `{current_task}`, `{current_phase}` -- from `.state.loop` - `{verdict}` -- the JSON verdict from the verifier (orchestrate role only) - `{artifact_content}` -- reads the file at `extras["artifact"]` (the implement output path) and substitutes its full text; empty string if the file is missing (verify role only) 3. **Temp file write**: writes the substituted content to `/outputs/tickN--prompt.md`. 4. **Return**: the temp file path replaces `{prompt}` in the harness command template. This means the harness receives a fully-resolved prompt file with all context baked in -- no runtime token substitution needed inside the harness. The loop-local override path (`/`) allows per-loop prompt customization without modifying the framework prompts. ## 9. Self-Improvement Loop Template `templates/loops/self-improvement/loop.json`: ```json { "name": "self-improvement", "description": "Ticks against status.py --audit on the framework's own repo", "work_source": {"kind": "audit", "project": "~/.automaton/"}, "roles": { "implement": {"prompt": "loop-implement.md", "tier": 16000}, "verify": {"prompt": "loop-verifier.md", "tier": 8000}, "orchestrate": {"prompt": "loop-orchestrate.md", "tier": 4000} }, "brakes": { "max_iterations": 10, "max_budget_usd": null, "score_plateau_window": 3 }, "blast_radius": {"worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}, "acceptance_criteria": [ "Audit findings resolved (no outstanding Cat-1/Cat-2/Cat-4 violations on the resolved task)", "All R-numbers from the task SPEC.md implemented", "Tests pass with no regressions", "Pipeline driven to complete" ], "harness": { "command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"], "prompt_var": "{prompt}", "cwd_var": "{cwd}", "output_var": "{output}" }, "schedule": {"kind": "native", "interval_seconds": 3600} } ``` Installs default-on at `install.sh` time: `status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` then `status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"`. Both commands use `|| true` so the framework works even if loop creation fails. `update.sh` bootstraps the loop idempotently for existing users (checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`). Disabling: `status.py --pause-loop self-improvement --project ~/.automaton/`. ## 10. Tests (`tests/test_loops.py`) Required by AGENTS.md ("tests required for any new Python code"). Covers: 1. `test_loop_tick_pass` — fixture task in `implement`, mock verifier returns PASS → loop ticks, task transitions to `code_review`, loop status stays `running`. 2. `test_loop_halt_iterations` — set `max_iterations: 1`, mock verifier returns non-PASS → halts as `iterations_exhausted`. 3. `test_loop_halt_verifier_failed` — three consecutive flat scores → halts as `verifier_failed`. 4. `test_loop_halt_drift_detected` — fixture worktree edited outside `file_scope` → next `--check-gate` halts as `drift_detected`. 5. `test_loop_approve_resume` — halted loop, `--approve --loop` clears halt, `resumed_count == 1`, next tick allowed. 6. `test_loop_paused_does_not_tick` — `--pause-loop` → `--check-gate` returns not-ok; `--mode tick` exits 0 without doing work. 7. `test_loop_untracked_refused` — no `.state.loop` → `--check-gate` returns error, `--mode tick` refuses. 8. `test_self_improvement_installs_default_on` — fresh install → `.automaton/loops/self-improvement/` exists, schedule unit present, status `running`. 9. `test_loop_no_tiny_context` — fixture with `tier: 8000` but `vram_detect` reports ≤16k available → runner refuses before invoking verifier. 10. `test_loop_idempotent_after_crash` — mid-tick crash simulated → `.state.loop` unchanged, next tick proceeds normally. Test fixtures use `tmp_path` and stub `subprocess.run` for the harness invocations — no live LLM calls in CI. ## 11. Per-Iteration Context Budget (Tier 1 fix) `tech.md` requirement: the runner computes a per-tick context budget before invoking any role. Formula: ``` available_kb = vram_detect.py --json | .available_context_kb tier_kb = role's tier from loop.json (capped to available_kb) floor_kb = 16000 # D13 if tier_kb < floor_kb: refuse("context window below 16k floor") ``` Double-headroom bug (vram_detect.py:642+:654): the existing `vram_detect.py` applies headroom twice. Fix in task `fix-context-sizing`: remove the inner application, keep only the outer one. The result is what the runner reads. `max(0, ...)` clamp + fake 8k/6k defaults (vram_detect.py:651, :698, :699): replace with `None` + explicit refuse-when-zero in loop mode. Non-loop callers unaffected. ## 12. Non-Goals (v1) — explicit restatement - No parallel mode (D6) - No auto-approve (D4) - No budget auto-halt on local models (D3 — informational only) - No new pip dependencies; stdlib only for new code - No network fetches except user-supplied git URL (D11) - No framework drafting its own designs (Scope 3 deferred) ## 13. LOCKED — v1 implementation order (8 bootstrap tasks) Per the locked plan, the human creates 8 tasks up front. Implementation order respects dependencies: | # | Task | Depends on | |---|---|---| | 1 | `fix-context-sizing` | (none) | | 2 | `add-status-brakes` | 1 (brakes need a working context budget) | | 3 | `add-loop-runner` | 2 (runner calls the brakes) | | 4 | `add-goal-mode` | 3 (`/goal` extends the runner) | | 5 | `add-blast-radius-scheduler` | 2 + 3 | | 6 | `add-loop-templates-onboarding` | 3 (templates reference the runner) | | 7 | `add-self-improvement-loop` | 2, 3, 6 (default-on template + install hook) | | 8 | `fix-install-update-flow` | (parallel, no dep on others) | After task 7 lands → write `design/context-sizing/` skeleton + BACKLOG → first loop tick picks up Tier 2 work → handoff.