- task.py: Task dataclass gains models: dict[str, str], loaded from
.state.models in discover_tasks()
- app.py: models dict included in all task API responses
- dashboard.js: model badges rendered between artifacts and subtask
progress on kanban cards; ROLE_LABELS map for readable tooltips
- styles.css: .task-card-models and .model-badge styles
- design/loops/technical.md: document {model} substitution token,
per-role model field, and model-divergence brake gate (gate #7)
25 KiB
Loop Engineering — Technical Design
Companion to functional.md. This file is the implementation contract: every line here is what the bootstrap tasks implement. Deviations require a [unreleased] CHANGELOG entry and a design doc update.
1. File Map (what v1 adds)
~/.automaton/
├── scripts/
│ ├── loop-runner.py # NEW — entry point, --mode tick
│ └── status.py # EXTENDED — new flags (see §3)
├── prompts/
│ ├── loop-implement.md # NEW — minimal implement-role prompt
│ ├── loop-verifier.md # NEW — minimal verifier-role prompt, emits JSON
│ └── loop-orchestrate.md # NEW — orchestrator-role prompt, calls status.py
├── templates/loops/
│ ├── ci-triage/ # NEW — example loop template
│ │ └── loop.json
│ └── self-improvement/ # NEW — default-on loop template
│ └── loop.json
├── .automaton/loops/<name>/ # NEW (per-loop state, created by --create-loop)
│ ├── loop.json # copied from template
│ ├── .state.loop # NEW file — loop state (see §2)
│ ├── .state.log # tick log, append-only
│ ├── automaton-loop-tick.sh # generated by --install-schedule (self-documenting name)
│ └── worktree/ # git worktree (unless --no-worktree)
└── tests/
└── test_loops.py # NEW — end-to-end coverage
.automaton/loops/ is per-project — under the project's .automaton/, not ~/.automaton/loops/. For framework self-hosting the project is ~/.automaton/ itself, so loops live at ~/.automaton/.automaton/loops/. The exception is the self-improvement loop which is the framework's own — it lives at ~/.automaton/.automaton/loops/self-improvement/ when running on the framework repo.
2. .state.loop Schema
Per-loop state is stored at <loop_dir>/.state.loop as JSON. Single source of truth for loop runtime state. Loops without .state.loop are UNTRACKED — mirror of the v2.0 task .state rule.
{
"schema_version": 1,
"name": "self-improvement",
"status": "running | halted | paused | complete",
"halt_reason": "iterations_exhausted | budget_exhausted | verifier_failed | drift_detected | human_intervention | null",
"iteration_count": 0,
"resumed_count": 0,
"last_tick_at": "1970-01-01T00:00:00Z",
"last_verdict": null,
"score_history": [],
"current_task": null,
"worktree_branch": null,
"worktree_path": null
}
iteration_countincrements on every tick that the loop actually runs work. A tick that finds no work does not increment (and does not produce a verdict).resumed_countincrements when a human issues--approve --loop <name>after a halt. This is distinct so consumers can distinguish "ran out of iterations" from "was resumed."score_historyis a capped list (last N entries, where N =score_plateau_windowfromloop.json). Older entries evicted FIFO.schema_versionlets future v1.1+ code migrate without guessing.
3. status.py New Flags
All loop-aware commands route through status.py — no second enforcement surface.
status.py --create-loop <name> --from-template <t> Create loop dir + loop.json from template
status.py --install-schedule <name> [--interval S] Generate OS-native unit + automaton-loop-tick.sh
status.py --pause-loop <name> Set status=paused, disable schedule
status.py --resume-loop <name> Set status=running (after manual pause; DOES NOT clear halt state)
status.py --approve --loop <name> Clears halt state, increments resumed_count (D4)
status.py --can-continue <name> Check gate: returns OK / HALTED / PAUSED / COMPLETE in JSON
status.py --check-gate <name> [--task <t>] Pre-tick gate: brakes + scope + budget (see §4)
status.py --can-edit --project <p> [--task <t>] [--file <path>] [--loop <name>] [--loop-worktree]
Existing semantics preserved; --loop adds worktree scope
status.py --transition <phase> --task <t> UNCHANGED, but refuses if a halted loop owns the task
status.py --audit UNCHANGED output, + loops section listing statuses
status.py --loop-list List loops and their status
status.py --version Print framework version (from config.md)
--approve --loop <name> is the only way to clear a halt. --resume-loop only clears paused (user-initiated), never a halt.
4. Brake Gate Checks (--check-gate)
Called at the top of every tick by loop-runner.py. Returns JSON:
{
"ok": false,
"reason": "halted:verifier_failed",
"halt_reason": "verifier_failed",
"remaining_iterations": 0,
"remaining_budget_usd": null,
"task_phase": "implement",
"task_in_halt_loop": true,
"out_of_scope_files": []
}
Gate checks, in order:
- Loop status — must be
running. Anything else halts the tick immediately. - Iteration count —
iteration_count < max_iterationsfromloop.json. - Budget — if
max_budget_usdis set (informational, remote-only), check the harness's reported cost (best-effort: read from acost.jsonthe harness writes; absence is non-fatal). Below 16k context = hard refuse (D13). - Task phase — if
current_taskis set, that task's.statemust still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. tohuman_interventionby some other path), halt ashuman_intervention. - Worktree drift — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via
git diff --name-only main...HEADrestricted tofile_scope), halt asdrift_detected. - Score plateau — last N entries in
score_historyare flat or monotonically decreasing (where N =score_plateau_window). Trip → halt asverifier_failed. - Model divergence — in multi-LLM mode (2+ models in
models.json), checks that the loop's implement and verify roles use different models. If they share the same model, halt ashuman_intervention(this prevents same-model verification / rubber-stamping within a loop tick). Single-LLM mode is exempt. Model is resolved fromroles[<role>].modelif set, otherwise the manifest default.
All halts atomically set status=halted, halt_reason=<reason>, write to .state.log, and call --pause-loop's schedule-disable step (see §6).
5. Verifier Prompt Shape (loop-verifier.md)
Intentionally minimal — a fresh-context LLM should be able to grade a single artifact without loading the framework's 11-file context.
# Loop Verifier — {loop_name}
You are grading one artifact for loop `{loop_name}`.
## Task context (read-only)
{task_brief} # from .automaton/tasks/<t>/RESEARCH.md or DESIGN.md, capped at 4k tokens
## Artifact under review
{artifact_content} # the file(s) the Implement: role just wrote
## Last tick's hint
{next_hint} # capped at 1k tokens, may be empty
## What to check
{acceptance_criteria} # from loop.json, capped at 2k tokens
## Output (strict JSON, no prose)
{
"pass": <bool>,
"score": <0.0-1.0>,
"reasons": ["..."],
"next_hint": "..."
}
Score rubric:
- 1.0 = acceptance_criteria fully satisfied, no defects
- 0.7 = functionally complete, minor defects not in criteria
- 0.4 = partial progress, criteria partially addressed
- 0.0 = no useful progress, or artifact is empty/missing
Tier budget (from loop.json): 4k minimum, 16k floor (D13). Below 16k the runner refuses before invoking the verifier — no tiny-context verifier ever runs.
Implement: prompt is similar but emits the artifact to a path, not JSON. Orchestrate: prompt loads the verdict JSON and calls exactly one status.py operation (no edits).
6. Scheduler Unit Generation (--install-schedule)
Platform detection via platform.system():
-
Darwin →
~/Library/LaunchAgents/com.automaton.loop.<name>.plistwithStartInterval=interval_seconds. The plist'sProgramArgumentscallsautomaton-loop-tick.sh(generated in the loop dir, see below).--pause-looprenames the plist to.disabled(Launch Agents don't honor a disabled bit portably).--resume-looprenames it back andlaunchctl loads it. -
Linux →
crontab -lis read, lines for this loop removed, new line added (*/N minutes * * * * <automaton-loop-tick.sh>),crontab -written back.--pause-loopremoves the line;--resume-loopre-adds it. -
Windows →
schtasks /create /tn "AutomatonLoop_<name>" /tr "<automaton-loop-tick.bat>" /sc minute /mo <N> /f.--pause-loopcallsschtasks /change /tn ... /disable;--resume-loopcalls/enable.automaton-loop-tick.sh(generated in the loop dir, chmod +x) is 3 lines. The name is self-documenting: when the scheduler unit (plist ProgramArguments / cron line / schtasks/tr) references the file path, the filename alone conveys "this is automaton's loop-tick entry point" — no need to parse the script to know its role.
#!/usr/bin/env bash
cd "<project_root>"
python3 "<framework>/scripts/loop-runner.py" --mode tick --loop "<name>"
This keeps the OS-specific unit trivial. All logic (gate check, role invocation, verdict parse, brakes) lives in Python.
--daemon opt-in: loop-runner.py --mode daemon --loop <name> runs a time.sleep(interval) loop calling --mode tick per iteration. For CI / shared servers without cron. v1 supports it; default is native.
7. loop-runner.py --mode tick Flow
1. parse --loop <name> → load .state.loop + loop.json
2. status.py --check-gate <name> --json → gate
if !ok:
log halt, exit 0 (clean exit; do not crash the scheduler)
3. find_work(work_source):
audit → status.py --audit --json, pick highest-severity unresolved
backlog → read design/<area>/BACKLOG.md, pick top not-done item
single → use current_task from .state.loop
3.5 claim task (cross-loop ownership check — v1.1):
if candidate != state.current_task:
status.py --claim-loop-task <name> --task <candidate> --project <p>
if non-zero exit → SKIP "task_claimed_by_other_loop" (transient; next tick retries)
4. ensure worktree exists (if worktree=true):
if worktree_path is null or path missing:
git worktree add .automaton/loops/<name>/worktree -b loop/<name>
(retry without -b if branch already exists)
record worktree_path + worktree_branch in .state.loop
if not a git repo or git unavailable: fall back to project root (WARNING)
5. spawn Implement: session with loop-implement.md,
cwd = worktree_path (or project root if --no-worktree)
captures artifact path
6. spawn Verify: session with loop-verifier.md,
reads artifact, emits verdict JSON
7. parse verdict (strict JSON, accept comments / ```json fences)
on parse failure → halt as verifier_failed, no retry in v1
on parse success, coerce defensively (see §6b below):
* `pass` accepts bool or `"true"`/`"false"` strings (case-insensitive,
whitespace-stripped). Other strings fall through to `bool(...)`.
* `score` is clamped to `[0, 1]`. NaN / ±Infinity / non-numeric
types default to `0.5` (neutral midpoint).
8. append score to score_history (cap = score_plateau_window)
9. spawn Orchestrate: session with loop-orchestrate.md
inputs: verdict, current_task, current_phase
executes exactly one status.py call: transition, approve (not auto; orchestrator refuses auto-approve), or escalate to human_intervention
9.5 release on terminal phase (v1.1):
re-read task .state; if phase is complete or human_intervention:
state.current_task = None (released for other loops)
10. write verdict to .state.log, increment iteration_count, update last_tick_at
11. if verdict.pass == true:
transition task to next phase (orchestrator decides which)
if task reached complete: status=complete in .state.loop
--mode tick is idempotent in the failure case: a crash mid-tick does not advance iteration_count and does not corrupt .state.loop (atomic write via tmp file, same pattern as _write_state).
Lock serialization (v1.1 — add-state-loop-lock)
The runner holds a cross-process file lock (_loop_lock) over the entire tick critical section — from the _gate subprocess call through the step-10 state write. The lock file is <loop_path>/.state.lock (per-loop granularity). POSIX uses fcntl.flock(LOCK_EX); Windows uses msvcrt.locking(LK_LOCK, 1). Blocking acquire, no timeout in v1.1 (operators notice a wedged tick via --loop-list stale last_tick_at).
The same _loop_lock wraps the read-modify-write blocks in cmd_pause_loop, cmd_resume_loop, cmd_approve_loop, and cmd_check_gate (in status.py). --create-loop is intentionally unwrapped — there is no prior state to race against.
Subprocess-deadlock avoidance (D-L6): the runner's _gate call spawns status.py --check-gate as a subprocess. If status.py's cmd_check_gate also acquired _loop_lock, it would deadlock waiting on the parent runner's held flock. To avoid this, the runner passes $AUTOMATON_NO_LOOP_LOCK=1 in that subprocess's env ONLY (scoped to the _gate subprocess; harness subprocesses do NOT inherit it). status.py's _loop_lock checks the env var; if set, it yields without flocking (trusting the caller's outer lock). Standalone CLI users don't set the env var, so --check-gate invoked manually locks normally and serializes against --pause-loop etc.
Same bypass for claim subprocess: --claim-loop-task (step 3.5) is also spawned from inside the runner's _loop_lock. The runner passes $AUTOMATON_NO_LOOP_LOCK=1 in the claim subprocess env for the same reason — cmd_claim_loop_task acquires _loop_lock in status.py, but the runner already holds it. The bypass env is set only for this subprocess; standalone --claim-loop-task invocations (e.g. from --can-continue or future operator tooling) lock normally.
Forbidding re-entry: _loop_lock is NOT re-entrant across processes. Audit callsites to ensure no nested _loop_lock within the same with block. All v1.1 callsites are flat — no nested locks.
Harness contract implication: harnesses invoked via harness.command should NOT call loop-control commands (--pause-loop, --resume-loop, --approve --loop, --check-gate) from inside a tick — that would deadlock waiting on the parent runner's lock. Task commands (--transition, --approve --task, --can-edit, --scope-check, --task, --claim) do NOT touch .state.lock and are safe. Future work: add this to contracts/harness-integration.md.
Outputs retention (v1.1 — add-outputs-retention)
To bound outputs/ directory growth (O5 from add-loop-runner/BUG_REPORT.md), the runner runs GC after step 10 inside _loop_lock, retaining only the last N tick groups (default 20). Configured via loop.json:
"outputs": {
"retention": 20
}
retention= 0 disables GC (unlimited, v1 behavior). Negative coerces to 0 with WARNING.- GC iterates
outputs/, parsestick{N}-prefix via^tick(\d+)-regex, computescutoff = max_seen - retention + 1, deletes files with tick index < cutoff. - Non-tick files (no
tick{N}-prefix) are preserved. - GC failure (permissions, file-not-found mid-iteration) is logged as WARNING and swallowed — never crashes the tick.
8. Harness Invocation
loop-runner.py invokes the user's harness via a single configured command in loop.json:
"harness": {
"command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"],
"prompt_var": "{prompt}",
"cwd_var": "{cwd}",
"output_var": "{output}"
}
v1.1's default harness.command is opencode run -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the command list (single argv element per token, no shell expansion):
{model}-- the model assigned to the role (fromroles[<role>].modelor manifest default). Passed viaextras["model"]. If the role has no model assignment, the token is left unsubstituted.{prompt}-- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path.{prompt_content}-- the resolved prompt file's text content as a single argv element. Safe undersubprocess.runlist mode; no shell quoting needed. Used by the default command sinceopencode runtakes the message as a positional argument and has no--prompt-fileflag.{cwd}-- the working directory the harness should run in (the loop's project root or worktree).{output},{artifact}-- role-specific extras (the implement output path handed to verify).{verdict},{current_task},{current_phase}, etc. -- other runtime extras; see_resolve_promptbelow.{model}-- the model assigned to the role being invoked (fromroles[<role>].modelinloop.json, or the manifest default). The runner passes it via theextras["model"]key. If the role has no explicit model,{model}is left as-is (no substitution). This allows per-role model pinning without hardcoding the model name inharness.command.
The default command does NOT hardcode a --model flag; the spawned opencode run inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set harness.command in their loop.json:
"harness": {"command": ["opencode", "run", "--model", "local-mlx/...", "--dir", "{cwd}", "{prompt_content}"]}
The runner core has zero knowledge of which harness is invoked (D8: framework never inspects model/provider). Concrete adapters for non-opencode harnesses (Pi Dev, aider, Cursor, Cline, Copilot) are out of scope for v1 -- users override harness.command to match their harness's CLI shape. Examples:
// Pi Dev
"harness": {"command": ["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]}
// aider
"harness": {"command": ["aider", "--message", "{prompt_content}", "--yes"]}
// Generic shell wrapper (any tool that reads prompt from stdin)
"harness": {"command": ["sh", "-c", "cat {prompt} | my-tool --cwd {cwd}"]}
No new harness adapter is written in v1.1.
Prompt Resolution and Token Substitution
Before building the harness command, the runner calls _resolve_prompt(prompt_ref, extras, loop_path, tick_num, role):
- File search: checks
<loop_path>/<prompt_ref>first (loop-local override), then~/.automaton/prompts/<prompt_ref>(framework default). If neither exists, returns the rawprompt_refstring (backward compat -- the harness receives the raw ref). - Content substitution: reads the prompt file and substitutes content-level tokens in the prompt text:
{task_brief},{acceptance_criteria},{next_hint}-- from the work source and.state.loop{current_task},{current_phase}-- from.state.loop{verdict}-- the JSON verdict from the verifier (orchestrate role only){artifact_content}-- reads the file atextras["artifact"](the implement output path) and substitutes its full text; empty string if the file is missing (verify role only)
- Temp file write: writes the substituted content to
<loop_path>/outputs/tickN-<role>-prompt.md. - Return: the temp file path replaces
{prompt}in the harness command template.
This means the harness receives a fully-resolved prompt file with all context baked in -- no runtime token substitution needed inside the harness. The loop-local override path (<loop_path>/<prompt_ref>) allows per-loop prompt customization without modifying the framework prompts.
9. Self-Improvement Loop Template
templates/loops/self-improvement/loop.json:
{
"name": "self-improvement",
"description": "Ticks against status.py --audit on the framework's own repo",
"work_source": {"kind": "audit", "project": "~/.automaton/"},
"roles": {
"implement": {"prompt": "loop-implement.md", "tier": 16000},
"verify": {"prompt": "loop-verifier.md", "tier": 8000},
"orchestrate": {"prompt": "loop-orchestrate.md", "tier": 4000}
},
"brakes": {
"max_iterations": 10,
"max_budget_usd": null,
"score_plateau_window": 3
},
"blast_radius": {"worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]},
"acceptance_criteria": [
"Audit findings resolved (no outstanding Cat-1/Cat-2/Cat-4 violations on the resolved task)",
"All R-numbers from the task SPEC.md implemented",
"Tests pass with no regressions",
"Pipeline driven to complete"
],
"harness": {
"command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"],
"prompt_var": "{prompt}",
"cwd_var": "{cwd}",
"output_var": "{output}"
},
"schedule": {"kind": "native", "interval_seconds": 3600}
}
Each role in roles accepts an optional "model" field to pin a specific model for that role (e.g. "implement": {"prompt": "loop-implement.md", "tier": 16000, "model": "model-a"}). When set, the runner passes model=<value> in the harness extras for that role, enabling {model} substitution in harness.command. This is how multi-LLM loops prevent same-model verification — see CONFLICT_MATRIX in status.py.
Installs default-on at install.sh time: status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR" then status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR". Both commands use || true so the framework works even if loop creation fails. update.sh bootstraps the loop idempotently for existing users (checks if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]). Disabling: status.py --pause-loop self-improvement --project ~/.automaton/.
10. Tests (tests/test_loops.py)
Required by AGENTS.md ("tests required for any new Python code"). Covers:
test_loop_tick_pass— fixture task inimplement, mock verifier returns PASS → loop ticks, task transitions tocode_review, loop status staysrunning.test_loop_halt_iterations— setmax_iterations: 1, mock verifier returns non-PASS → halts asiterations_exhausted.test_loop_halt_verifier_failed— three consecutive flat scores → halts asverifier_failed.test_loop_halt_drift_detected— fixture worktree edited outsidefile_scope→ next--check-gatehalts asdrift_detected.test_loop_approve_resume— halted loop,--approve --loopclears halt,resumed_count == 1, next tick allowed.test_loop_paused_does_not_tick—--pause-loop→--check-gatereturns not-ok;--mode tickexits 0 without doing work.test_loop_untracked_refused— no.state.loop→--check-gatereturns error,--mode tickrefuses.test_self_improvement_installs_default_on— fresh install →.automaton/loops/self-improvement/exists, schedule unit present, statusrunning.test_loop_no_tiny_context— fixture withtier: 8000butvram_detectreports ≤16k available → runner refuses before invoking verifier.test_loop_idempotent_after_crash— mid-tick crash simulated →.state.loopunchanged, next tick proceeds normally.
Test fixtures use tmp_path and stub subprocess.run for the harness invocations — no live LLM calls in CI.
11. Per-Iteration Context Budget (Tier 1 fix)
tech.md requirement: the runner computes a per-tick context budget before invoking any role. Formula:
available_kb = vram_detect.py --json | .available_context_kb
tier_kb = role's tier from loop.json (capped to available_kb)
floor_kb = 16000 # D13
if tier_kb < floor_kb: refuse("context window below 16k floor")
Double-headroom bug (vram_detect.py:642+:654): the existing vram_detect.py applies headroom twice. Fix in task fix-context-sizing: remove the inner application, keep only the outer one. The result is what the runner reads.
max(0, ...) clamp + fake 8k/6k defaults (vram_detect.py:651, :698, :699): replace with None + explicit refuse-when-zero in loop mode. Non-loop callers unaffected.
12. Non-Goals (v1) — explicit restatement
- No parallel mode (D6)
- No auto-approve (D4)
- No budget auto-halt on local models (D3 — informational only)
- No new pip dependencies; stdlib only for new code
- No network fetches except user-supplied git URL (D11)
- No framework drafting its own designs (Scope 3 deferred)
13. LOCKED — v1 implementation order (8 bootstrap tasks)
Per the locked plan, the human creates 8 tasks up front. Implementation order respects dependencies:
| # | Task | Depends on |
|---|---|---|
| 1 | fix-context-sizing |
(none) |
| 2 | add-status-brakes |
1 (brakes need a working context budget) |
| 3 | add-loop-runner |
2 (runner calls the brakes) |
| 4 | add-goal-mode |
3 (/goal extends the runner) |
| 5 | add-blast-radius-scheduler |
2 + 3 |
| 6 | add-loop-templates-onboarding |
3 (templates reference the runner) |
| 7 | add-self-improvement-loop |
2, 3, 6 (default-on template + install hook) |
| 8 | fix-install-update-flow |
(parallel, no dep on others) |
After task 7 lands → write design/context-sizing/ skeleton + BACKLOG → first loop tick picks up Tier 2 work → handoff.