Files
Lap Tran f980ccfe27 feat(dashboard): model badges on kanban cards + design doc update
- task.py: Task dataclass gains models: dict[str, str], loaded from
  .state.models in discover_tasks()
- app.py: models dict included in all task API responses
- dashboard.js: model badges rendered between artifacts and subtask
  progress on kanban cards; ROLE_LABELS map for readable tooltips
- styles.css: .task-card-models and .model-badge styles
- design/loops/technical.md: document {model} substitution token,
  per-role model field, and model-divergence brake gate (gate #7)
2026-06-26 13:26:00 -04:00

25 KiB

Loop Engineering — Technical Design

Companion to functional.md. This file is the implementation contract: every line here is what the bootstrap tasks implement. Deviations require a [unreleased] CHANGELOG entry and a design doc update.

1. File Map (what v1 adds)

~/.automaton/
├── scripts/
│   ├── loop-runner.py            # NEW — entry point, --mode tick
│   └── status.py                 # EXTENDED — new flags (see §3)
├── prompts/
│   ├── loop-implement.md         # NEW — minimal implement-role prompt
│   ├── loop-verifier.md          # NEW — minimal verifier-role prompt, emits JSON
│   └── loop-orchestrate.md       # NEW — orchestrator-role prompt, calls status.py
├── templates/loops/
│   ├── ci-triage/                # NEW — example loop template
│   │   └── loop.json
│   └── self-improvement/         # NEW — default-on loop template
│       └── loop.json
├── .automaton/loops/<name>/      # NEW (per-loop state, created by --create-loop)
│   ├── loop.json                 # copied from template
│   ├── .state.loop               # NEW file — loop state (see §2)
│   ├── .state.log                # tick log, append-only
│   ├── automaton-loop-tick.sh    # generated by --install-schedule (self-documenting name)
│   └── worktree/                 # git worktree (unless --no-worktree)
└── tests/
    └── test_loops.py             # NEW — end-to-end coverage

.automaton/loops/ is per-project — under the project's .automaton/, not ~/.automaton/loops/. For framework self-hosting the project is ~/.automaton/ itself, so loops live at ~/.automaton/.automaton/loops/. The exception is the self-improvement loop which is the framework's own — it lives at ~/.automaton/.automaton/loops/self-improvement/ when running on the framework repo.

2. .state.loop Schema

Per-loop state is stored at <loop_dir>/.state.loop as JSON. Single source of truth for loop runtime state. Loops without .state.loop are UNTRACKED — mirror of the v2.0 task .state rule.

{
  "schema_version": 1,
  "name": "self-improvement",
  "status": "running | halted | paused | complete",
  "halt_reason": "iterations_exhausted | budget_exhausted | verifier_failed | drift_detected | human_intervention | null",
  "iteration_count": 0,
  "resumed_count": 0,
  "last_tick_at": "1970-01-01T00:00:00Z",
  "last_verdict": null,
  "score_history": [],
  "current_task": null,
  "worktree_branch": null,
  "worktree_path": null
}
  • iteration_count increments on every tick that the loop actually runs work. A tick that finds no work does not increment (and does not produce a verdict).
  • resumed_count increments when a human issues --approve --loop <name> after a halt. This is distinct so consumers can distinguish "ran out of iterations" from "was resumed."
  • score_history is a capped list (last N entries, where N = score_plateau_window from loop.json). Older entries evicted FIFO.
  • schema_version lets future v1.1+ code migrate without guessing.

3. status.py New Flags

All loop-aware commands route through status.py — no second enforcement surface.

status.py --create-loop <name> --from-template <t>   Create loop dir + loop.json from template
status.py --install-schedule <name> [--interval S]   Generate OS-native unit + automaton-loop-tick.sh
status.py --pause-loop <name>                       Set status=paused, disable schedule
status.py --resume-loop <name>                      Set status=running (after manual pause; DOES NOT clear halt state)
status.py --approve --loop <name>                   Clears halt state, increments resumed_count (D4)
status.py --can-continue <name>                     Check gate: returns OK / HALTED / PAUSED / COMPLETE in JSON
status.py --check-gate <name> [--task <t>]          Pre-tick gate: brakes + scope + budget (see §4)
status.py --can-edit --project <p> [--task <t>] [--file <path>] [--loop <name>] [--loop-worktree]
                                                     Existing semantics preserved; --loop adds worktree scope
status.py --transition <phase> --task <t>           UNCHANGED, but refuses if a halted loop owns the task
status.py --audit                                   UNCHANGED output, + loops section listing statuses
status.py --loop-list                                List loops and their status
status.py --version                                  Print framework version (from config.md)

--approve --loop <name> is the only way to clear a halt. --resume-loop only clears paused (user-initiated), never a halt.

4. Brake Gate Checks (--check-gate)

Called at the top of every tick by loop-runner.py. Returns JSON:

{
  "ok": false,
  "reason": "halted:verifier_failed",
  "halt_reason": "verifier_failed",
  "remaining_iterations": 0,
  "remaining_budget_usd": null,
  "task_phase": "implement",
  "task_in_halt_loop": true,
  "out_of_scope_files": []
}

Gate checks, in order:

  1. Loop status — must be running. Anything else halts the tick immediately.
  2. Iteration count — iteration_count < max_iterations from loop.json.
  3. Budget — if max_budget_usd is set (informational, remote-only), check the harness's reported cost (best-effort: read from a cost.json the harness writes; absence is non-fatal). Below 16k context = hard refuse (D13).
  4. Task phase — if current_task is set, that task's .state must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to human_intervention by some other path), halt as human_intervention.
  5. Worktree drift — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via git diff --name-only main...HEAD restricted to file_scope), halt as drift_detected.
  6. Score plateau — last N entries in score_history are flat or monotonically decreasing (where N = score_plateau_window). Trip → halt as verifier_failed.
  7. Model divergence — in multi-LLM mode (2+ models in models.json), checks that the loop's implement and verify roles use different models. If they share the same model, halt as human_intervention (this prevents same-model verification / rubber-stamping within a loop tick). Single-LLM mode is exempt. Model is resolved from roles[<role>].model if set, otherwise the manifest default.

All halts atomically set status=halted, halt_reason=<reason>, write to .state.log, and call --pause-loop's schedule-disable step (see §6).

5. Verifier Prompt Shape (loop-verifier.md)

Intentionally minimal — a fresh-context LLM should be able to grade a single artifact without loading the framework's 11-file context.

# Loop Verifier — {loop_name}

You are grading one artifact for loop `{loop_name}`.

## Task context (read-only)
{task_brief}           # from .automaton/tasks/<t>/RESEARCH.md or DESIGN.md, capped at 4k tokens

## Artifact under review
{artifact_content}    # the file(s) the Implement: role just wrote

## Last tick's hint
{next_hint}            # capped at 1k tokens, may be empty

## What to check
{acceptance_criteria}  # from loop.json, capped at 2k tokens

## Output (strict JSON, no prose)
{
  "pass": <bool>,
  "score": <0.0-1.0>,
  "reasons": ["..."],
  "next_hint": "..."
}

Score rubric:
- 1.0 = acceptance_criteria fully satisfied, no defects
- 0.7 = functionally complete, minor defects not in criteria
- 0.4 = partial progress, criteria partially addressed
- 0.0 = no useful progress, or artifact is empty/missing

Tier budget (from loop.json): 4k minimum, 16k floor (D13). Below 16k the runner refuses before invoking the verifier — no tiny-context verifier ever runs.

Implement: prompt is similar but emits the artifact to a path, not JSON. Orchestrate: prompt loads the verdict JSON and calls exactly one status.py operation (no edits).

6. Scheduler Unit Generation (--install-schedule)

Platform detection via platform.system():

  • Darwin → ~/Library/LaunchAgents/com.automaton.loop.<name>.plist with StartInterval = interval_seconds. The plist's ProgramArguments calls automaton-loop-tick.sh (generated in the loop dir, see below). --pause-loop renames the plist to .disabled (Launch Agents don't honor a disabled bit portably). --resume-loop renames it back and launchctl loads it.

  • Linux → crontab -l is read, lines for this loop removed, new line added (*/N minutes * * * * <automaton-loop-tick.sh>), crontab - written back. --pause-loop removes the line; --resume-loop re-adds it.

  • Windows → schtasks /create /tn "AutomatonLoop_<name>" /tr "<automaton-loop-tick.bat>" /sc minute /mo <N> /f. --pause-loop calls schtasks /change /tn ... /disable; --resume-loop calls /enable. automaton-loop-tick.sh (generated in the loop dir, chmod +x) is 3 lines. The name is self-documenting: when the scheduler unit (plist ProgramArguments / cron line / schtasks /tr) references the file path, the filename alone conveys "this is automaton's loop-tick entry point" — no need to parse the script to know its role.

#!/usr/bin/env bash
cd "<project_root>"
python3 "<framework>/scripts/loop-runner.py" --mode tick --loop "<name>"

This keeps the OS-specific unit trivial. All logic (gate check, role invocation, verdict parse, brakes) lives in Python.

--daemon opt-in: loop-runner.py --mode daemon --loop <name> runs a time.sleep(interval) loop calling --mode tick per iteration. For CI / shared servers without cron. v1 supports it; default is native.

7. loop-runner.py --mode tick Flow

1. parse --loop <name> → load .state.loop + loop.json
2. status.py --check-gate <name> --json → gate
   if !ok:
       log halt, exit 0 (clean exit; do not crash the scheduler)
3. find_work(work_source):
   audit     → status.py --audit --json, pick highest-severity unresolved
   backlog   → read design/<area>/BACKLOG.md, pick top not-done item
   single    → use current_task from .state.loop
3.5 claim task (cross-loop ownership check — v1.1):
   if candidate != state.current_task:
       status.py --claim-loop-task <name> --task <candidate> --project <p>
       if non-zero exit → SKIP "task_claimed_by_other_loop" (transient; next tick retries)
4. ensure worktree exists (if worktree=true):
   if worktree_path is null or path missing:
       git worktree add .automaton/loops/<name>/worktree -b loop/<name>
       (retry without -b if branch already exists)
       record worktree_path + worktree_branch in .state.loop
   if not a git repo or git unavailable: fall back to project root (WARNING)
5. spawn Implement: session with loop-implement.md,
   cwd = worktree_path (or project root if --no-worktree)
   captures artifact path
6. spawn Verify: session with loop-verifier.md,
   reads artifact, emits verdict JSON
7. parse verdict (strict JSON, accept comments / ```json fences)
   on parse failure → halt as verifier_failed, no retry in v1
   on parse success, coerce defensively (see §6b below):
     * `pass` accepts bool or `"true"`/`"false"` strings (case-insensitive,
       whitespace-stripped). Other strings fall through to `bool(...)`.
     * `score` is clamped to `[0, 1]`. NaN / ±Infinity / non-numeric
       types default to `0.5` (neutral midpoint).
8. append score to score_history (cap = score_plateau_window)
9. spawn Orchestrate: session with loop-orchestrate.md
   inputs: verdict, current_task, current_phase
   executes exactly one status.py call: transition, approve (not auto; orchestrator refuses auto-approve), or escalate to human_intervention
9.5 release on terminal phase (v1.1):
   re-read task .state; if phase is complete or human_intervention:
       state.current_task = None (released for other loops)
10. write verdict to .state.log, increment iteration_count, update last_tick_at
11. if verdict.pass == true:
       transition task to next phase (orchestrator decides which)
       if task reached complete: status=complete in .state.loop

--mode tick is idempotent in the failure case: a crash mid-tick does not advance iteration_count and does not corrupt .state.loop (atomic write via tmp file, same pattern as _write_state).

Lock serialization (v1.1 — add-state-loop-lock)

The runner holds a cross-process file lock (_loop_lock) over the entire tick critical section — from the _gate subprocess call through the step-10 state write. The lock file is <loop_path>/.state.lock (per-loop granularity). POSIX uses fcntl.flock(LOCK_EX); Windows uses msvcrt.locking(LK_LOCK, 1). Blocking acquire, no timeout in v1.1 (operators notice a wedged tick via --loop-list stale last_tick_at).

The same _loop_lock wraps the read-modify-write blocks in cmd_pause_loop, cmd_resume_loop, cmd_approve_loop, and cmd_check_gate (in status.py). --create-loop is intentionally unwrapped — there is no prior state to race against.

Subprocess-deadlock avoidance (D-L6): the runner's _gate call spawns status.py --check-gate as a subprocess. If status.py's cmd_check_gate also acquired _loop_lock, it would deadlock waiting on the parent runner's held flock. To avoid this, the runner passes $AUTOMATON_NO_LOOP_LOCK=1 in that subprocess's env ONLY (scoped to the _gate subprocess; harness subprocesses do NOT inherit it). status.py's _loop_lock checks the env var; if set, it yields without flocking (trusting the caller's outer lock). Standalone CLI users don't set the env var, so --check-gate invoked manually locks normally and serializes against --pause-loop etc.

Same bypass for claim subprocess: --claim-loop-task (step 3.5) is also spawned from inside the runner's _loop_lock. The runner passes $AUTOMATON_NO_LOOP_LOCK=1 in the claim subprocess env for the same reason — cmd_claim_loop_task acquires _loop_lock in status.py, but the runner already holds it. The bypass env is set only for this subprocess; standalone --claim-loop-task invocations (e.g. from --can-continue or future operator tooling) lock normally.

Forbidding re-entry: _loop_lock is NOT re-entrant across processes. Audit callsites to ensure no nested _loop_lock within the same with block. All v1.1 callsites are flat — no nested locks.

Harness contract implication: harnesses invoked via harness.command should NOT call loop-control commands (--pause-loop, --resume-loop, --approve --loop, --check-gate) from inside a tick — that would deadlock waiting on the parent runner's lock. Task commands (--transition, --approve --task, --can-edit, --scope-check, --task, --claim) do NOT touch .state.lock and are safe. Future work: add this to contracts/harness-integration.md.

Outputs retention (v1.1 — add-outputs-retention)

To bound outputs/ directory growth (O5 from add-loop-runner/BUG_REPORT.md), the runner runs GC after step 10 inside _loop_lock, retaining only the last N tick groups (default 20). Configured via loop.json:

"outputs": {
  "retention": 20
}
  • retention = 0 disables GC (unlimited, v1 behavior). Negative coerces to 0 with WARNING.
  • GC iterates outputs/, parses tick{N}- prefix via ^tick(\d+)- regex, computes cutoff = max_seen - retention + 1, deletes files with tick index < cutoff.
  • Non-tick files (no tick{N}- prefix) are preserved.
  • GC failure (permissions, file-not-found mid-iteration) is logged as WARNING and swallowed — never crashes the tick.

8. Harness Invocation

loop-runner.py invokes the user's harness via a single configured command in loop.json:

"harness": {
  "command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"],
  "prompt_var": "{prompt}",
  "cwd_var": "{cwd}",
  "output_var": "{output}"
}

v1.1's default harness.command is opencode run -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the command list (single argv element per token, no shell expansion):

  • {model} -- the model assigned to the role (from roles[<role>].model or manifest default). Passed via extras["model"]. If the role has no model assignment, the token is left unsubstituted.
  • {prompt} -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path.
  • {prompt_content} -- the resolved prompt file's text content as a single argv element. Safe under subprocess.run list mode; no shell quoting needed. Used by the default command since opencode run takes the message as a positional argument and has no --prompt-file flag.
  • {cwd} -- the working directory the harness should run in (the loop's project root or worktree).
  • {output}, {artifact} -- role-specific extras (the implement output path handed to verify).
  • {verdict}, {current_task}, {current_phase}, etc. -- other runtime extras; see _resolve_prompt below.
  • {model} -- the model assigned to the role being invoked (from roles[<role>].model in loop.json, or the manifest default). The runner passes it via the extras["model"] key. If the role has no explicit model, {model} is left as-is (no substitution). This allows per-role model pinning without hardcoding the model name in harness.command.

The default command does NOT hardcode a --model flag; the spawned opencode run inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set harness.command in their loop.json:

"harness": {"command": ["opencode", "run", "--model", "local-mlx/...", "--dir", "{cwd}", "{prompt_content}"]}

The runner core has zero knowledge of which harness is invoked (D8: framework never inspects model/provider). Concrete adapters for non-opencode harnesses (Pi Dev, aider, Cursor, Cline, Copilot) are out of scope for v1 -- users override harness.command to match their harness's CLI shape. Examples:

// Pi Dev
"harness": {"command": ["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]}

// aider
"harness": {"command": ["aider", "--message", "{prompt_content}", "--yes"]}

// Generic shell wrapper (any tool that reads prompt from stdin)
"harness": {"command": ["sh", "-c", "cat {prompt} | my-tool --cwd {cwd}"]}

No new harness adapter is written in v1.1.

Prompt Resolution and Token Substitution

Before building the harness command, the runner calls _resolve_prompt(prompt_ref, extras, loop_path, tick_num, role):

  1. File search: checks <loop_path>/<prompt_ref> first (loop-local override), then ~/.automaton/prompts/<prompt_ref> (framework default). If neither exists, returns the raw prompt_ref string (backward compat -- the harness receives the raw ref).
  2. Content substitution: reads the prompt file and substitutes content-level tokens in the prompt text:
    • {task_brief}, {acceptance_criteria}, {next_hint} -- from the work source and .state.loop
    • {current_task}, {current_phase} -- from .state.loop
    • {verdict} -- the JSON verdict from the verifier (orchestrate role only)
    • {artifact_content} -- reads the file at extras["artifact"] (the implement output path) and substitutes its full text; empty string if the file is missing (verify role only)
  3. Temp file write: writes the substituted content to <loop_path>/outputs/tickN-<role>-prompt.md.
  4. Return: the temp file path replaces {prompt} in the harness command template.

This means the harness receives a fully-resolved prompt file with all context baked in -- no runtime token substitution needed inside the harness. The loop-local override path (<loop_path>/<prompt_ref>) allows per-loop prompt customization without modifying the framework prompts.

9. Self-Improvement Loop Template

templates/loops/self-improvement/loop.json:

{
  "name": "self-improvement",
  "description": "Ticks against status.py --audit on the framework's own repo",
  "work_source": {"kind": "audit", "project": "~/.automaton/"},
  "roles": {
    "implement": {"prompt": "loop-implement.md", "tier": 16000},
    "verify": {"prompt": "loop-verifier.md", "tier": 8000},
    "orchestrate": {"prompt": "loop-orchestrate.md", "tier": 4000}
  },
  "brakes": {
    "max_iterations": 10,
    "max_budget_usd": null,
    "score_plateau_window": 3
  },
  "blast_radius": {"worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]},
  "acceptance_criteria": [
    "Audit findings resolved (no outstanding Cat-1/Cat-2/Cat-4 violations on the resolved task)",
    "All R-numbers from the task SPEC.md implemented",
    "Tests pass with no regressions",
    "Pipeline driven to complete"
  ],
  "harness": {
    "command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"],
    "prompt_var": "{prompt}",
    "cwd_var": "{cwd}",
    "output_var": "{output}"
  },
  "schedule": {"kind": "native", "interval_seconds": 3600}
}

Each role in roles accepts an optional "model" field to pin a specific model for that role (e.g. "implement": {"prompt": "loop-implement.md", "tier": 16000, "model": "model-a"}). When set, the runner passes model=<value> in the harness extras for that role, enabling {model} substitution in harness.command. This is how multi-LLM loops prevent same-model verification — see CONFLICT_MATRIX in status.py.

Installs default-on at install.sh time: status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR" then status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR". Both commands use || true so the framework works even if loop creation fails. update.sh bootstraps the loop idempotently for existing users (checks if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]). Disabling: status.py --pause-loop self-improvement --project ~/.automaton/.

10. Tests (tests/test_loops.py)

Required by AGENTS.md ("tests required for any new Python code"). Covers:

  1. test_loop_tick_pass — fixture task in implement, mock verifier returns PASS → loop ticks, task transitions to code_review, loop status stays running.
  2. test_loop_halt_iterations — set max_iterations: 1, mock verifier returns non-PASS → halts as iterations_exhausted.
  3. test_loop_halt_verifier_failed — three consecutive flat scores → halts as verifier_failed.
  4. test_loop_halt_drift_detected — fixture worktree edited outside file_scope → next --check-gate halts as drift_detected.
  5. test_loop_approve_resume — halted loop, --approve --loop clears halt, resumed_count == 1, next tick allowed.
  6. test_loop_paused_does_not_tick — --pause-loop → --check-gate returns not-ok; --mode tick exits 0 without doing work.
  7. test_loop_untracked_refused — no .state.loop → --check-gate returns error, --mode tick refuses.
  8. test_self_improvement_installs_default_on — fresh install → .automaton/loops/self-improvement/ exists, schedule unit present, status running.
  9. test_loop_no_tiny_context — fixture with tier: 8000 but vram_detect reports ≤16k available → runner refuses before invoking verifier.
  10. test_loop_idempotent_after_crash — mid-tick crash simulated → .state.loop unchanged, next tick proceeds normally.

Test fixtures use tmp_path and stub subprocess.run for the harness invocations — no live LLM calls in CI.

11. Per-Iteration Context Budget (Tier 1 fix)

tech.md requirement: the runner computes a per-tick context budget before invoking any role. Formula:

available_kb = vram_detect.py --json | .available_context_kb
tier_kb      = role's tier from loop.json (capped to available_kb)
floor_kb     = 16000   # D13
if tier_kb < floor_kb: refuse("context window below 16k floor")

Double-headroom bug (vram_detect.py:642+:654): the existing vram_detect.py applies headroom twice. Fix in task fix-context-sizing: remove the inner application, keep only the outer one. The result is what the runner reads.

max(0, ...) clamp + fake 8k/6k defaults (vram_detect.py:651, :698, :699): replace with None + explicit refuse-when-zero in loop mode. Non-loop callers unaffected.

12. Non-Goals (v1) — explicit restatement

  • No parallel mode (D6)
  • No auto-approve (D4)
  • No budget auto-halt on local models (D3 — informational only)
  • No new pip dependencies; stdlib only for new code
  • No network fetches except user-supplied git URL (D11)
  • No framework drafting its own designs (Scope 3 deferred)

13. LOCKED — v1 implementation order (8 bootstrap tasks)

Per the locked plan, the human creates 8 tasks up front. Implementation order respects dependencies:

# Task Depends on
1 fix-context-sizing (none)
2 add-status-brakes 1 (brakes need a working context budget)
3 add-loop-runner 2 (runner calls the brakes)
4 add-goal-mode 3 (/goal extends the runner)
5 add-blast-radius-scheduler 2 + 3
6 add-loop-templates-onboarding 3 (templates reference the runner)
7 add-self-improvement-loop 2, 3, 6 (default-on template + install hook)
8 fix-install-update-flow (parallel, no dep on others)

After task 7 lands → write design/context-sizing/ skeleton + BACKLOG → first loop tick picks up Tier 2 work → handoff.