- task.py: Task dataclass gains models: dict[str, str], loaded from
.state.models in discover_tasks()
- app.py: models dict included in all task API responses
- dashboard.js: model badges rendered between artifacts and subtask
progress on kanban cards; ROLE_LABELS map for readable tooltips
- styles.css: .task-card-models and .model-badge styles
- design/loops/technical.md: document {model} substitution token,
per-role model field, and model-divergence brake gate (gate #7)
Loop Engineering — Design Index
Status: v1 locked 2026-06-22. Implementation in progress.
What this is
The Automaton framework's loop system: a state-enforcement layer that lets pre-approved work run unattended in capped iterations, with mandatory human approval at every halt. Built on top of the existing status.py phase machine — no second enforcement surface.
Why
Per-session manual driving doesn't scale against the framework's growing backlog (Tier 2 context-sizing cleanup, audit violations, design drift). The framework should be its own first customer: dogfood loops on the framework's own repo.
Documents
functional.md— what v1 does, roles, the five deaths, blast radius, schedules, success criteria. Read this first.technical.md— the implementation contract: file map,.state.loopschema,status.pyflags, gate checks, runner flow, test coverage. Read this if you're implementing.BACKLOG.md— v1.1 and deferred items (Scope 2 design-update loop, Scope 3 self-designing, parallel mode, dashboard panel). The self-improvement loop's work queue.- Sibling design:
../framework/— framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). Has its ownBACKLOG.mdconsumable viawork_source.area = "framework".
v1 scope (locked)
Seven implementation items + one cleanup task:
fix-context-sizing— six Tier 1 fixes invram_detect.py,decompose.md, newloop-verifier.md,config.mdrole sectionadd-status-brakes—.state.loopschema,--can-continue,--check-gate,--version,--create-loop,--install-schedule, transition refusalsadd-loop-runner—scripts/loop-runner.py --mode tickadd-goal-mode— verifier session, graded JSON, score circuit-breakeradd-blast-radius-scheduler—--can-edit --loop-worktree, worktree creation,platform.system()dispatchadd-loop-templates-onboarding—templates/loops/ci-triage/,prompts/loop-{implement,verifier,orchestrate}.md,prompts/onboarding.mdloop branchadd-self-improvement-loop—templates/loops/self-improvement/, default-on at installfix-install-update-flow— user-supplied git URL,.venvcwd bug, Windows venv path, hook copy-vs-symlink, missing--version
The 8 tasks above are the last tasks a human creates by hand. After task 7 lands, the self-improvement loop creates subsequent tasks from BACKLOG.md and --audit output.
Locked decisions (referenced as (Dn) in functional/technical)
| ID | Decision |
|---|---|
| D1 | Native OS scheduler unit + opt-in --daemon fallback |
| D2 | Per-loop git worktree at .automaton/loops/<name>/worktree/; --no-worktree opt-out |
| D3 | max_iterations universal; max_budget_usd optional, informational, remote-only |
| D4 | Human --approve --loop mandatory at halt; no auto-approve in v1 |
| D6 | Parallel mode opt-in --mode parallel; off by default |
| D7 | Graded verifier feedback {pass, score, reasons, next_hint}; circuit-breaker on flat scores |
| D8 | Framework never inspects model capability / size / provider |
| D11 | Install requires user-supplied git URL; refuse with irreversibility warning if absent |
| D12 | Strict session divergence always enforced; model divergence only when Verify: != Implement: |
| D13 | Detection best-effort portable; user override authoritative; 16k floor hard refuse |
| D16 | Tier 1 (six context-sizing fixes) folded into loop v1 |
| D17 | Tier 2 context-sizing cleanup = sibling design design/context-sizing/, driven by first loop workstream |
| D20 | Scopes 1+2 in v1.1; Scope 3 (self-designing) deferred |
| D21 | Self-improvement loop default-on at install (in v1) |
| D22 | Design proposals live at design/<area>/proposals/ as markdown |
| D23 | Comprehension debt tracked via last-read-sha per file; dashboard flags drift |
| D24 | Framework dogfooding is an explicit design principle |
| D25 | Bootstrap: human + AI hand-write runner, brakes, first verifier; loops pick up Tier 2 work |
| D28 | Install fixes part of loop work |
Post-handoff (v1.1+)
After task 7 lands and the first loop tick picks up Tier 2 work from design/context-sizing/BACKLOG.md:
- Scope 2 (design-update loop) → keeps
design/docs in sync with code - Dashboard "Loops" panel + four-deaths audit table
test_design_drift.pydisciplinecontracts/loop-integration.mdlast-read-shacomprehension-debt tracking- The framework creating its own tasks from
BACKLOG.md+--auditoutput
Items deferred indefinitely: parallel mode as default (D6 stays opt-in), Scope 3 (framework drafting its own designs), auto-approve (D4).