Files
automaton/design/loops/README.md
T
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00

70 lines
4.7 KiB
Markdown

# Loop Engineering — Design Index
Status: **v1 locked 2026-06-22**. Implementation in progress.
## What this is
The Automaton framework's loop system: a state-enforcement layer that lets pre-approved work run unattended in capped iterations, with mandatory human approval at every halt. Built on top of the existing `status.py` phase machine — no second enforcement surface.
## Why
Per-session manual driving doesn't scale against the framework's growing backlog (Tier 2 context-sizing cleanup, audit violations, design drift). The framework should be its own first customer: dogfood loops on the framework's own repo.
## Documents
- [`functional.md`](functional.md) — what v1 does, roles, the five deaths, blast radius, schedules, success criteria. **Read this first.**
- [`technical.md`](technical.md) — the implementation contract: file map, `.state.loop` schema, `status.py` flags, gate checks, runner flow, test coverage. **Read this if you're implementing.**
- [`BACKLOG.md`](BACKLOG.md) — v1.1 and deferred items (Scope 2 design-update loop, Scope 3 self-designing, parallel mode, dashboard panel). The self-improvement loop's work queue.
- **Sibling design**: [`../framework/`](../framework/) — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). Has its own `BACKLOG.md` consumable via `work_source.area = "framework"`.
## v1 scope (locked)
Seven implementation items + one cleanup task:
1. `fix-context-sizing` — six Tier 1 fixes in `vram_detect.py`, `decompose.md`, new `loop-verifier.md`, `config.md` role section
2. `add-status-brakes` — `.state.loop` schema, `--can-continue`, `--check-gate`, `--version`, `--create-loop`, `--install-schedule`, transition refusals
3. `add-loop-runner` — `scripts/loop-runner.py --mode tick`
4. `add-goal-mode` — verifier session, graded JSON, score circuit-breaker
5. `add-blast-radius-scheduler` — `--can-edit --loop-worktree`, worktree creation, `platform.system()` dispatch
6. `add-loop-templates-onboarding` — `templates/loops/ci-triage/`, `prompts/loop-{implement,verifier,orchestrate}.md`, `prompts/onboarding.md` loop branch
7. `add-self-improvement-loop` — `templates/loops/self-improvement/`, default-on at install
8. `fix-install-update-flow` — user-supplied git URL, `.venv` cwd bug, Windows venv path, hook copy-vs-symlink, missing `--version`
The 8 tasks above are **the last tasks a human creates by hand**. After task 7 lands, the self-improvement loop creates subsequent tasks from `BACKLOG.md` and `--audit` output.
## Locked decisions (referenced as `(Dn)` in functional/technical)
| ID | Decision |
|---|---|
| D1 | Native OS scheduler unit + opt-in `--daemon` fallback |
| D2 | Per-loop git worktree at `.automaton/loops/<name>/worktree/`; `--no-worktree` opt-out |
| D3 | `max_iterations` universal; `max_budget_usd` optional, informational, remote-only |
| D4 | Human `--approve --loop` mandatory at halt; no auto-approve in v1 |
| D6 | Parallel mode opt-in `--mode parallel`; off by default |
| D7 | Graded verifier feedback `{pass, score, reasons, next_hint}`; circuit-breaker on flat scores |
| D8 | Framework never inspects model capability / size / provider |
| D11 | Install requires user-supplied git URL; refuse with irreversibility warning if absent |
| D12 | Strict session divergence always enforced; model divergence only when `Verify:` != `Implement:` |
| D13 | Detection best-effort portable; user override authoritative; 16k floor hard refuse |
| D16 | Tier 1 (six context-sizing fixes) folded into loop v1 |
| D17 | Tier 2 context-sizing cleanup = sibling design `design/context-sizing/`, driven by first loop workstream |
| D20 | Scopes 1+2 in v1.1; Scope 3 (self-designing) deferred |
| D21 | Self-improvement loop default-on at install (in v1) |
| D22 | Design proposals live at `design/<area>/proposals/` as markdown |
| D23 | Comprehension debt tracked via `last-read-sha` per file; dashboard flags drift |
| D24 | Framework dogfooding is an explicit design principle |
| D25 | Bootstrap: human + AI hand-write runner, brakes, first verifier; loops pick up Tier 2 work |
| D28 | Install fixes part of loop work |
## Post-handoff (v1.1+)
After task 7 lands and the first loop tick picks up Tier 2 work from `design/context-sizing/BACKLOG.md`:
- Scope 2 (design-update loop) → keeps `design/` docs in sync with code
- Dashboard "Loops" panel + four-deaths audit table
- `test_design_drift.py` discipline
- `contracts/loop-integration.md`
- `last-read-sha` comprehension-debt tracking
- The framework creating its own tasks from `BACKLOG.md` + `--audit` output
Items deferred indefinitely: parallel mode as default (D6 stays opt-in), Scope 3 (framework drafting its own designs), auto-approve (D4).