Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
CI / build (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,168 @@
|
||||
# Loop Engineering — Functional Design
|
||||
|
||||
Status: v1 (locked 2026-06-22). Supersedes any prior informal loop discussions.
|
||||
|
||||
Audience: framework maintainers (currently: the human and one AI assistant). After handoff the self-improvement loop is also an audience — designs must be legible to a fresh-context LLM verifier.
|
||||
|
||||
## 1. Problem
|
||||
|
||||
Automaton currently runs **one phase per session**: a human issues one command, the agent completes one phase, the session ends. The next phase needs a new session and a new manual command. This produces correct, disciplined work but cannot run unattended.
|
||||
|
||||
A growing backlog (Tier 2 context-sizing cleanup, audit violations, design-doc drift) makes per-session manual driving unsustainable. The framework should be its own first customer: dogfood the loop system on the framework's own repo.
|
||||
|
||||
## 2. Goals
|
||||
|
||||
v1 — **the framework takes over implementation work that already has an approved design**:
|
||||
|
||||
1. Run a single task through every phase unattended, capped by `max_iterations` and brake gates.
|
||||
2. Halt predictably on the five deaths (see §6) and require human `--approve` to resume.
|
||||
3. Stay model-agnostic: never inspect model capability, provider, or size.
|
||||
4. Stay harness-agnostic: enforce brakes inside `status.py`, not in the harness.
|
||||
5. Stay OS-portable: detect platform via `platform.system()`, generate native scheduler units.
|
||||
6. Ship a self-improvement loop template, *default-on at install*, that ticks against `~/.automaton/tasks/` using `status.py --audit` as its trigger source. This is the seed of self-management.
|
||||
|
||||
## 3. Non-Goals (v1)
|
||||
|
||||
- The framework **drafting its own designs**. v1 loops run against pre-authored designs in `design/<area>/`. Self-designing loops are Scope 3, deferred indefinitely.
|
||||
- Parallel multi-task execution (`--mode parallel`). Off by default, never required for v1.
|
||||
- A daemon runtime. v1 uses an OS-native scheduler as the tick source; `--daemon` is opt-in.
|
||||
- A new agent harness. The loop runner is a thin Python driver that shells out to the existing harness the user already uses.
|
||||
- Network-fetched dependencies. New code is Python stdlib only. No new pip installs.
|
||||
|
||||
## 4. Loop Definition
|
||||
|
||||
A **loop** is a configured instance of the generic loop runner, bound to:
|
||||
|
||||
- a **work source** (where it finds the next unit of work)
|
||||
- a **verifier** (how it grades the result)
|
||||
- a **role configuration** (which LLM session plays each role)
|
||||
- a **blast radius** (where its edits are allowed to land)
|
||||
- a **schedule** (when its ticks fire)
|
||||
|
||||
A **tick** is one execution of the loop: find work → produce artifact → verify → decide. A tick is *not* a phase transition. Ticks compose on top of the existing phase machine: a tick may cause one or more `--transition` calls, or may cause zero.
|
||||
|
||||
## 5. Roles
|
||||
|
||||
Every loop defines three session roles. Roles are *sessions*, not models — a single model can fill multiple roles. A role is identified by the prompt file fed into it and the harness session that runs it.
|
||||
|
||||
| Role | Prompt | Responsibility |
|
||||
|---|---|---|
|
||||
| `Implement:` | `prompts/loop-implement.md` | Makes the artifact for this tick. Bound to the task's current phase. |
|
||||
| `Verify:` | `prompts/loop-verifier.md` | Grades the artifact and emits the JSON verdict that feeds the next decision. |
|
||||
| `Orchestrate:` | `prompts/loop-orchestrate.md` | Decides whether to continue the tick, advance the phase, halt, or escalate. Reads the verifier's verdict and applies brakes. |
|
||||
|
||||
**Conflict-of-interest rule:** `Verify:` and `Implement:` must never be the same *session*. If a user's harness cannot run two sessions, they fall back to session-only divergence — still safe, still functional. The framework never inspects whether the two sessions use the same *model* — that decision belongs to the user, not to us (D12).
|
||||
|
||||
**Tiered context budgets:** Each role has a context budget computed from `vram_detect.py`, capped to its tier. See `technical.md` §4 for the exact tiers.
|
||||
|
||||
## 6. The Five Deaths (halt conditions requiring human `--approve`)
|
||||
|
||||
A loop halts — and refuses its next tick — when any of the following fires. Resuming requires `status.py --approve --loop <name>` from a human. No auto-approve path (D4).
|
||||
|
||||
1. **`iterations_exhausted`** — tick count reached `max_iterations` without a terminating verdict.
|
||||
2. **`budget_exhausted`** — wall-clock or token budget (informational, remote-only) hit its cap.
|
||||
3. **`verifier_failed`** — verifier returned `pass: false` and the score circuit-breaker tripped (flat score across N consecutive ticks, D7).
|
||||
4. **`drift_detected`** — the artifact written this tick is out of scope of the loop's blast radius, or the task's `.state` is no longer one of the loop's expected phases. (The loop did something it wasn't allowed to do.)
|
||||
5. **`human_intervention`** — the task itself transitioned to `human_intervention` via the normal state machine (e.g. referee verdict was BLOCKED).
|
||||
|
||||
When a loop is halted, its schedule unit self-disables until `--approve` clears it. Implementation detail (how the scheduler disables itself) is in `technical.md` §6.
|
||||
|
||||
## 7. Blast Radius
|
||||
|
||||
A loop runs in a **per-loop git worktree** at `.automaton/loops/<name>/worktree/`, branched from the project's main branch on first tick. All edits from the loop land in the worktree; merging back to main is a human action.
|
||||
|
||||
`--no-worktree` is an opt-out for users who want loops to edit the primary checkout directly (e.g. in CI environments with no git state). v1 documents this flag; the default is worktree-on.
|
||||
|
||||
`status.py --can-edit` is extended: edits are allowed only if the path is under the loop's worktree (or the loop has `--no-worktree`). Pre-existing `--can-edit` semantics for harness integration are preserved.
|
||||
|
||||
## 8. Schedules
|
||||
|
||||
Tick sources, in priority order:
|
||||
|
||||
1. **Native OS unit** (default). `status.py --install-schedule <name>` emits:
|
||||
- macOS: a `~/Library/LaunchAgents/com.automaton.loop.<name>.plist`
|
||||
- Linux: crontab line via `crontab -l | ... | crontab -`
|
||||
- Windows: a `schtasks /create` invocation
|
||||
2. **`--daemon`** (opt-in). Runs `loop-runner.py --mode tick` on a Python `sleep` loop. Useful in CI containers and on shared servers without cron access.
|
||||
3. **Manual**: `python loop-runner.py --mode tick --loop <name>`. Always available.
|
||||
|
||||
A loop's `--install-schedule` produces a small platform dispatcher script at `.automaton/loops/<name>/automaton-loop-tick.sh` that re-enters `loop-runner.py --mode tick`. The OS unit calls `automaton-loop-tick.sh`. This keeps the native unit trivial and inspects-free; all logic lives in Python.
|
||||
|
||||
## 9. Loop Configuration File
|
||||
|
||||
Each loop is configured at `.automaton/loops/<name>/loop.json`. Exact schema is in `technical.md` §2. Key fields:
|
||||
|
||||
- `name`, `description`
|
||||
- `work_source`: `{kind: "audit" | "backlog" | "single", project: optional, area: optional}`
|
||||
- `acceptance_criteria`: optional string OR list of strings — fed to the verifier prompt as `{acceptance_criteria}` (capped at 2k tokens); the loop's "goal"
|
||||
- `roles`: `{implement: {...}, verify: {...}, orchestrate: {...}}` — role → prompt + tier
|
||||
- `brakes`: `{max_iterations: N, max_budget_usd: optional, score_plateau_window: N}`
|
||||
- `blast_radius`: `{worktree: bool, file_scope: [paths], base_branch: str (optional, default "main")}`
|
||||
- `outputs`: `{retention: N}` — keep last N tick groups in `outputs/` (default 20, 0 = unlimited, v1.1 `add-outputs-retention`)
|
||||
- `schedule`: `{kind: "native" | "daemon" | "manual", interval_seconds: N}`
|
||||
|
||||
## 10. Verifier Contract
|
||||
|
||||
`Verify:` returns a JSON object of **exactly** this shape (graded verdict, P6):
|
||||
|
||||
```json
|
||||
{
|
||||
"pass": false,
|
||||
"score": 0.62,
|
||||
"reasons": ["..."],
|
||||
"next_hint": "..."
|
||||
}
|
||||
```
|
||||
|
||||
- `pass` (bool) definitive. The runner also accepts the strings `"true"`/`"false"` (case-insensitive, whitespace-stripped) for defensive compatibility; other non-bool types defer to `bool(...)` (v1.1 `harden-parse-verdict`).
|
||||
- `score` (0.0–1.0) is the signal for the circuit-breaker: if N consecutive ticks have a flat or monotonically-decreasing score, halt as `verifier_failed`. The runner clamps any out-of-range numeric score to `[0, 1]`; NaN, ±Infinity, and non-numeric values default to `0.5` (v1.1 `harden-parse-verdict`).
|
||||
- `reasons` is human-legible, written to the loop's tick log.
|
||||
- `next_hint` is fed back into the next tick's `Implement:` session as the only persistent context. Capped at 1k tokens by the runner; longer hints are truncated.
|
||||
|
||||
The verifier prompt is intentionally minimal (see `technical.md` §5) — a fresh-context LLM verifier should not need to load the framework's 11-file context to grade a single artifact.
|
||||
|
||||
## 11. Self-Improvement Loop (Scope 1, default-on)
|
||||
|
||||
A pre-shipped loop template at `templates/loops/self-improvement/` ticks against `status.py --audit` output for the framework's *own* repo (`~/.automaton/`). Its `Implement:` role picks the highest-severity audit violation, drafts a fix in a worktree, transitions the affected task (or creates a new task via `status.py --create-task` when the violation is a new issue), runs its phases, and hands off to a human reviewer.
|
||||
|
||||
At install time, `install.sh` calls `status.py --create-loop self-improvement --from-template self-improvement` and `--install-schedule self-improvement --interval 3600`. Users can disable with `status.py --pause-loop self-improvement` (added in v1).
|
||||
|
||||
This is the literal seed of self-management: a loop that runs on the framework's own audit output. It is *not* drafting designs — it runs against documented audit cleanup work only.
|
||||
|
||||
## 12. Cross-Loop Task Claim (v1.1)
|
||||
|
||||
Prevents two loops from racing to claim the same task (closes `add-status-brakes/ADVERSARIAL_BUG_REPORT.md` A2). The runner calls `status.py --claim-loop-task <loop> --task <taskname>` inside `_loop_lock` after `_find_work` picks a candidate (step 3.5 in tick flow). The claim command scans all running/paused loops' `.state.loop.current_task` and refuses if another loop already owns the target task. Self-ownership is idempotent. The claim subprocess uses the same `$AUTOMATON_NO_LOOP_LOCK` env bypass as `--check-gate` to avoid deadlock on the parent's held flock.
|
||||
|
||||
**Release**: after the orchestrate subprocess (step 9.5), the runner re-reads the task's `.state`. If the phase is `complete` or `human_intervention`, `current_task` is cleared to `None` — the task is released for other loops. Non-terminal phases leave the claim intact.
|
||||
|
||||
**Design decisions**:
|
||||
- Claim is a `status.py` subprocess, not an in-process helper (single authority for loop state).
|
||||
- Cross-loop scan is advisory (no cross-loop lock). Race window is one tick; self-healing on next tick.
|
||||
- `paused` and `halted` loops retain their claim (resumed loop resumes without re-claiming).
|
||||
- Release is the runner's responsibility, not the orchestrator's. Runner re-reads task state after orchestrate; terminal → clear.
|
||||
|
||||
## 13. v1.1 (deferred from this scope)
|
||||
|
||||
- A **design-update loop** template (Scope 2) that keeps `design/<area>/*.md` in sync with code
|
||||
- `test_design_drift.py` discipline + dashboard "Loops" panel
|
||||
- `contracts/loop-integration.md`
|
||||
- `last-read-sha` comprehension-debt tracking
|
||||
- Tier 2 context-sizing cleanup (the first real work-stream the loops pick up)
|
||||
|
||||
These are listed in `design/loops/BACKLOG.md` so the self-improvement loop has a work queue from day one, but they are not implemented by v1.
|
||||
|
||||
## 14. Harness Compatibility
|
||||
|
||||
v1 loops work with any harness that can (a) run a session against a given prompt and (b) write the resulting artifact to a path. The `Orchestrate:` role is the only role that ever calls `status.py`; this means `status.py` integration is bounded to one session per tick. Concrete harness adapters are out of scope for v1 — loops are driven by `loop-runner.py` invoking the user's existing harness as a subprocess.
|
||||
|
||||
## 15. Success Criteria for v1
|
||||
|
||||
1. A single capped loop runs unattended for ≤ N ticks without human involvement and halts cleanly on one of the five deaths.
|
||||
2. A human `--approve` resumes a halted loop, and the resume increments `__resumed_count` not `__iteration_count`.
|
||||
3. End-to-end test `tests/test_loops.py` passes, covering: one terminating ticket (PASS), one `verifier_failed` halt, one `iterations_exhausted` halt, one `drift_detected` halt, one `--approve` resume.
|
||||
4. The self-improvement loop installs default-on and ticks once against a seeded audit violation in a temp fixture, producing a worktree + a transition not an edit to main.
|
||||
5. `pytest tests/ -v` is still green for the framework's pre-existing test suite (no regression).
|
||||
|
||||
## 16. Locked Decision Index
|
||||
|
||||
All decisions referenced by `(Dn)` above are recorded in v1 scope conversation (this session). They are non-negotiable for v1 implementation. Changes require a design doc update and a new `[unreleased]` changelog entry.
|
||||
Reference in New Issue
Block a user