Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
CI / build (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,397 @@
|
||||
# Loop Engineering — Technical Design
|
||||
|
||||
Companion to `functional.md`. This file is the implementation contract: every line here is what the bootstrap tasks implement. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update.
|
||||
|
||||
## 1. File Map (what v1 adds)
|
||||
|
||||
```
|
||||
~/.automaton/
|
||||
├── scripts/
|
||||
│ ├── loop-runner.py # NEW — entry point, --mode tick
|
||||
│ └── status.py # EXTENDED — new flags (see §3)
|
||||
├── prompts/
|
||||
│ ├── loop-implement.md # NEW — minimal implement-role prompt
|
||||
│ ├── loop-verifier.md # NEW — minimal verifier-role prompt, emits JSON
|
||||
│ └── loop-orchestrate.md # NEW — orchestrator-role prompt, calls status.py
|
||||
├── templates/loops/
|
||||
│ ├── ci-triage/ # NEW — example loop template
|
||||
│ │ └── loop.json
|
||||
│ └── self-improvement/ # NEW — default-on loop template
|
||||
│ └── loop.json
|
||||
├── .automaton/loops/<name>/ # NEW (per-loop state, created by --create-loop)
|
||||
│ ├── loop.json # copied from template
|
||||
│ ├── .state.loop # NEW file — loop state (see §2)
|
||||
│ ├── .state.log # tick log, append-only
|
||||
│ ├── automaton-loop-tick.sh # generated by --install-schedule (self-documenting name)
|
||||
│ └── worktree/ # git worktree (unless --no-worktree)
|
||||
└── tests/
|
||||
└── test_loops.py # NEW — end-to-end coverage
|
||||
```
|
||||
|
||||
`.automaton/loops/` is **per-project** — under the project's `.automaton/`, not `~/.automaton/loops/`. For framework self-hosting the project is `~/.automaton/` itself, so loops live at `~/.automaton/.automaton/loops/`. The exception is the self-improvement loop which is the framework's own — it lives at `~/.automaton/.automaton/loops/self-improvement/` when running on the framework repo.
|
||||
|
||||
## 2. `.state.loop` Schema
|
||||
|
||||
Per-loop state is stored at `<loop_dir>/.state.loop` as JSON. Single source of truth for loop runtime state. Loops without `.state.loop` are UNTRACKED — mirror of the v2.0 task `.state` rule.
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 1,
|
||||
"name": "self-improvement",
|
||||
"status": "running | halted | paused | complete",
|
||||
"halt_reason": "iterations_exhausted | budget_exhausted | verifier_failed | drift_detected | human_intervention | null",
|
||||
"iteration_count": 0,
|
||||
"resumed_count": 0,
|
||||
"last_tick_at": "1970-01-01T00:00:00Z",
|
||||
"last_verdict": null,
|
||||
"score_history": [],
|
||||
"current_task": null,
|
||||
"worktree_branch": null,
|
||||
"worktree_path": null
|
||||
}
|
||||
```
|
||||
|
||||
- `iteration_count` increments on every tick that the loop **actually runs work**. A tick that finds no work does not increment (and does not produce a verdict).
|
||||
- `resumed_count` increments when a human issues `--approve --loop <name>` after a halt. This is distinct so consumers can distinguish "ran out of iterations" from "was resumed."
|
||||
- `score_history` is a capped list (last N entries, where N = `score_plateau_window` from `loop.json`). Older entries evicted FIFO.
|
||||
- `schema_version` lets future v1.1+ code migrate without guessing.
|
||||
|
||||
## 3. `status.py` New Flags
|
||||
|
||||
All loop-aware commands route through `status.py` — no second enforcement surface.
|
||||
|
||||
```
|
||||
status.py --create-loop <name> --from-template <t> Create loop dir + loop.json from template
|
||||
status.py --install-schedule <name> [--interval S] Generate OS-native unit + automaton-loop-tick.sh
|
||||
status.py --pause-loop <name> Set status=paused, disable schedule
|
||||
status.py --resume-loop <name> Set status=running (after manual pause; DOES NOT clear halt state)
|
||||
status.py --approve --loop <name> Clears halt state, increments resumed_count (D4)
|
||||
status.py --can-continue <name> Check gate: returns OK / HALTED / PAUSED / COMPLETE in JSON
|
||||
status.py --check-gate <name> [--task <t>] Pre-tick gate: brakes + scope + budget (see §4)
|
||||
status.py --can-edit --project <p> [--task <t>] [--file <path>] [--loop <name>] [--loop-worktree]
|
||||
Existing semantics preserved; --loop adds worktree scope
|
||||
status.py --transition <phase> --task <t> UNCHANGED, but refuses if a halted loop owns the task
|
||||
status.py --audit UNCHANGED output, + loops section listing statuses
|
||||
status.py --loop-list List loops and their status
|
||||
status.py --version Print framework version (from config.md)
|
||||
```
|
||||
|
||||
`--approve --loop <name>` is the **only** way to clear a halt. `--resume-loop` only clears `paused` (user-initiated), never a halt.
|
||||
|
||||
## 4. Brake Gate Checks (`--check-gate`)
|
||||
|
||||
Called at the top of every tick by `loop-runner.py`. Returns JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"ok": false,
|
||||
"reason": "halted:verifier_failed",
|
||||
"halt_reason": "verifier_failed",
|
||||
"remaining_iterations": 0,
|
||||
"remaining_budget_usd": null,
|
||||
"task_phase": "implement",
|
||||
"task_in_halt_loop": true,
|
||||
"out_of_scope_files": []
|
||||
}
|
||||
```
|
||||
|
||||
Gate checks, in order:
|
||||
|
||||
1. **Loop status** — must be `running`. Anything else halts the tick immediately.
|
||||
2. **Iteration count** — `iteration_count < max_iterations` from `loop.json`.
|
||||
3. **Budget** — if `max_budget_usd` is set (informational, remote-only), check the harness's reported cost (best-effort: read from a `cost.json` the harness writes; absence is non-fatal). Below 16k context = hard refuse (D13).
|
||||
4. **Task phase** — if `current_task` is set, that task's `.state` must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to `human_intervention` by some other path), halt as `human_intervention`.
|
||||
5. **Worktree drift** — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via `git diff --name-only main...HEAD` restricted to `file_scope`), halt as `drift_detected`.
|
||||
6. **Score plateau** — last N entries in `score_history` are flat or monotonically decreasing (where N = `score_plateau_window`). Trip → halt as `verifier_failed`.
|
||||
|
||||
All halts atomically set `status=halted`, `halt_reason=<reason>`, write to `.state.log`, and call `--pause-loop`'s schedule-disable step (see §6).
|
||||
|
||||
## 5. Verifier Prompt Shape (`loop-verifier.md`)
|
||||
|
||||
Intentionally minimal — a fresh-context LLM should be able to grade a single artifact without loading the framework's 11-file context.
|
||||
|
||||
```
|
||||
# Loop Verifier — {loop_name}
|
||||
|
||||
You are grading one artifact for loop `{loop_name}`.
|
||||
|
||||
## Task context (read-only)
|
||||
{task_brief} # from .automaton/tasks/<t>/RESEARCH.md or DESIGN.md, capped at 4k tokens
|
||||
|
||||
## Artifact under review
|
||||
{artifact_content} # the file(s) the Implement: role just wrote
|
||||
|
||||
## Last tick's hint
|
||||
{next_hint} # capped at 1k tokens, may be empty
|
||||
|
||||
## What to check
|
||||
{acceptance_criteria} # from loop.json, capped at 2k tokens
|
||||
|
||||
## Output (strict JSON, no prose)
|
||||
{
|
||||
"pass": <bool>,
|
||||
"score": <0.0-1.0>,
|
||||
"reasons": ["..."],
|
||||
"next_hint": "..."
|
||||
}
|
||||
|
||||
Score rubric:
|
||||
- 1.0 = acceptance_criteria fully satisfied, no defects
|
||||
- 0.7 = functionally complete, minor defects not in criteria
|
||||
- 0.4 = partial progress, criteria partially addressed
|
||||
- 0.0 = no useful progress, or artifact is empty/missing
|
||||
```
|
||||
|
||||
Tier budget (from `loop.json`): **4k minimum, 16k floor** (D13). Below 16k the runner refuses before invoking the verifier — no tiny-context verifier ever runs.
|
||||
|
||||
`Implement:` prompt is similar but emits the artifact to a path, not JSON. `Orchestrate:` prompt loads the verdict JSON and calls exactly one `status.py` operation (no edits).
|
||||
|
||||
## 6. Scheduler Unit Generation (`--install-schedule`)
|
||||
|
||||
Platform detection via `platform.system()`:
|
||||
- **Darwin** → `~/Library/LaunchAgents/com.automaton.loop.<name>.plist` with `StartInterval` = `interval_seconds`. The plist's `ProgramArguments` calls `automaton-loop-tick.sh` (generated in the loop dir, see below). `--pause-loop` renames the plist to `.disabled` (Launch Agents don't honor a disabled bit portably). `--resume-loop` renames it back and `launchctl load`s it.
|
||||
|
||||
- **Linux** → `crontab -l` is read, lines for this loop removed, new line added (`*/N minutes * * * * <automaton-loop-tick.sh>`), `crontab -` written back. `--pause-loop` removes the line; `--resume-loop` re-adds it.
|
||||
|
||||
- **Windows** → `schtasks /create /tn "AutomatonLoop_<name>" /tr "<automaton-loop-tick.bat>" /sc minute /mo <N> /f`. `--pause-loop` calls `schtasks /change /tn ... /disable`; `--resume-loop` calls `/enable`.
|
||||
`automaton-loop-tick.sh` (generated in the loop dir, chmod +x) is 3 lines. The name is self-documenting: when the scheduler unit (plist ProgramArguments / cron line / schtasks `/tr`) references the file path, the filename alone conveys "this is automaton's loop-tick entry point" — no need to parse the script to know its role.
|
||||
|
||||
```bash
|
||||
#!/usr/bin/env bash
|
||||
cd "<project_root>"
|
||||
python3 "<framework>/scripts/loop-runner.py" --mode tick --loop "<name>"
|
||||
```
|
||||
|
||||
This keeps the OS-specific unit trivial. All logic (gate check, role invocation, verdict parse, brakes) lives in Python.
|
||||
|
||||
`--daemon` opt-in: `loop-runner.py --mode daemon --loop <name>` runs a `time.sleep(interval)` loop calling `--mode tick` per iteration. For CI / shared servers without cron. v1 supports it; default is native.
|
||||
|
||||
## 7. `loop-runner.py --mode tick` Flow
|
||||
|
||||
```
|
||||
1. parse --loop <name> → load .state.loop + loop.json
|
||||
2. status.py --check-gate <name> --json → gate
|
||||
if !ok:
|
||||
log halt, exit 0 (clean exit; do not crash the scheduler)
|
||||
3. find_work(work_source):
|
||||
audit → status.py --audit --json, pick highest-severity unresolved
|
||||
backlog → read design/<area>/BACKLOG.md, pick top not-done item
|
||||
single → use current_task from .state.loop
|
||||
3.5 claim task (cross-loop ownership check — v1.1):
|
||||
if candidate != state.current_task:
|
||||
status.py --claim-loop-task <name> --task <candidate> --project <p>
|
||||
if non-zero exit → SKIP "task_claimed_by_other_loop" (transient; next tick retries)
|
||||
4. ensure worktree exists (if worktree=true):
|
||||
if worktree_path is null or path missing:
|
||||
git worktree add .automaton/loops/<name>/worktree -b loop/<name>
|
||||
(retry without -b if branch already exists)
|
||||
record worktree_path + worktree_branch in .state.loop
|
||||
if not a git repo or git unavailable: fall back to project root (WARNING)
|
||||
5. spawn Implement: session with loop-implement.md,
|
||||
cwd = worktree_path (or project root if --no-worktree)
|
||||
captures artifact path
|
||||
6. spawn Verify: session with loop-verifier.md,
|
||||
reads artifact, emits verdict JSON
|
||||
7. parse verdict (strict JSON, accept comments / ```json fences)
|
||||
on parse failure → halt as verifier_failed, no retry in v1
|
||||
on parse success, coerce defensively (see §6b below):
|
||||
* `pass` accepts bool or `"true"`/`"false"` strings (case-insensitive,
|
||||
whitespace-stripped). Other strings fall through to `bool(...)`.
|
||||
* `score` is clamped to `[0, 1]`. NaN / ±Infinity / non-numeric
|
||||
types default to `0.5` (neutral midpoint).
|
||||
8. append score to score_history (cap = score_plateau_window)
|
||||
9. spawn Orchestrate: session with loop-orchestrate.md
|
||||
inputs: verdict, current_task, current_phase
|
||||
executes exactly one status.py call: transition, approve (not auto; orchestrator refuses auto-approve), or escalate to human_intervention
|
||||
9.5 release on terminal phase (v1.1):
|
||||
re-read task .state; if phase is complete or human_intervention:
|
||||
state.current_task = None (released for other loops)
|
||||
10. write verdict to .state.log, increment iteration_count, update last_tick_at
|
||||
11. if verdict.pass == true:
|
||||
transition task to next phase (orchestrator decides which)
|
||||
if task reached complete: status=complete in .state.loop
|
||||
```
|
||||
|
||||
`--mode tick` is **idempotent in the failure case**: a crash mid-tick does not advance iteration_count and does not corrupt `.state.loop` (atomic write via tmp file, same pattern as `_write_state`).
|
||||
|
||||
### Lock serialization (v1.1 — `add-state-loop-lock`)
|
||||
|
||||
The runner holds a cross-process file lock (`_loop_lock`) over the entire tick critical section — from the `_gate` subprocess call through the step-10 state write. The lock file is `<loop_path>/.state.lock` (per-loop granularity). POSIX uses `fcntl.flock(LOCK_EX)`; Windows uses `msvcrt.locking(LK_LOCK, 1)`. Blocking acquire, no timeout in v1.1 (operators notice a wedged tick via `--loop-list` stale `last_tick_at`).
|
||||
|
||||
The same `_loop_lock` wraps the read-modify-write blocks in `cmd_pause_loop`, `cmd_resume_loop`, `cmd_approve_loop`, and `cmd_check_gate` (in `status.py`). `--create-loop` is intentionally unwrapped — there is no prior state to race against.
|
||||
|
||||
**Subprocess-deadlock avoidance (D-L6)**: the runner's `_gate` call spawns `status.py --check-gate` as a subprocess. If status.py's `cmd_check_gate` also acquired `_loop_lock`, it would deadlock waiting on the parent runner's held flock. To avoid this, the runner passes `$AUTOMATON_NO_LOOP_LOCK=1` in that subprocess's env ONLY (scoped to the `_gate` subprocess; harness subprocesses do NOT inherit it). status.py's `_loop_lock` checks the env var; if set, it yields without flocking (trusting the caller's outer lock). Standalone CLI users don't set the env var, so `--check-gate` invoked manually locks normally and serializes against `--pause-loop` etc.
|
||||
|
||||
**Same bypass for claim subprocess**: `--claim-loop-task` (step 3.5) is also spawned from inside the runner's `_loop_lock`. The runner passes `$AUTOMATON_NO_LOOP_LOCK=1` in the claim subprocess env for the same reason — `cmd_claim_loop_task` acquires `_loop_lock` in status.py, but the runner already holds it. The bypass env is set only for this subprocess; standalone `--claim-loop-task` invocations (e.g. from `--can-continue` or future operator tooling) lock normally.
|
||||
|
||||
**Forbidding re-entry**: `_loop_lock` is NOT re-entrant across processes. Audit callsites to ensure no nested `_loop_lock` within the same `with` block. All v1.1 callsites are flat — no nested locks.
|
||||
|
||||
**Harness contract implication**: harnesses invoked via `harness.command` should NOT call loop-control commands (`--pause-loop`, `--resume-loop`, `--approve --loop`, `--check-gate`) from inside a tick — that would deadlock waiting on the parent runner's lock. Task commands (`--transition`, `--approve --task`, `--can-edit`, `--scope-check`, `--task`, `--claim`) do NOT touch `.state.lock` and are safe. Future work: add this to `contracts/harness-integration.md`.
|
||||
|
||||
### Outputs retention (v1.1 — `add-outputs-retention`)
|
||||
|
||||
To bound `outputs/` directory growth (O5 from `add-loop-runner/BUG_REPORT.md`), the runner runs GC after step 10 inside `_loop_lock`, retaining only the last N tick groups (default 20). Configured via `loop.json`:
|
||||
|
||||
```json
|
||||
"outputs": {
|
||||
"retention": 20
|
||||
}
|
||||
```
|
||||
|
||||
- `retention` = 0 disables GC (unlimited, v1 behavior). Negative coerces to 0 with WARNING.
|
||||
- GC iterates `outputs/`, parses `tick{N}-` prefix via `^tick(\d+)-` regex, computes `cutoff = max_seen - retention + 1`, deletes files with tick index < cutoff.
|
||||
- Non-tick files (no `tick{N}-` prefix) are preserved.
|
||||
- GC failure (permissions, file-not-found mid-iteration) is logged as WARNING and swallowed — never crashes the tick.
|
||||
|
||||
## 8. Harness Invocation
|
||||
|
||||
`loop-runner.py` invokes the user's harness via a single configured command in `loop.json`:
|
||||
|
||||
```json
|
||||
"harness": {
|
||||
"command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"],
|
||||
"prompt_var": "{prompt}",
|
||||
"cwd_var": "{cwd}",
|
||||
"output_var": "{output}"
|
||||
}
|
||||
```
|
||||
|
||||
v1.1's default `harness.command` is `opencode run` -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the `command` list (single argv element per token, no shell expansion):
|
||||
|
||||
- `{prompt}` -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path.
|
||||
- `{prompt_content}` -- the resolved prompt file's text content as a single argv element. Safe under `subprocess.run` list mode; no shell quoting needed. Used by the default command since `opencode run` takes the message as a positional argument and has no `--prompt-file` flag.
|
||||
- `{cwd}` -- the working directory the harness should run in (the loop's project root or worktree).
|
||||
- `{output}`, `{artifact}` -- role-specific extras (the implement output path handed to verify).
|
||||
- `{verdict}`, `{current_task}`, `{current_phase}`, etc. -- other runtime extras; see `_resolve_prompt` below.
|
||||
|
||||
The default command does NOT hardcode a `--model` flag; the spawned `opencode run` inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set `harness.command` in their `loop.json`:
|
||||
|
||||
```json
|
||||
"harness": {"command": ["opencode", "run", "--model", "local-mlx/...", "--dir", "{cwd}", "{prompt_content}"]}
|
||||
```
|
||||
|
||||
The runner core has zero knowledge of which harness is invoked (D8: framework never inspects model/provider). Concrete adapters for non-opencode harnesses (Pi Dev, aider, Cursor, Cline, Copilot) are out of scope for v1 -- users override `harness.command` to match their harness's CLI shape. Examples:
|
||||
|
||||
```json
|
||||
// Pi Dev
|
||||
"harness": {"command": ["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]}
|
||||
|
||||
// aider
|
||||
"harness": {"command": ["aider", "--message", "{prompt_content}", "--yes"]}
|
||||
|
||||
// Generic shell wrapper (any tool that reads prompt from stdin)
|
||||
"harness": {"command": ["sh", "-c", "cat {prompt} | my-tool --cwd {cwd}"]}
|
||||
```
|
||||
|
||||
No new harness adapter is written in v1.1.
|
||||
|
||||
### Prompt Resolution and Token Substitution
|
||||
|
||||
Before building the harness command, the runner calls `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)`:
|
||||
|
||||
1. **File search**: checks `<loop_path>/<prompt_ref>` first (loop-local override), then `~/.automaton/prompts/<prompt_ref>` (framework default). If neither exists, returns the raw `prompt_ref` string (backward compat -- the harness receives the raw ref).
|
||||
2. **Content substitution**: reads the prompt file and substitutes content-level tokens in the prompt text:
|
||||
- `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` -- from the work source and `.state.loop`
|
||||
- `{current_task}`, `{current_phase}` -- from `.state.loop`
|
||||
- `{verdict}` -- the JSON verdict from the verifier (orchestrate role only)
|
||||
- `{artifact_content}` -- reads the file at `extras["artifact"]` (the implement output path) and substitutes its full text; empty string if the file is missing (verify role only)
|
||||
3. **Temp file write**: writes the substituted content to `<loop_path>/outputs/tickN-<role>-prompt.md`.
|
||||
4. **Return**: the temp file path replaces `{prompt}` in the harness command template.
|
||||
|
||||
This means the harness receives a fully-resolved prompt file with all context baked in -- no runtime token substitution needed inside the harness. The loop-local override path (`<loop_path>/<prompt_ref>`) allows per-loop prompt customization without modifying the framework prompts.
|
||||
|
||||
## 9. Self-Improvement Loop Template
|
||||
|
||||
`templates/loops/self-improvement/loop.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "self-improvement",
|
||||
"description": "Ticks against status.py --audit on the framework's own repo",
|
||||
"work_source": {"kind": "audit", "project": "~/.automaton/"},
|
||||
"roles": {
|
||||
"implement": {"prompt": "loop-implement.md", "tier": 16000},
|
||||
"verify": {"prompt": "loop-verifier.md", "tier": 8000},
|
||||
"orchestrate": {"prompt": "loop-orchestrate.md", "tier": 4000}
|
||||
},
|
||||
"brakes": {
|
||||
"max_iterations": 10,
|
||||
"max_budget_usd": null,
|
||||
"score_plateau_window": 3
|
||||
},
|
||||
"blast_radius": {"worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]},
|
||||
"acceptance_criteria": [
|
||||
"Audit findings resolved (no outstanding Cat-1/Cat-2/Cat-4 violations on the resolved task)",
|
||||
"All R-numbers from the task SPEC.md implemented",
|
||||
"Tests pass with no regressions",
|
||||
"Pipeline driven to complete"
|
||||
],
|
||||
"harness": {
|
||||
"command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"],
|
||||
"prompt_var": "{prompt}",
|
||||
"cwd_var": "{cwd}",
|
||||
"output_var": "{output}"
|
||||
},
|
||||
"schedule": {"kind": "native", "interval_seconds": 3600}
|
||||
}
|
||||
```
|
||||
|
||||
Installs default-on at `install.sh` time: `status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` then `status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"`. Both commands use `|| true` so the framework works even if loop creation fails. `update.sh` bootstraps the loop idempotently for existing users (checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`). Disabling: `status.py --pause-loop self-improvement --project ~/.automaton/`.
|
||||
|
||||
## 10. Tests (`tests/test_loops.py`)
|
||||
|
||||
Required by AGENTS.md ("tests required for any new Python code"). Covers:
|
||||
|
||||
1. `test_loop_tick_pass` — fixture task in `implement`, mock verifier returns PASS → loop ticks, task transitions to `code_review`, loop status stays `running`.
|
||||
2. `test_loop_halt_iterations` — set `max_iterations: 1`, mock verifier returns non-PASS → halts as `iterations_exhausted`.
|
||||
3. `test_loop_halt_verifier_failed` — three consecutive flat scores → halts as `verifier_failed`.
|
||||
4. `test_loop_halt_drift_detected` — fixture worktree edited outside `file_scope` → next `--check-gate` halts as `drift_detected`.
|
||||
5. `test_loop_approve_resume` — halted loop, `--approve --loop` clears halt, `resumed_count == 1`, next tick allowed.
|
||||
6. `test_loop_paused_does_not_tick` — `--pause-loop` → `--check-gate` returns not-ok; `--mode tick` exits 0 without doing work.
|
||||
7. `test_loop_untracked_refused` — no `.state.loop` → `--check-gate` returns error, `--mode tick` refuses.
|
||||
8. `test_self_improvement_installs_default_on` — fresh install → `.automaton/loops/self-improvement/` exists, schedule unit present, status `running`.
|
||||
9. `test_loop_no_tiny_context` — fixture with `tier: 8000` but `vram_detect` reports ≤16k available → runner refuses before invoking verifier.
|
||||
10. `test_loop_idempotent_after_crash` — mid-tick crash simulated → `.state.loop` unchanged, next tick proceeds normally.
|
||||
|
||||
Test fixtures use `tmp_path` and stub `subprocess.run` for the harness invocations — no live LLM calls in CI.
|
||||
|
||||
## 11. Per-Iteration Context Budget (Tier 1 fix)
|
||||
|
||||
`tech.md` requirement: the runner computes a per-tick context budget before invoking any role. Formula:
|
||||
|
||||
```
|
||||
available_kb = vram_detect.py --json | .available_context_kb
|
||||
tier_kb = role's tier from loop.json (capped to available_kb)
|
||||
floor_kb = 16000 # D13
|
||||
if tier_kb < floor_kb: refuse("context window below 16k floor")
|
||||
```
|
||||
|
||||
Double-headroom bug (vram_detect.py:642+:654): the existing `vram_detect.py` applies headroom twice. Fix in task `fix-context-sizing`: remove the inner application, keep only the outer one. The result is what the runner reads.
|
||||
|
||||
`max(0, ...)` clamp + fake 8k/6k defaults (vram_detect.py:651, :698, :699): replace with `None` + explicit refuse-when-zero in loop mode. Non-loop callers unaffected.
|
||||
|
||||
## 12. Non-Goals (v1) — explicit restatement
|
||||
|
||||
- No parallel mode (D6)
|
||||
- No auto-approve (D4)
|
||||
- No budget auto-halt on local models (D3 — informational only)
|
||||
- No new pip dependencies; stdlib only for new code
|
||||
- No network fetches except user-supplied git URL (D11)
|
||||
- No framework drafting its own designs (Scope 3 deferred)
|
||||
|
||||
## 13. LOCKED — v1 implementation order (8 bootstrap tasks)
|
||||
|
||||
Per the locked plan, the human creates 8 tasks up front. Implementation order respects dependencies:
|
||||
|
||||
| # | Task | Depends on |
|
||||
|---|---|---|
|
||||
| 1 | `fix-context-sizing` | (none) |
|
||||
| 2 | `add-status-brakes` | 1 (brakes need a working context budget) |
|
||||
| 3 | `add-loop-runner` | 2 (runner calls the brakes) |
|
||||
| 4 | `add-goal-mode` | 3 (`/goal` extends the runner) |
|
||||
| 5 | `add-blast-radius-scheduler` | 2 + 3 |
|
||||
| 6 | `add-loop-templates-onboarding` | 3 (templates reference the runner) |
|
||||
| 7 | `add-self-improvement-loop` | 2, 3, 6 (default-on template + install hook) |
|
||||
| 8 | `fix-install-update-flow` | (parallel, no dep on others) |
|
||||
|
||||
After task 7 lands → write `design/context-sizing/` skeleton + BACKLOG → first loop tick picks up Tier 2 work → handoff.
|
||||
Reference in New Issue
Block a user