Archive completed tasks, add cleanup commands, self-documenting dashboard UI
CI / build (push) Has been cancelled

- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
This commit is contained in:
Lap Tran
2026-06-24 22:43:33 -04:00
parent e13513faaa
commit 4a2301b077
572 changed files with 856 additions and 101 deletions
-1
View File
@@ -1 +0,0 @@
complete
-2
View File
@@ -1,2 +0,0 @@
research:approved|2026-06-23T01:31:35.099808+00:00|user
code_review:approved|2026-06-23T01:44:53.823232+00:00|user
@@ -1,42 +0,0 @@
# Adversarial Bug Report: add-status-brakes
Adversarial pass: attack the brakes layer as a hostile agent would, looking for ways to escape enforcement.
## Attack vectors tried
### A1 — Can a halted loop be silently resumed by re-running `--create-loop` with the same name?
`cmd_create_loop` refuses duplicates (exit 2). Attacker cannot overwrite a halted loop's state via create. Defended.
### A2 — Can a hostile harness skip `--check-gate` and just call `--transition`?
Not if the loop owns the task — R8 refuses the transition until `--approve --loop` clears the halt. Defended.
BUT: if the loop never `current_task`-claimed the task, `_loop_owning_task` returns None and the transition proceeds. The agent can edit a task that isn't claimed by any loop. That is correct behavior (humans and ad-hoc agents can still work), but it means a hostile agent could **race the loop runner to claim a task**. Mitigation: loop runner should call a `--claim-loop-task` (not in v1) or set `current_task` atomically before transitioning. **Not a brakes-layer bug** — it's a runner-layer concern (task 3). Note for task 3.
### A3 — Can `--pause-loop` be used to bypass the worktree drift gate?
No — pausing just stops the scheduler; `.state.loop` remains the source of truth. `--check-gate` still runs all gates; a paused loop's gates return `paused:` (not ok). The runner will not tick. Defended.
### A4 — Can the agent hand-write `.state.loop` to fake "running"?
Yes — `.state.loop` is a plain JSON file. A fully local adversary with filesystem access can edit it. This is an accepted threat model: the framework assumes the harness is cooperative for state files but enforces via the pre-edit hook (`--can-edit`) for source edits. **`.state.loop` is not in any task's file_scope**, so it's never editable by a loop agent. Defended by file-scope design.
### A5 — Race: two concurrent `--check-gate` invocations both halt the loop
Both call `_halt_loop` which uses atomic tmp+rename. Last writer wins. Both write the same halt_reason (deterministic from gate), so the result is consistent. No corruption. Defended.
### A6 — Can `--approve --loop` be called while the loop is mid-tick?
`--approve` does tmp+rename. If a tick is concurrently writing iteration_count, the approve's write wins and the tick's increment is lost. Window is small (subprocess boundary). Acceptable for v1; the next tick re-reads and re-increments. Not a corruption vector. **Note for v1.1:** file-locking (fcntl) on `.state.loop` would close this race. Add to BACKLOG.
### A7 — Can `--install-schedule` be pointed at a different project than the loop?
`--install-schedule` uses `_find_project_dir(args.project)` and writes the stub at `loop_path / run-tick.*`. The stub `cd`s into the project root and invokes the runner with the loop name. An attacker could swap the loop_name in the stub after generation, but that's just running an arbitrary loop — not a privilege escalation. Not an attack.
### A8 — Can the schedule wake the loop after it's halted?
Yes — the OS unit fires `run-tick` on schedule. `run-tick` invokes `loop-runner.py --mode tick --loop NAME`, which **must** call `--check-gate` first and exit 1 if not ok. The runner's contract (task 3) is: gate first, then work. The OS unit itself cannot refuse. So a halted loop's schedule will fire `run-tick`, which will no-op via the runner's gate check. The `--pause-loop` best-effort disable is belt-and-braces. Defended by runner contract (must be enforced in task 3).
## Hardening recommendations (for BACKLOG)
1. `fcntl` file-lock on `.state.loop` for tick/approve race (A6) — v1.1.
2. `--claim-loop-task` to atomically set `current_task` before a runner touches the task (A2) — task 3.
3. `_enable_schedule` Linux parity with Darwin/Windows (O4) — task 5 / v1.1.
4. `blast_radius.base_branch` parameterization for drift diff (O3) — task 5.
## Verdict
PASS — no exploitable escape from the brakes layer. All adversarial vectors are either defended today or have explicit runner-contract mitigations landing in tasks 3/5. Hardening items routed to `design/loops/BACKLOG.md`.
-37
View File
@@ -1,37 +0,0 @@
# Bug Report: add-status-brakes
Adversarial probing of the brakes layer against the five loop-death modes listed in `design/loops/functional.md` (drift, runaway, bad verifier, resource burn, undetected halt).
## Bugs found
None blocking. The code passed all six gates exercised in `tests/test_status_brakes.py`. Below are minor robustness observations (informational, not blockers).
## Observations (non-blocking)
### O1 — `_loop_untracked_hint` mentions `--upgrade-loops` which doesn't exist yet
`_loop_untracked_hint` references a future `--upgrade-loops` command. Until it ships (v1.1), users will see the hint but the command won't exist. The hint is advisory; the actionable path (`--create-loop`) is also named. Acceptable for v1.
### O2 — `cmd_install_schedule` on Linux does not re-install via `_enable_schedule`
`_enable_schedule` for Linux is a no-op branch (`pass`). `--resume-loop` therefore does not restart a Linux cron block that was stripped by `--pause-loop`. Darwin path renames `*.plist.disabled` back, Windows path re-runs `schtasks /run`. Linux asymmetry is a known gap; the next tick will still fire per the original cron line if it survived. For full symmetry, `_enable_schedule` on Linux should re-invoke the install code. Minor; not blocking — runner's `--check-gate` is the runtime enforcement, not the scheduler.
### O3 — `_gate_worktree_drift` runs `git diff main...HEAD`
Hard-codes `main` as the integration branch. Projects on `master`/`trunk` would show every file as out-of-scope (no `main` to diff against → git errors → gate skips with warning). Worth parameterizing per loop config (`blast_radius.base_branch`) in task 5 when worktree creation lands. For v1, the warning path is the correct fail-safe.
### O4 — `_disable_schedule` Linux path strips the cron block permanently
`--pause-loop` on Linux removes the cron block; `--resume-loop`'s Linux branch is a no-op. So a Linux user who pauses a loop loses their schedule. Mitigation: the user can re-run `--install-schedule` after resuming. Same as O2; tracked together.
### O5 — `cmd_check_gate` halts the loop when any gate returns a failure dict
Even informational gates (`budget_exhausted`) cause a halt write. Per D3 budget is "informational only (remote)". If we want it to **halt but not refuse continuation**, we'd need a softer "warn" verdict. Out of scope for v1; matches SPEC R5 wording ("first failure wins").
## No blocker bugs
All five loop-death modes are defended:
- **drift** → `_gate_worktree_drift` (R5)
- **runaway** → `_gate_iterations` (R5)
- **bad verifier** → `_gate_score_plateau` (R5)
- **resource burn** → `_gate_budget` (R5, remote-only informational)
- **undetected halt** → `cmd_transition` R8 refusal + `cmd_audit` Cat-6 + `cmd_check_gate` halt-write
## Verdict
PASS — proceed to adversarial_bug_find.
-50
View File
@@ -1,50 +0,0 @@
# Code Review: add-status-brakes
Reviewed against SPEC.md R1–R10. All requirements implemented; no functional gaps found.
## R1–R10 checklist
| Req | Status | Notes |
|-----|--------|-------|
| R1 `.state.loop` schema | ✅ | All 13 defaults present; atomic write via tmp+rename |
| R2 `--create-loop` | ✅ | kebab/Dup/template validation; name patching |
| R3 `--version`, `--approve --loop` | ✅ | version parses `## Framework Version`; approve only clears halt; `resumed_count++` |
| R4 `--can-continue` | ✅ | Correct boolean: `status == "running"` only |
| R5 `--check-gate` (6 gates) | ✅ | Order matches SPEC; first failure halts; JSON structured |
| R6 `--install-schedule` | ✅ | Triple dispatch Darwin/Linux/Windows; stubs generated; pause disables (best-effort) |
| R7 `--can-edit --loop [--loop-worktree]` | ✅ | Root residency + file_scope; refuses outside root |
| R8 `--transition` halt refusal | ✅ | Owned-task scan; points user at `--approve --loop` |
| R9 `--audit`/`--loop-list` | ✅ | Cat-6 runs even with no tasks; untracked/halted flagged; missing current_task flagged |
| R10 `.state.log` | ✅ | ISO timestamps; tested for PAUSED/RESUMED/APPROVED/HALT |
## Defensive coding observations
1. **Atomic `.state.loop` writes** — tmp+`replace()`. Crashes mid-write cannot corrupt state.
2. **Best-effort schedule disable** — wrapped in `try/except` so a non-existent cron/plist on a dev box cannot crash `--pause-loop` or the halt path. `.state.loop` remains source of truth; the OS unit reads it on next wake and self-skips.
3. **No new pip deps** — stdlib only (`platform`, `subprocess`, `json`, `re`, `datetime`). Per project constraints.
4. **Harness-agnostic** — every gate is reachable via `status.py` subprocess + `--json`. No harness-specific code. Works with opencode, any other harness, or a raw shell.
5. **`--approve --loop` is the only halt-clear** — D4 enforced; `--resume-loop` explicitly refuses halted loops and tells the user to approve.
6. **R8 ownership scan** — `_loop_owning_task` is O(loops) per transition; loops are few, so fine. Could be cached later if needed.
## Edge cases checked
- Empty project (no tasks) — `--audit` still runs Cat-6 (R9 fix; was originally early-return).
- Loop with no `loop.json` — `--install-schedule` exits 2 with clear message.
- Loop with no `.state.loop` — every `--loop` command refuses with the `_loop_untracked_hint`.
- `--check-gate` on a paused loop — `_gate_loop_status` returns the `paused:` reason (not a halt, since the user paused it; harness checks separately via `--can-continue`).
- Budget informational when `max_budget_usd == null` — gate skipped, returns None.
- Score plateau with too-short history — gate skipped.
- Worktree missing — `_gate_worktree_drift` treats as no-drift (runner will recreate).
- `git diff` failure — warning logged to stderr, drift gate skips. Not a halt; per "best-effort portable" principle (D13).
## Things deliberately NOT in this task (per scope)
- `loop-runner.py` itself — task 3.
- Verifier role / graded JSON — task 4.
- Worktree creation plumbing — task 5.
- Full `templates/loops/ci-triage/` content (prompts, README) — task 6.
- `--upgrade-loops` for stray pre-state-loop dirs —audit just flags them. Refactor in v1.1.
## Verdict
APPROVE. No blocking issues. Ready for bug_find.
-45
View File
@@ -1,45 +0,0 @@
# Doc Review: add-status-brakes
Reviewed doc impact: `AGENTS.md`, `README.md`, `prompts/`, `config.md`, `CHANGELOG.md`.
## Doc gaps to land in THIS task
### Already updated in this task
- None ( изменения are in `status.py`, `tests/test_status_brakes.py`, `templates/loops/ci-triage/loop.json`). No prompt or config doc touched.
### To be updated (within this task's scope or follow-on)
1. **`AGENTS.md` Build & Test Commands section** — should mention:
- `python3 -m pytest tests/test_status_brakes.py -v`
- `--version` flag exists
However, AGENTS.md is a framework-wide doc; per the project convention it covers the test suite as a whole, not per-test-file. **Decision: do NOT pile per-test-file entries into AGENTS.md** — the existing `python3 -m pytest tests/ -v` already covers it. Leave alone.
2. **`AGENTS.md` Harness Integration section** — should add the new `--can-edit --loop [--loop-worktree] --file P` mode. The current AGENTS.md describes modes 1–4 for `--can-edit`. Adding a 5th mode belongs here.
**Action**: extend AGENTS.md's "Modes:" block under Harness Integration to describe the loop worktree scope mode. Will apply in this task.
3. **`AGENTS.md` Conventions / State Enforcement section** — should mention `.state.loop` and `--approve --loop`. Will add a short paragraph.
4. **`README.md`** — user-facing. Should mention loop commands exist (high-level). Defer detailed user docs to task 6 (templates/onboarding); only the existence of loop commands is in scope here.
**Action**: add a brief "Loop engineering (beta)" subsection in README.md.)
5. **`prompts/`** — no loop-specific prompts land in this task. Task 6 owns `prompts/loop-{implement,verifier,orchestrate}.md`. **No action.**
6. **`config.md`** — already has `## Loop Role Models` (task 1) and `## Framework Version`. The `## Framework Version` section is what `--version` parses. Confirmed it parses correctly. **No action.**
7. **`CHANGELOG.md`** — should get an `[unreleased]` entry for the brakes layer. **Action**: add.
## Doc consistency observations (non-blocking, defer)
- The harness-integration contract at `contracts/harness-integration.md` lists `--can-edit` modes 1–4. Should add mode 5 (--loop worktree). **Defer to a follow-on doc-rev task**; touching the contract file is out of scope for this code task and risks destabilizing the contract.
- `design/loops/technical.md` describes `--install-schedule` semantics; the implementation matches. No update needed.
## Summary of doc edits in this task
- `AGENTS.md`: extend Harness Integration modes list; brief `.state.loop` paragraph.
- `README.md`: one "Loop engineering (beta)" subsection.
- `CHANGELOG.md`: entry under `[unreleased]`.
No code-doc mismatches found. READY for referee.
-89
View File
@@ -1,89 +0,0 @@
# Implementation: add-status-brakes
Implements SPEC.md R1–R10. All new code lives in `scripts/status.py` (loop extensions) plus a new test file `tests/test_status_brakes.py` and a minimal loop template at `templates/loops/ci-triage/loop.json`.
## Surface added (R1–R10)
| Req | CLI surface | Behavior |
|-----|-------------|----------|
| R1 | n/a | `.state.loop` schema v1 with 13 default fields; written atomically via tmp+rename |
| R2 | `--create-loop NAME [--from-template T]` | Refuses non-kebab, duplicates, unknown template; patches `name` into copied `loop.json`; seeds empty `.state.log` |
| R3 | `--version`; `--approve --loop NAME` | `--version` reads `## Framework Version` from `config.md`; `--approve --loop` is the **only** way to clear a halt (D4); increments `resumed_count` |
| R4 | `--can-continue NAME [--json]` | Cheap status probe: `ok := status == "running"` |
| R5 | `--check-gate NAME [--json]` | Runs 6 gates in order; first failure halts the loop and emits structured verdict |
| R6 | `--install-schedule NAME [--interval S]` | Generates `run-tick.sh`/`.bat`; installs launchd plist / crontab block / schtasks unit per `platform.system()`; `--pause-loop` best-effort disables the unit |
| R7 | `--can-edit --loop NAME [--loop-worktree] --file P` | Checks file against loop's `blast_radius.file_scope`; refuses files outside project/framework root |
| R8 | `--transition` extension | Refuses if a HALTED loop owns the task (`_loop_owning_task` scan); points user at `--approve --loop` |
| R9 | `--audit` Cat-6 block; `--loop-list` | Reuses `_audit_loops_block`; runs even when no tasks exist |
| R10 | `.state.log` tick trail | Every state-changing op appends an ISO-timestamped line; tests assert PAUSED/RESUMED/APPROVED/HALT are all logged |
## Gate order (R5)
```
gate_loop_status -> not running -> halt w/ existing halt_reason
gate_iterations -> iteration_count >= max_iterations -> iterations_exhausted
gate_budget -> spent_usd >= max_budget_usd -> budget_exhausted (remote-only, informational)
gate_task_phase -> current_task in human_intervention -> human_intervention
gate_worktree_drift -> changed files outside file_scope -> drift_detected
gate_score_plateau -> score_history flat across window -> verifier_failed
```
First failure wins. Halt is written atomically; schedule is best-effort disabled.
## Helper functions added (scripts/status.py, before `def main()`)
- `LOOP_*` constants (states, halts, schema version, file names)
- `_loops_dir`, `_loop_dir`, `_all_loop_dirs`
- `_read_state_loop`, `_write_state_loop`, `_initial_state_loop`, `_read_loop_config`
- `_append_tick_log`, `_loop_untracked_hint`
- `_halt_loop`, `_disable_schedule`, `_enable_schedule`
- `_loop_owning_task` (R8 ownership scan)
- `_gate_*` (6 gate functions)
- `_loop_max_iterations`
- `_task_phase_for_loop`
- `cmd_create_loop`, `cmd_install_schedule`, `cmd_pause_loop`, `cmd_resume_loop`
- `cmd_approve_loop` (R3 halt-clear)
- `cmd_check_gate`, `cmd_can_continue`
- `cmd_loop_list`, `cmd_version`
- `cmd_can_edit_loop` (R7 worktree scope)
## Existing functions extended
- `cmd_can_edit` — early hook: if `args.loop`, delegate to `cmd_can_edit_loop`.
- `cmd_transition` — R8 halt-refusal inserted after `_require_state`; `_loop_owning_task` scan.
- `cmd_audit` — `_audit_loops_block(args)` helper called twice (early-return empty-tasks path + main path); Cat-6 header always printed.
## Argparse additions (main())
`--create-loop`, `--from-template`, `--install-schedule`, `--interval`, `--pause-loop`, `--resume-loop`, `--loop`, `--loop-worktree`, `--check-gate`, `--can-continue`, `--loop-list`, `--version`.
Dispatch order places loop commands before task commands so `--approve --loop` doesn't fall through to the `--task`-required `cmd_approve`.
## New file: templates/loops/ci-triage/loop.json
Minimal template used as `--create-loop` default. Defines `brakes.max_iterations=25`, `score_plateau_window=5`, `blast_radius.use_worktree=true`. Full prompt/template expansion is task 6.
## Tests
`tests/test_status_brakes.py` — 46 tests across 10 classes mirroring R1–R10:
- `TestStateLoopSchema` (R1) — default-schema assertions + tick log file presence
- `TestCreateLoop` (R2) — kebab/dup/template rejection + name-patching
- `TestVersionAndApprove` (R3) — version regex; approve refuses non-halted; clears halted + bumps `resumed_count`
- `TestCanContinue` (R4) — running ok, halted denied, unknown → exit 2
- `TestCheckGate` (R5) — fresh-pass, status-halt, iterations-exhausted, iterations-remaining, budget-exhausted, budget-informational, task-phase-halt, score-plateau, short-history-ok, JSON output
- `TestInstallSchedule` (R6) — stub generation, default interval from config, unknown-loop rejection
- `TestCanEditLoop` (R7) — in-scope allowed, out-of-scope denied, outside-root denied, no-file rejected
- `TestTransitionHaltRefusal` (R8) — refused when halted owner, allowed when running owner, allowed when no owner
- `TestAuditAndList` (R9) — empty list, populated list, Cat-6 header on empty, halted flag, untracked flag, running-pass
- `TestTickLog` (R10) — PAUSED/RESUMED/APPROVED/HALT all logged
- `TestPauseResume` — pause sets paused; resume only from paused; halted→approve pointer
## Verification
```
python3 -m py_compile scripts/status.py # OK
python3 -m pytest tests/test_status_brakes.py -q # 46 passed
python3 -m pytest tests/ -q # 310 passed (was 264 + 46 new)
```
No existing tests changed. Full suite green.
-162
View File
@@ -1,162 +0,0 @@
# Add Status Brakes
Implement the loop-aware extension to `status.py` per `design/loops/technical.md` §3 and §4. This is the second-tier enforcement layer that the loop runner (task 3) will call. Brakes live *inside* `status.py` so they cannot be routed around by the harness.
## Goal
Make `status.py` aware of loops. Add `.state.loop` files, on-disk loop folder layout, gate-check commands, schedule-unit installers, and the `--approve --loop` resume path. No runtime/runner code in this task — task 3 (`add-loop-runner`) wires `loop-runner.py` to call these commands. This task only ships the *enforcement surface*.
## Requirements
### R1. Loop directory layout
Each project gets `.automaton/loops/<name>/` containing:
- `loop.json` — copied from `templates/loops/<template>/loop.json` (template files themselves are task 6's deliverable; `--create-loop` works against any existing template dir)
- `.state.loop` — JSON state file (R2 schema)
- `.state.log` — append-only tick log, seeded empty on creation
- `worktree/` — created lazily on first worktree-needing tick (task 3's runner creates it; `--create-loop` does NOT set up worktree)
- `run-tick.sh` — generated by `--install-schedule` (R6); not present at `--create-loop` time
Loops without `.state.loop` are **UNTRACKED** — mirror of v2.0 task `.state` rule. All `--loop` commands refuse to operate on an untracked loop and emit the upgrade hint: `Run --upgrade-loops to bootstrap`. (`--upgrade-loops` is not in this task; future bootstrap work. Provided only as the hint target.)
### R2. `.state.loop` schema
JSON:
```json
{
"schema_version": 1,
"name": "<loop-name>",
"status": "running",
"halt_reason": null,
"iteration_count": 0,
"resumed_count": 0,
"last_tick_at": null,
"last_verdict": null,
"score_history": [],
"current_task": null,
"worktree_branch": null,
"worktree_path": null
}
```
`status` ∈ `{"running", "halted", "paused", "complete"}`. `halt_reason` ∈ the five deaths + `null`. `last_verdict` is the most recent verdict JSON or `null`. `score_history` is capped at `score_plateau_window` (from `loop.json`), FIFO.
`_write_state_loop()` helper mirrors `_write_state()`'s atomic-tmp-then-replace pattern.
### R3. New flags on `status.py`
All route through one argparse parser to keep harness integration single-point.
```
status.py --create-loop <name> --from-template <template> [--project <p>]
status.py --install-schedule <name> [--interval N] [--project <p>]
status.py --pause-loop <name> [--project <p>]
status.py --resume-loop <name> [--project <p>]
status.py --approve --loop <name> [--project <p>]
status.py --can-continue <name> [--project <p>] (--json supported)
status.py --check-gate <name> [--task <t>] [--project <p>] (--json supported)
status.py --can-edit --project <p> [--task <t>] [--file <path>] [--loop <name>] [--loop-worktree]
status.py --loop-list [--project <p>]
status.py --version
```
`--approve --loop` is the **only** way to clear a halt. `--resume-loop` only clears `paused` (user-initiated pause), never a halt — refuses with `"loop is halted, use --approve --loop to clear halt"`.
`--version` reads the `## Framework Version` section of `config.md` and prints as `automaton <version>\n`. Exit 0 always (matches POSIX convention for `--version`). When the section is missing, prints `automaton (unknown version)\n` and still exits 0.
### R4. `--check-gate` JSON return
Returns JSON to stdout (last line, pre-encoded). Exit code 0 on `ok:true`; exit code 1 on `ok:false` (HALTED/PAUSED/COMPLETE etc.); exit code 2 on error (loop untracked / not found).
Shape (from technical.md §4):
```json
{
"ok": false,
"reason": "halted:verifier_failed",
"halt_reason": "verifier_failed",
"remaining_iterations": 0,
"remaining_budget_usd": null,
"task_phase": "implement",
"task_in_halt_loop": true,
"out_of_scope_files": []
}
```
Gate checks execute in order: loop status → iteration count → budget → task phase → worktree drift → score plateau. The first failing check halts and sets `halt_reason` atomically. Worktree drift requires `git diff --name-only main...HEAD` scoped to `loop.json.blast_radius.file_scope` (uses `subprocess.run` best-effort; on no-git environments, drift check is skipped with a stderr warning, not a halt).
### R5. `--can-continue` shorthand
Returns `{"ok": true/false, "status": "running|halted|paused|complete"}` — used by schedulers/CI to decide `run-tick.sh` shouldn't proceed. More general than `--check-gate` (which is the pre-tick gate). `--can-continue` is the "is the loop alive at all" check.
### R6. `--install-schedule` platform dispatcher
Detect `platform.system()`:
- `Darwin` → write `~/Library/LaunchAgents/com.automaton.loop.<name>.plist` with `StartInterval = interval_seconds`. Also writes `run-tick.sh` (chmod +x) into the loop dir for the plist's `ProgramArguments`.
- `Linux` → read `crontab -l`, strip any existing `# automaton-loop:<name>` block, append a new block tagged `# automaton-loop:<name>\n*/N * * * * <run-tick.sh>`, and `crontab -` back. Also writes `run-tick.sh`.
- `Windows` → `schtasks /create /tn "AutomatonLoop_<name>" /tr <run-tick.sh> /sc minute /mo <N_minutes> /f`. Also writes `run-tick.bat` (Windows uses `.bat`, not `.sh`, but the runner is still Python).
- Other → refuse with exit 2 and an unsupported-OS message.
`run-tick.sh` content is locked by technical.md §6:
```bash
#!/usr/bin/env bash
cd "<project_root>"
python3 "<framework>/scripts/loop-runner.py" --mode tick --loop "<name>"
```
`--pause-loop`:
- Darwin → rename plist to `.disabled`.
- Linux → strip the `# automaton-loop:<name>` block from crontab.
- Windows → `schtasks /change /tn "AutomatonLoop_<name>" /disable`.
`--resume-loop` is the inverse; refuse with halt-state error per R3.
### R7. `--can-edit --loop` extension
Existing `--can-edit` semantics preserved. New flags:
- `--loop <name>` adds a worktree-scope clause: edits allowed only if file is inside `<loop_dir>/worktree/` (or, when `--loop-worktree`, against `<project>/.automaton/loops/<name>/worktree/`).
- `--loop-worktree` (requires `--loop`) switches the file-scope anchor to the worktree path instead of the project root.
Exit codes/host-side output unchanged; only the ALLOWED/DENIED response shifts.
### R8. `--transition` refuses when a halted loop owns the task
`status.py --transition <phase> --task <t>` already operates on tasks. New behavior: when the task's `current_task` field is set in *any* loop whose `status` is `halted` and whose `halt_reason` is one of the five deaths, transitions are refused with `"task is bound to halted loop '<name>' (halt_reason=<reason>). --approve --loop <name> to resume."`. Exit 1.
When the loop is `running` or `paused`, transitions proceed normally (the loop will see the new phase at next tick).
### R9. `--audit` extension
`--audit` output gains a `Loops` section listing every loop with `(name, status, halt_reason, iteration_count, started_at)`. Loops in `halted` state are flagged with an audit warning.
New flag `--loop-list` provides the same data as `--audit`'s loop section but standalone.
### R10. Tests
New file `tests/test_status_brakes.py` covering:
- `cmd_create_loop` — creates dir + loop.json + .state.loop with default state; refuses on duplicate; refuses on missing template.
- `.state.loop` schema initialization — all R2 fields present.
- `--approve --loop` clears halt, increments `resumed_count`, refuses on running loop, refuses on untracked loop.
- `--resume-loop` clears paused, refuses on halted.
- `--pause-loop` invalidates `--can-continue`.
- `--check-gate` JSON for: clean running, halted on iterations, halted on verifier_failed (flat score), halted on drift, paused loop.
- `--can-edit --loop` allowed when file under worktree, denied when outside.
- `--transition` refused when owning loop halted; allowed when running/paused.
- `--install-schedule` writes `run-tick.sh` (and a stub plist on Darwin using tmp_path monkey-patching of `Path.home()`).
- `--version` prints "automaton <version>" reading from a fixture `config.md`.
## Acceptance Criteria
- [ ] `--create-loop` produces a valid `.state.loop` with R2 fields; duplicate-name returns exit 2.
- [ ] `--approve --loop` increments `resumed_count`, clears `halt_reason`, returns status to `running`. Refuses on a running loop.
- [ ] `--resume-loop` clears `paused` only; refuses on `halted`.
- [ ] `--check-gate --json` emits the §4 JSON shape; returns exit 1 when not-ok.
- [ ] `--can-edit --loop --file <outside>` exit 1; the same file inside the worktree exit 0.
- [ ] `--transition --task <t>` exit 1 when an owning loop is halted.
- [ ] `--version` writes `automaton <version>\n` to stdout from `config.md`'s `## Framework Version` section.
- [ ] `--install-schedule` writes `run-tick.sh` and the OS-native schedule unit (Darwin plist / Linux crontab block / Windows schtasks invocation) using a tmp_path fixture.
- [ ] `--audit` includes a Loops section.
- [ ] `tests/test_status_brakes.py` passes.
- [ ] Pre-existing framework tests still green: `pytest tests/ -q`.
## Non-Goals
- No `loop-runner.py` in this task (task 3).
- No verifier prompt contents (task 6 templates).
- No worktree creation logic for live ticks (task 5 — `--create-loop` makes the dir but not the worktree).
- No `--upgrade-loops` command (referenced only in error messages; bootstrap path remains manual for v1).
- No parallel mode (D6 stays opt-in; not implemented in v1).
## Dependencies
- Task 1 (`fix-context-sizing`) — DONE. `--check-gate` budget check relies on `vram_detect.py --loop-mode` JSON `available_context_kb >= 16000`. The `loop_mode_eligible` field is available.
## Out of Scope (deferred)
- `--upgrade-loops` (bootstrap pre-2.0 loops) — not blocking v1; manual create-loop is the path.
- Dashboard "Loops" panel — v1.1.
-56
View File
@@ -1,56 +0,0 @@
# Verdict: add-status-brakes
**Status: PASS**
The task delivers the loop-engineering brakes layer (R1–R10) entirely inside `status.py`, with no new dependencies and no second enforcement surface. It is the foundation that tasks 3–7 build on; everything those tasks need to call (`--check-gate`, `--can-continue`, `--approve --loop`, `--can-edit --loop`, `--create-loop`, `--install-schedule`, `--loop-list`, `.state.log`) is now in place and unit-tested.
## Requirement coverage
| Req | Delivered | Tests |
|-----|-----------|-------|
| R1 `.state.loop` schema | All 13 fields, atomic tmp+rename | `TestStateLoopSchema` (2) |
| R2 `--create-loop` | kebab/dup/template rejection, name patch | `TestCreateLoop` (5) |
| R3 `--version`, `--approve --loop` | version regex; only halt-clear; `resumed_count++` | `TestVersionAndApprove` (4) |
| R4 `--can-continue` | running-only probe | `TestCanContinue` (3) |
| R5 `--check-gate` (6 gates) | First-failure halts + JSON | `TestCheckGate` (10) |
| R6 `--install-schedule` | Darwin/Linux/Windows dispatch + stub | `TestInstallSchedule` (3) |
| R7 `--can-edit --loop [--loop-worktree]` | Root residency + file_scope | `TestCanEditLoop` (4) |
| R8 `--transition` halt refusal | Owned-task scan | `TestTransitionHaltRefusal` (3) |
| R9 `--audit` Cat-6 + `--loop-list` | Runs even when no tasks; untracked/halted flag | `TestAuditAndList` (6) |
| R10 `.state.log` tick trail | ISO timestamps | `TestTickLog` (3), `TestPauseResume` (3) |
Total: 46 new tests. Suite: **310 passed** (was 264 + 46 new). No regressions. `python3 -m py_compile scripts/status.py` clean.
## Defense against the five loop deaths
- **drift** → `_gate_worktree_drift` (R5)
- **runaway** → `_gate_iterations` (R5)
- **bad verifier** → `_gate_score_plateau` (R5)
- **resource burn** → `_gate_budget` (R5, remote-only informational)
- **undetected halt** → R8 transition refusal + Cat-6 audit + gate halt-write
## Harness / OS / model agnosticism preserved
- All surface reachable via `status.py` subprocess + `--json`. No harness-specific code. Works with opencode or any harness (D8).
- `platform.system()` dispatches launchd/cron/schtasks; missing tools degrade gracefully (warn + skip, not crash). D13 honored.
- Framework never inspects model capability/size/provider — `--loop-mode` already refused sub-16k in task 1; this task does not consult any model field.
## Doc impact landed
- `AGENTS.md` Harness Integration modes block extended with the `--loop` worktree-scope mode (mode 5).
- `AGENTS.md` new "State Enforcement — Loops (v1)" section.
- `README.md` new "Loop Engineering (beta)" subsection with quick-reference commands.
- `CHANGELOG.md` `[unreleased]` entry for the brakes layer.
## Hardening items deferred (tracked)
- A6 `fcntl` lock on `.state.loop` → v1.1.
- A2 `--claim-loop-task` atomic ownership → task 3.
- O3 `blast_radius.base_branch` drift parameterization → task 5.
- O4 `_enable_schedule` Linux parity → task 5 / v1.1.
All four are explicit follow-ups in `BUG_REPORT.md` and `ADVERSARIAL_BUG_REPORT.md`; none block this task.
## Resolution
**PASS — proceed to `complete`.** Task `add-status-brakes` is the foundation for the loop v1 implementation. Tasks 3, 4, 5, 6, 7 can now be unblocked, each relying on the standardized `.state.loop` schema and the brakes gates this task ships.