Files
automaton/tasks/add-status-brakes/VERDICT.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

56 lines
3.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Verdict: add-status-brakes
**Status: PASS**
The task delivers the loop-engineering brakes layer (R1–R10) entirely inside `status.py`, with no new dependencies and no second enforcement surface. It is the foundation that tasks 3–7 build on; everything those tasks need to call (`--check-gate`, `--can-continue`, `--approve --loop`, `--can-edit --loop`, `--create-loop`, `--install-schedule`, `--loop-list`, `.state.log`) is now in place and unit-tested.
## Requirement coverage
| Req | Delivered | Tests |
|-----|-----------|-------|
| R1 `.state.loop` schema | All 13 fields, atomic tmp+rename | `TestStateLoopSchema` (2) |
| R2 `--create-loop` | kebab/dup/template rejection, name patch | `TestCreateLoop` (5) |
| R3 `--version`, `--approve --loop` | version regex; only halt-clear; `resumed_count++` | `TestVersionAndApprove` (4) |
| R4 `--can-continue` | running-only probe | `TestCanContinue` (3) |
| R5 `--check-gate` (6 gates) | First-failure halts + JSON | `TestCheckGate` (10) |
| R6 `--install-schedule` | Darwin/Linux/Windows dispatch + stub | `TestInstallSchedule` (3) |
| R7 `--can-edit --loop [--loop-worktree]` | Root residency + file_scope | `TestCanEditLoop` (4) |
| R8 `--transition` halt refusal | Owned-task scan | `TestTransitionHaltRefusal` (3) |
| R9 `--audit` Cat-6 + `--loop-list` | Runs even when no tasks; untracked/halted flag | `TestAuditAndList` (6) |
| R10 `.state.log` tick trail | ISO timestamps | `TestTickLog` (3), `TestPauseResume` (3) |
Total: 46 new tests. Suite: **310 passed** (was 264 + 46 new). No regressions. `python3 -m py_compile scripts/status.py` clean.
## Defense against the five loop deaths
- **drift** → `_gate_worktree_drift` (R5)
- **runaway** → `_gate_iterations` (R5)
- **bad verifier** → `_gate_score_plateau` (R5)
- **resource burn** → `_gate_budget` (R5, remote-only informational)
- **undetected halt** → R8 transition refusal + Cat-6 audit + gate halt-write
## Harness / OS / model agnosticism preserved
- All surface reachable via `status.py` subprocess + `--json`. No harness-specific code. Works with opencode or any harness (D8).
- `platform.system()` dispatches launchd/cron/schtasks; missing tools degrade gracefully (warn + skip, not crash). D13 honored.
- Framework never inspects model capability/size/provider — `--loop-mode` already refused sub-16k in task 1; this task does not consult any model field.
## Doc impact landed
- `AGENTS.md` Harness Integration modes block extended with the `--loop` worktree-scope mode (mode 5).
- `AGENTS.md` new "State Enforcement — Loops (v1)" section.
- `README.md` new "Loop Engineering (beta)" subsection with quick-reference commands.
- `CHANGELOG.md` `[unreleased]` entry for the brakes layer.
## Hardening items deferred (tracked)
- A6 `fcntl` lock on `.state.loop` → v1.1.
- A2 `--claim-loop-task` atomic ownership → task 3.
- O3 `blast_radius.base_branch` drift parameterization → task 5.
- O4 `_enable_schedule` Linux parity → task 5 / v1.1.
All four are explicit follow-ups in `BUG_REPORT.md` and `ADVERSARIAL_BUG_REPORT.md`; none block this task.
## Resolution
**PASS — proceed to `complete`.** Task `add-status-brakes` is the foundation for the loop v1 implementation. Tasks 3, 4, 5, 6, 7 can now be unblocked, each relying on the standardized `.state.loop` schema and the brakes gates this task ships.