- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
2.1 KiB
2.1 KiB
Doc Review: harden-parse-verdict
Docs touched
design/loops/technical.md§7 — tick-flow step 7 (parse verdict): added 4-line inline block documenting the defensive coercion (pass string acceptance; score clamp + NaN/inf/non-numeric → 0.5).design/loops/functional.md§10 — Verifier Contract: annotatedpass(bool) definitive + runner accepts"true"/"false"strings; annotatedscore(0.0–1.0) clamp + NaN/inf/non-numeric → 0.5 neutral.CHANGELOG.md— new[unreleased]"Fixed —parse_verdictdefensive coercion" block above the existingadd-state-loop-lockandfix-harness-command-templateblocks.
Docs NOT touched (intentional)
AGENTS.md: parse_verdict is not a user-visible CLI surface; the hardening doesn't change phase enforcement,.state.loop, or any contract that harness integrators need to know. The Verifier Contract lives indesign/loops/functional.md§10; AGENTS.md already points to design docs at the top. No edit.README.md: user-facing README doesn't enumerateparse_verdictinternals; loop monitoring table mentions verdicts as a concept, not the parser. No edit.prompts/loop-verifier.md: contract was alreadybool pass+score 0.0–1.0. The hardening is belt-and-suspenders against malformed output, not a contract change. The prompt's strict-JSON directive stays authoritative. No edit.templates/loops/self-improvement/loop.json: no schema change. No edit.
Cross-references
tasks/add-loop-runner/BUG_REPORT.mdO6 — the original finding — now closed by this task. The CHANGELOG entry explicitly references it.tasks/add-loop-runner/ADVERSARIAL_BUG_REPORT.mdA6 — the score- clamping observation — also closed by this task. The CHANGELOG entry references the clamping.tasks/harden-parse-verdict/BUG_REPORT.mdO3 — notes that clamping improves plateau detection (a tighterscore_historyrange makes plateau more honest). Cross-referenced from the CHANGELOG.
Verdict
Docs are in sync with the implementation. Proceed to referee.