- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
3.3 KiB
3.3 KiB
Code Review: harden-parse-verdict
SPEC coverage
| Requirement | Status |
|---|---|
R1 — Coerce pass from string or bool (3-way dispatch: true / false / other) |
✓ — three-way branch with case-insensitive match + bool(raw_pass.strip()) fallback |
R2 — Clamp score to [0, 1] via max(0, min(1, x)) |
✓ |
R3 — Defensive non-numeric score (try/except TypeError, ValueError) |
✓ — caught and defaulted to 0.5 |
| R4 — Backwards compat (dict shape unchanged; strict emitters unaffected) | ✓ — verdict still has pass, score, reasons, next_hint keys; strict JSON emitters get identical results to v1 |
R5 — No new pip deps (math stdlib) |
✓ |
Code readability
- Three-way branch is more verbose than the v1 single-line
bool(...)but the case-intent is clearer: the comment "case-insensitive. The string 'true' → True; the string 'false' → False. Any other non-empty string → fall through to the existingbool(...)semantics" in the SPEC is preserved exactly by the if/elif/else. - The score-clamp block is two statements (try/except, then isfinite
check, then clamp). The order matters: the TypeError/ValueError from
float(None)orfloat("great")must be caught BEFORE themath.isfinitecall; elsemath.isfinite(None)raises TypeError uncaught. The order in the implementation is correct (try/except wraps the float call; isfinite only sees a finite-or-NaN float).
Defensive correctness check
bool(None)→ False (ifdatais{"pass": None}; treated as no-pass → False; matches pre-fixbool(None)= False; no regression).bool(0)→ False (if verifier emits"pass": 0); preserved.bool(1)→ True;bool([])False;bool({})False; all preserved — no regression for non-string types.- String
" tRuE "strips via.strip().lower()→ "true" → True. - String
"\nfalse"strips → "false" → False. Edge case covered.
Cross-script impact
parse_verdictis local toloop-runner.py; not duplicated tostatus.py. The change is contained._gate_score_plateauinstatus.pyconsumesscore_history(with clamped values via the runner's atomic write oflast_verdict) — already assumed[0, 1]. The clamp guarantees it._write_state_looptimestamps store JSON; clamped scores round-trip cleanly (no serialization loss).
Tests spot-check
test_other_truthy_string_pass(the bug found inline): verifies that the SPEC R1'sbool(...)fallback clause is honored — a"yes"string yieldsTrue. Pre-fix v1 behavior preserved.test_pass_with_surrounding_whitespace: covers an edge case (" true ") the SPEC didn't explicitly enumerate but is sensible.test_score_none_value_to_neutral:data.get("score", 0.0)returnsNone(key exists with None value);float(None)raises TypeError → caught → 0.5. Not in SPEC's explicit test plan but is a natural consequence of R3's TypeError coverage. Good defensive test.test_score_infinity_to_neutral: coversmath.isfinite(Infinity)→ False path. Added aftertest_score_nan_to_neutral; both prove the isfinite check.- All 22 tests pass.
Verdict
PASS — implementer followed SPEC; inline bug found and fixed during test; the fix matches SPEC R1's three-way dispatch wording exactly. Proceed to bug_find.