Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

3.3 KiB

Code Review: harden-parse-verdict

SPEC coverage

Requirement Status
R1 — Coerce pass from string or bool (3-way dispatch: true / false / other) ✓ — three-way branch with case-insensitive match + bool(raw_pass.strip()) fallback
R2 — Clamp score to [0, 1] via max(0, min(1, x)) ✓
R3 — Defensive non-numeric score (try/except TypeError, ValueError) ✓ — caught and defaulted to 0.5
R4 — Backwards compat (dict shape unchanged; strict emitters unaffected) ✓ — verdict still has pass, score, reasons, next_hint keys; strict JSON emitters get identical results to v1
R5 — No new pip deps (math stdlib) ✓

Code readability

  • Three-way branch is more verbose than the v1 single-line bool(...) but the case-intent is clearer: the comment "case-insensitive. The string 'true' → True; the string 'false' → False. Any other non-empty string → fall through to the existing bool(...) semantics" in the SPEC is preserved exactly by the if/elif/else.
  • The score-clamp block is two statements (try/except, then isfinite check, then clamp). The order matters: the TypeError/ValueError from float(None) or float("great") must be caught BEFORE the math.isfinite call; else math.isfinite(None) raises TypeError uncaught. The order in the implementation is correct (try/except wraps the float call; isfinite only sees a finite-or-NaN float).

Defensive correctness check

  • bool(None) → False (if data is {"pass": None}; treated as no-pass → False; matches pre-fix bool(None) = False; no regression).
  • bool(0) → False (if verifier emits "pass": 0); preserved.
  • bool(1) → True; bool([]) False; bool({}) False; all preserved — no regression for non-string types.
  • String " tRuE " strips via .strip().lower() → "true" → True.
  • String "\nfalse" strips → "false" → False. Edge case covered.

Cross-script impact

  • parse_verdict is local to loop-runner.py; not duplicated to status.py. The change is contained.
  • _gate_score_plateau in status.py consumes score_history (with clamped values via the runner's atomic write of last_verdict) — already assumed [0, 1]. The clamp guarantees it.
  • _write_state_loop timestamps store JSON; clamped scores round-trip cleanly (no serialization loss).

Tests spot-check

  • test_other_truthy_string_pass (the bug found inline): verifies that the SPEC R1's bool(...) fallback clause is honored — a "yes" string yields True. Pre-fix v1 behavior preserved.
  • test_pass_with_surrounding_whitespace: covers an edge case (" true ") the SPEC didn't explicitly enumerate but is sensible.
  • test_score_none_value_to_neutral: data.get("score", 0.0) returns None (key exists with None value); float(None) raises TypeError → caught → 0.5. Not in SPEC's explicit test plan but is a natural consequence of R3's TypeError coverage. Good defensive test.
  • test_score_infinity_to_neutral: covers math.isfinite(Infinity) → False path. Added after test_score_nan_to_neutral; both prove the isfinite check.
  • All 22 tests pass.

Verdict

PASS — implementer followed SPEC; inline bug found and fixed during test; the fix matches SPEC R1's three-way dispatch wording exactly. Proceed to bug_find.