Files
automaton/tasks/harden-parse-verdict/IMPLEMENTATION.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

4.0 KiB

Implementation: harden-parse-verdict

SCOPE

Closed add-loop-runner/BUG_REPORT.md O6 (pass-string coercion bug) plus the un-noted sibling issue (no score clamping). Pure-function change to parse_verdict in scripts/loop-runner.py. No CLI surface change; no schema change; no new deps (math is stdlib).

FILES TOUCHED

  • scripts/loop-runner.py
    • Added import math.
    • parse_verdict(text): replaced verdict["pass"] = bool(data.get("pass")) with explicit string-vs-bool dispatch:
      • bool passed through → bool(True) = True; bool(False) = False (unchanged).
      • "true" (any case, leading/trailing whitespace stripped) → True.
      • "false" (any case, leading/trailing whitespace stripped) → False.
      • Any other string → bool(raw_pass.strip()) (empty → False; non-empty → True). Preserves old bool(...) truthy semantics for "yes" / etc.
    • Replaced verdict["score"] = float(data.get("score", 0.0)) with:
      • try: score = float(data.get("score", 0.0)) / except (TypeError, ValueError): score = 0.5 (TypeError for non-numeric types like None/list/dict; ValueError for non-numeric strings like "great").
      • if not math.isfinite(score): score = 0.5 (catches NaN, Infinity, -Infinity returned by some Hermes-style recursive decoders).
      • score = max(0.0, min(1.0, score)) (clamp to [0, 1]).
    • Updated docstring to spell out the new contract: pass accepts bool OR "true"/"false" strings (case-insensitive); score is clamped to [0, 1] with NaN/non-finite → 0.5.

D-ITEMS Locked

  • D-V1: "true" / "false" strings → bool via case-insensitive equality. Other strings defer to current bool(...) for backwards compat (pass: "yes" stays truthy).
  • D-V2: score NaN / non-finite → 0.5.
  • D-V3: score non-numeric string → 0.5.
  • D-V4: No opt-out flag for clamping. Strict emitters unaffected.
  • D-V5: Tests are pure-functional; no subprocess.

BUG FOUND AND FIXED INLINE

While running tests/test_parse_verdict.py, the test test_other_truthy_string_pass failed on the first iteration. The implementation had:

if isinstance(raw_pass, str):
    verdict_pass = raw_pass.strip().lower() == "true"

This treats EVERY non-"true" string as False — including "yes", which used to be True via bool("yes"). SPEC R1 explicitly says non-true/false strings fall through to bool(...) for backwards compat. Fixed to the three-way branch:

if isinstance(raw_pass, str):
    lower = raw_pass.strip().lower()
    if lower == "true":
        verdict_pass = True
    elif lower == "false":
        verdict_pass = False
    else:
        verdict_pass = bool(raw_pass.strip())

This is consistent with SPEC R1 wording. All 22 tests pass after the fix.

TESTS

New file tests/test_parse_verdict.py — 22 tests across 5 classes:

  • TestStrictBaseline (2): bool pass true/false; preserves existing semantics.
  • TestPassStringCoercion (6): "true"/"false" strings, case-insensitive, surrounding whitespace, empty string, other truthy string.
  • TestScoreClamping (9): clamped high (1.5→1.0), clamped low (-0.3→0.0), edges (0.0, 1.0), NaN → 0.5, Infinity → 0.5, non-numeric string "great" → 0.5, numeric string "0.75" → 0.75, missing score → 0.0, None score → 0.5.
  • TestFenceBlockStillWorks (2): existing fence-block path with bool pass; fence path with string-pass + clamped score.
  • TestOptionalKeysPreserved (3): reasons+next_hint combination, missing reasons → [], non-list reasons coerced to [].

TEST COUNT

  • Baseline: 447 passed (post-add-state-loop-lock).
  • New: +22 in tests/test_parse_verdict.py.
  • Final: 469 passed, 0 regressions.

PIPELINE TO COMPLETION

research → research:awaiting_approval → research:approved → implement. Next: → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete.