Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

9.7 KiB

Harden parse_verdict

Small pure-function hardening task: close the add-loop-runner/BUG_REPORT.md O6 finding plus the un-noted sibling issue (no score clamping). Both shipped in v1 because the verifier prompt's contract layer was expected to enforce JSON-typed pass / numeric score in [0,1]; observed real LLM responses and the O6 finding show the runner should not trust the prompt contract alone.

This is a v1.1 hardening task. No CLI surface change; no new feature; no schema migration. Pure robustness inside scripts/loop-runner.py's parse_verdict.

Goal

parse_verdict(text) currently builds the verdict dict as:

verdict = {
    "pass": bool(data.get("pass")),
    "score": float(data.get("score", 0.0)),
}

Two issues:

  1. pass string coercion bug (O6): if the verifier emits {"pass": "false", "score": 0.1}, bool("false") returns True (non-empty string is truthy). The tick records pass=True; the --check-gate score-plateau brake, the tick log, and the orchestrator downstream all see a "passing" tick when the verifier said "failing". This is a silent correctness bug. Real LLMs do emit JSON booleans most of the time, but OpenAI-grade models occasionally emit "false" / "true" strings (quote-wrapped). The runner should accept both.

  2. Unclamped score: if the verifier emits "score": 1.5 or "score": -0.3 (out-of-contract), the value is stored as-is. The score_history cap and the score-plateau brake assume [0, 1]. A 1.5 value inflates the rolling-average computation; a -0.3 value causes gate_score_plateau to compute a negative trend that looks like degradation when none exists. The verifier prompt (prompts/loop-verifier.md) declares score is a float in [0, 1], but the runner should not rely on prompt-discipline alone.

Requirements

R1 — Coerce pass from string or bool

parse_verdict accepts the following as data["pass"]:

  • true / false (JSON bool) — already correct via json.loads.
  • "true" / "false" (JSON string) — case-insensitive. The string "true" → True; the string "false" → False. Any other non-empty string → fall through to the existing bool(...) semantics (i.e. truthy). Empty string → False.

Implementation shape:

raw_pass = data.get("pass")
if isinstance(raw_pass, str):
    verdict_pass = raw_pass.strip().lower() == "true"
else:
    verdict_pass = bool(raw_pass)

This handles true/false strings AND retains current behavior for actual booleans (True/False) AND numbers (0/1 — bool(0) returns False; current behavior is preserved).

R2 — Clamp score to [0, 1]

After parsing float(data.get("score", 0.0)), clamp:

score = float(data.get("score", 0.0))
score = max(0.0, min(1.0, score))

NaN handling: float("nan") would propagate. If the verifier emits a literal NaN (impossible in strict JSON; some Hermes-style models occasionally emit it via float('nan') in reflowed text), the max/min comparison returns NaN — both branches preserve NaN, NaN is not equal to NaN, and score-plateau gate would see a constant NaN history. Defensive: reject NaN / non-finite scores by treating them as 0.5 (the verifier emitted something unusable; the midpoint is a neutral default). Use math.isfinite:

import math
score = float(data.get("score", 0.0))
if not math.isfinite(score):
    score = 0.5
score = max(0.0, min(1.0, score))

R3 — Defensive non-numeric score

If data.get("score") is a string like "0.8", float(...) already handles it (Python's float accepts string numerics). If it's a non-numeric string, float(...) raises ValueError. Current code doesn't catch this; would propagate as an unhandled exception mid-tick (→ halt via the runner's finally → lock releases → tick log shows a HALT but errors aren't categorized as verifier_failed). Wrap the float conversion:

try:
    score = float(data.get("score", 0.0))
except (TypeError, ValueError):
    score = 0.5

R4 — Backwards compat

  • Verdict dict shape is unchanged: {"pass": bool, "score": float, "reasons": list[str], "next_hint": str}. Existing callers (cmd_tick, _gate_score_plateau indirectly via score_history, tick-log line format) are unaffected.
  • Strict-JSON emitters (true booleans, numeric scores in [0,1]) get identical results to current behavior.
  • The score = 0.5 defaults for NaN / non-numeric are new behavior; documented in the CHANGELOG.

R5 — No new pip deps

math is stdlib.

Test plan (tests/test_parse_verdict.py)

New file. Pure unit tests against parse_verdict; no subprocess, no fixtures, no tmp_path needed. All use the function directly with literal input strings.

  1. test_pass_true_bool — {"pass": true, "score": 0.8} → pass=True, score=0.8.
  2. test_pass_false_bool — {"pass": false, "score": 0.2} → pass=False, score=0.2.
  3. test_pass_true_string — {"pass": "true", "score": 0.9} → pass=True, score=0.9 (the O6 bug).
  4. test_pass_false_string — {"pass": "false", "score": 0.1} → pass=False, score=0.1 (the O6 bug).
  5. test_pass_string_case_insensitive — {"pass": "FALSE", "score": 0.1} → pass=False. {"pass": "True", "score": 0.9} → pass=True.
  6. test_score_clamped_high — {"pass": true, "score": 1.5} → pass=True, score=1.0.
  7. test_score_clamped_low — {"pass": true, "score": -0.3} → pass=True, score=0.0.
  8. test_score_nan_to_neutral — {"pass": true, "score": NaN} → pass=True, score=0.5. (Use float("nan") literal in the test JSON __import__('math').nan — actually use the string "NaN" to simulate.
  9. test_score_non_numeric_string — {"pass": true, "score": "great"} → pass=True, score=0.5.
  10. test_score_numeric_string_ok — {"pass": true, "score": "0.75"} → pass=True, score=0.75 (Python's float() already handles this; no regression).
  11. test_empty_pass_string — {"pass": "", "score": 0.5} → pass=False (per R1's bool(raw_pass) fallback for non-true/false strings: empty string → bool("") → False).
  12. test_other_truthy_string_pass — {"pass": "yes", "score": 0.5} → pass=True ("yes" is non-empty, non-true → bool("yes") is True). Backwards-compat with prior semantics.
  13. test_existing_fence_block_behavior — ```json {"pass": true, "score": 0.8} ``` → still parses; new clamp/coerce don't break the fence-extractor path.

Concrete code shape

import math

def parse_verdict(text: str) -> Optional[dict]:
    """Parse verifier JSON verdict. Accepts raw, fenced, or commented JSON.

    Required keys: pass (bool — also accepts "true"/"false" strings),
    score (float — clamped to [0, 1]; NaN/non-finite defaults to 0.5).
    Optional: reasons (list[str]), next_hint (str). Returns None on parse
    failure.
    """
    if not text or not text.strip():
        return None
    candidates = []
    fence_match = _FENCE_RE.search(text)
    if fence_match:
        candidates.append(fence_match.group(1))
    candidates.append(text)
    for body in candidates:
        body = _strip_comments(body).strip()
        if not body:
            continue
        try:
            data = json.loads(body)
        except json.JSONDecodeError:
            continue
        if not isinstance(data, dict):
            continue
        if "pass" not in data:
            continue
        raw_pass = data.get("pass")
        if isinstance(raw_pass, str):
            verdict_pass = raw_pass.strip().lower() == "true"
        else:
            verdict_pass = bool(raw_pass)
        try:
            score = float(data.get("score", 0.0))
        except (TypeError, ValueError):
            score = 0.5
        if not math.isfinite(score):
            score = 0.5
        score = max(0.0, min(1.0, score))
        verdict = {
            "pass": verdict_pass,
            "score": score,
        }
        if "reasons" in data and isinstance(data["reasons"], list):
            verdict["reasons"] = [str(r) for r in data["reasons"]]
        else:
            verdict["reasons"] = []
        if "next_hint" in data and isinstance(data["next_hint"], str):
            verdict["next_hint"] = data["next_hint"]
        return verdict
    return None

D-items (decisions locked for this task)

  • D-V1: "true" / "false" strings → bool via case-insensitive equality with "true". Other strings defer to current bool(...) for backwards compat (a verifier emitting pass: "yes" keeps current truthy behavior).
  • D-V2: score NaN / non-finite → 0.5 (neutral midpoint). This is arbitrary but defensible; documented in CHANGELOG.
  • D-V3: score non-numeric string → 0.5 (same neutral default). Documented.
  • D-V4: No CLI flag to opt out of clamping. Strict emitters in [0,1] are unaffected; loose emitters get a deterministic value rather than a raw one.
  • D-V5: Tests are pure-functional; no subprocess; no monkeypatch.

Non-goals

  • No parse_verdict rewrite in status.py's _parse_verdict_status_line (different function, different concern — parses VERDICT.md STATUS:PASS / FAIL strings; not in scope for this task).
  • No loop-verifier.md prompt changes (the prompt still asks for JSON booleans; the runner-side coercion is defense-in-depth). Verifier prompt changes are tracked separately.
  • No schema change to .state.loop score_history — existing floats already in [0,1] from prior ticks are unaffected; new ticks are clamped.
  • No verifier_failed halt prompt change.

Verification

  • python3 -m py_compile scripts/loop-runner.py
  • python3 -m pytest tests/test_parse_verdict.py -v
  • python3 -m pytest tests/ -q (full suite stays green; baseline 447 + new)