- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
9.7 KiB
Harden parse_verdict
Small pure-function hardening task: close the
add-loop-runner/BUG_REPORT.md O6 finding plus the un-noted sibling issue
(no score clamping). Both shipped in v1 because the verifier prompt's
contract layer was expected to enforce JSON-typed pass / numeric score
in [0,1]; observed real LLM responses and the O6 finding show the
runner should not trust the prompt contract alone.
This is a v1.1 hardening task. No CLI surface change; no new feature;
no schema migration. Pure robustness inside scripts/loop-runner.py's
parse_verdict.
Goal
parse_verdict(text) currently builds the verdict dict as:
verdict = {
"pass": bool(data.get("pass")),
"score": float(data.get("score", 0.0)),
}
Two issues:
-
passstring coercion bug (O6): if the verifier emits{"pass": "false", "score": 0.1},bool("false")returnsTrue(non-empty string is truthy). The tick recordspass=True; the--check-gatescore-plateau brake, the tick log, and the orchestrator downstream all see a "passing" tick when the verifier said "failing". This is a silent correctness bug. Real LLMs do emit JSON booleans most of the time, but OpenAI-grade models occasionally emit"false"/"true"strings (quote-wrapped). The runner should accept both. -
Unclamped
score: if the verifier emits"score": 1.5or"score": -0.3(out-of-contract), the value is stored as-is. Thescore_historycap and the score-plateau brake assume[0, 1]. A1.5value inflates the rolling-average computation; a-0.3value causesgate_score_plateauto compute a negative trend that looks like degradation when none exists. The verifier prompt (prompts/loop-verifier.md) declaresscoreis a float in[0, 1], but the runner should not rely on prompt-discipline alone.
Requirements
R1 — Coerce pass from string or bool
parse_verdict accepts the following as data["pass"]:
true/false(JSON bool) — already correct viajson.loads."true"/"false"(JSON string) — case-insensitive. The string"true"→True; the string"false"→False. Any other non-empty string → fall through to the existingbool(...)semantics (i.e. truthy). Empty string →False.
Implementation shape:
raw_pass = data.get("pass")
if isinstance(raw_pass, str):
verdict_pass = raw_pass.strip().lower() == "true"
else:
verdict_pass = bool(raw_pass)
This handles true/false strings AND retains current behavior for
actual booleans (True/False) AND numbers (0/1 — bool(0)
returns False; current behavior is preserved).
R2 — Clamp score to [0, 1]
After parsing float(data.get("score", 0.0)), clamp:
score = float(data.get("score", 0.0))
score = max(0.0, min(1.0, score))
NaN handling: float("nan") would propagate. If the verifier emits a
literal NaN (impossible in strict JSON; some Hermes-style models
occasionally emit it via float('nan') in reflowed text), the
max/min comparison returns NaN — both branches preserve NaN, NaN is
not equal to NaN, and score-plateau gate would see a constant NaN history.
Defensive: reject NaN / non-finite scores by treating them as 0.5 (the
verifier emitted something unusable; the midpoint is a neutral default).
Use math.isfinite:
import math
score = float(data.get("score", 0.0))
if not math.isfinite(score):
score = 0.5
score = max(0.0, min(1.0, score))
R3 — Defensive non-numeric score
If data.get("score") is a string like "0.8", float(...) already
handles it (Python's float accepts string numerics). If it's a
non-numeric string, float(...) raises ValueError. Current code doesn't
catch this; would propagate as an unhandled exception mid-tick (→ halt via
the runner's finally → lock releases → tick log shows a HALT but errors
aren't categorized as verifier_failed). Wrap the float conversion:
try:
score = float(data.get("score", 0.0))
except (TypeError, ValueError):
score = 0.5
R4 — Backwards compat
- Verdict dict shape is unchanged:
{"pass": bool, "score": float, "reasons": list[str], "next_hint": str}. Existing callers (cmd_tick,_gate_score_plateauindirectly viascore_history, tick-log line format) are unaffected. - Strict-JSON emitters (true booleans, numeric scores in
[0,1]) get identical results to current behavior. - The
score = 0.5defaults for NaN / non-numeric are new behavior; documented in the CHANGELOG.
R5 — No new pip deps
math is stdlib.
Test plan (tests/test_parse_verdict.py)
New file. Pure unit tests against parse_verdict; no subprocess, no
fixtures, no tmp_path needed. All use the function directly with literal
input strings.
test_pass_true_bool—{"pass": true, "score": 0.8}→pass=True, score=0.8.test_pass_false_bool—{"pass": false, "score": 0.2}→pass=False, score=0.2.test_pass_true_string—{"pass": "true", "score": 0.9}→pass=True, score=0.9(the O6 bug).test_pass_false_string—{"pass": "false", "score": 0.1}→pass=False, score=0.1(the O6 bug).test_pass_string_case_insensitive—{"pass": "FALSE", "score": 0.1}→pass=False.{"pass": "True", "score": 0.9}→pass=True.test_score_clamped_high—{"pass": true, "score": 1.5}→pass=True, score=1.0.test_score_clamped_low—{"pass": true, "score": -0.3}→pass=True, score=0.0.test_score_nan_to_neutral—{"pass": true, "score": NaN}→pass=True, score=0.5. (Usefloat("nan")literal in the test JSON__import__('math').nan— actually use the string"NaN"to simulate.test_score_non_numeric_string—{"pass": true, "score": "great"}→pass=True, score=0.5.test_score_numeric_string_ok—{"pass": true, "score": "0.75"}→pass=True, score=0.75(Python's float() already handles this; no regression).test_empty_pass_string—{"pass": "", "score": 0.5}→pass=False(per R1'sbool(raw_pass)fallback for non-true/falsestrings: empty string →bool("")→False).test_other_truthy_string_pass—{"pass": "yes", "score": 0.5}→pass=True("yes"is non-empty, non-true→bool("yes")isTrue). Backwards-compat with prior semantics.test_existing_fence_block_behavior—```json {"pass": true, "score": 0.8} ```→ still parses; new clamp/coerce don't break the fence-extractor path.
Concrete code shape
import math
def parse_verdict(text: str) -> Optional[dict]:
"""Parse verifier JSON verdict. Accepts raw, fenced, or commented JSON.
Required keys: pass (bool — also accepts "true"/"false" strings),
score (float — clamped to [0, 1]; NaN/non-finite defaults to 0.5).
Optional: reasons (list[str]), next_hint (str). Returns None on parse
failure.
"""
if not text or not text.strip():
return None
candidates = []
fence_match = _FENCE_RE.search(text)
if fence_match:
candidates.append(fence_match.group(1))
candidates.append(text)
for body in candidates:
body = _strip_comments(body).strip()
if not body:
continue
try:
data = json.loads(body)
except json.JSONDecodeError:
continue
if not isinstance(data, dict):
continue
if "pass" not in data:
continue
raw_pass = data.get("pass")
if isinstance(raw_pass, str):
verdict_pass = raw_pass.strip().lower() == "true"
else:
verdict_pass = bool(raw_pass)
try:
score = float(data.get("score", 0.0))
except (TypeError, ValueError):
score = 0.5
if not math.isfinite(score):
score = 0.5
score = max(0.0, min(1.0, score))
verdict = {
"pass": verdict_pass,
"score": score,
}
if "reasons" in data and isinstance(data["reasons"], list):
verdict["reasons"] = [str(r) for r in data["reasons"]]
else:
verdict["reasons"] = []
if "next_hint" in data and isinstance(data["next_hint"], str):
verdict["next_hint"] = data["next_hint"]
return verdict
return None
D-items (decisions locked for this task)
- D-V1:
"true"/"false"strings → bool via case-insensitive equality with"true". Other strings defer to currentbool(...)for backwards compat (a verifier emittingpass: "yes"keeps current truthy behavior). - D-V2:
scoreNaN / non-finite →0.5(neutral midpoint). This is arbitrary but defensible; documented in CHANGELOG. - D-V3:
scorenon-numeric string →0.5(same neutral default). Documented. - D-V4: No CLI flag to opt out of clamping. Strict emitters in
[0,1]are unaffected; loose emitters get a deterministic value rather than a raw one. - D-V5: Tests are pure-functional; no subprocess; no monkeypatch.
Non-goals
- No
parse_verdictrewrite instatus.py's_parse_verdict_status_line(different function, different concern — parsesVERDICT.mdSTATUS:PASS / FAIL strings; not in scope for this task). - No
loop-verifier.mdprompt changes (the prompt still asks for JSON booleans; the runner-side coercion is defense-in-depth). Verifier prompt changes are tracked separately. - No schema change to
.state.loopscore_history— existing floats already in[0,1]from prior ticks are unaffected; new ticks are clamped. - No
verifier_failedhalt prompt change.
Verification
python3 -m py_compile scripts/loop-runner.pypython3 -m pytest tests/test_parse_verdict.py -vpython3 -m pytest tests/ -q(full suite stays green; baseline 447 + new)