- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
6.2 KiB
Adversarial Bug Report: harden-parse-verdict
Probed parse_verdict with non-contract inputs. Each attack vector
hypothesized, tested, verdict given.
A1 — pass as Python types (None, list, dict, int) — type-confusion
Hypothesis: A verifier emitting non-string non-bool pass values
(e.g. {"pass": null}, {"pass": [false]}, {"pass": 0}) could yield
surprising verdicts.
Test: 9-row sweep via json.dumps (Python None → JSON null,
Python True/False → JSON true/false):
| Input | Output | Notes |
|---|---|---|
pass: null |
pass=False |
bool(None) = False; preserved from v1. |
pass: [] (empty list) |
pass=False |
bool([]) = False; preserved. |
pass: [false] (list with False) |
pass=True |
bool([False]) = True (non-empty list is truthy). Surprising but documented Python semantics. v1 returned same. Not a regression. |
pass: [true] |
pass=True |
Same. |
pass: {} (empty dict) |
pass=False |
bool({}) = False; preserved. |
pass: 0 (int) |
pass=False |
bool(0) = False; preserved. |
pass: 1 (int) |
pass=True |
bool(1) = True; preserved. |
pass: -1 (int) |
pass=True |
bool(-1) = True (non-zero); preserved. |
pass: 1.5 (float) |
pass=True |
bool(1.5) = True (non-zero); preserved. |
Verdict: PASS — no regression for any non-string non-bool type. RESET
behavior matches v1's bool(...) semantics. The SPEC's three-way
string/bool branch handles strings explicitly; everything else falls
through to v1's bool(...).
A2 — score as Python non-numeric types
Hypothesis: score: [], score: {}, score: [1, 2], score: "high"
should default to 0.5 per R3 (TypeError / ValueError caught).
Test:
| Input | Output |
|---|---|
score: [] |
score=0.5 (TypeError caught by float([])) |
score: {} |
score=0.5 (TypeError caught) |
score: [1, 2] |
score=0.5 (TypeError caught) |
score: "high" |
score=0.5 (ValueError caught) |
score: None |
score=0.5 (TypeError caught) |
Verdict: PASS — R3's except (TypeError, ValueError) catches all
non-numeric types; defaults to 0.5 (D-V3). Confirmed.
A3 — score as out-of-range numeric strings
Hypothesis: A verifier emitting score: "2.0" (an out-of-range
numeric STRING) bypasses the clamp because R3's except arm never fires
and R2's clamp applies after — but is the clamp correctly triggered?
Test:
| Input | Output |
|---|---|
score: "2.0" |
score=1.0 (parse to 2.0, clamp to 1.0) |
score: "-0.5" |
score=0.0 (parse to -0.5, clamp to 0.0) |
score: "0.75" |
score=0.75 (parse to 0.75, no clamping) |
Verdict: PASS — clamping applies to all numeric inputs regardless of
whether they came in as JSON numbers or numeric strings. Confirmed in
the SPEC test plan (test_score_numeric_string_ok and the inline fix).
A4 — score as JSON literal NaN / Infinity / -Infinity
Hypothesis: Some Hermes-style recursive decoders emit the bare
tokens NaN / Infinity / -Infinity (rejected by strict JSON but
accepted by Python's json.loads with the default parse_constant).
float(NaN) succeeds (returns math.nan). The math.isfinite check
catches it.
Test:
| Input | Output |
|---|---|
score: NaN (bare token) |
score=0.5 (isfinite catches; D-V2 default) |
score: Infinity (bare token) |
score=0.5 |
score: -Infinity (bare token) |
score=0.5 |
Verdict: PASS — D-V2 documented neutral default. The math.isfinite
guard fires before the clamp so the NaN doesn't propagate through max /
min.
A5 — score as JSON booleans (true / false)
Hypothesis: A verifier erroneously using "score": true instead of
"score": 0.8 would yield float(True) = 1.0 in Python (no exception),
then clamp to 1.0 (no change). The result is "the verifier said pass
with a perfect score" — incorrect but not a crash. Is this OK?
Test:
| Input | Output |
|---|---|
score: true |
score=1.0 (float(True) → max(0, min(1, 1.0)) → 1.0) |
score: false |
score=0.0 (float(False) → 0.0) |
Verdict: PASS — float(True) is well-defined in Python. A verifier
mis-typing score: true produces a deterministic 1.0 (not a crash; not
NaN). Score-plateau gate will see consistent 1.0 across ticks → halt as
score_plateau. Reasonable downstream behavior; documented quirk.
A6 — pass short strings ("t", "T", "f")
Hypothesis: A verifier abbreviating pass: "t" or pass: "T" might
be misread as True (since SPEC only says "true"/"false" exact match
maps to True/False). Per SPEC R1, other strings fall through to
bool(...), which is truthy for non-empty.
Test:
| Input | Output |
|---|---|
pass: "t" |
pass=True (abstract: bool("t") = True; not "true") |
pass: "T" |
pass=True |
pass: "f" |
pass=True (truthy; surprising!) |
Verdict: PASS — documented behavior. Risk: a verifier emitting
pass: "f" intending "false" gets True. Same as v1. The SPEC's
contract is to use full true/false strings or JSON booleans. This
abbreviated-string case is undocumented but not a regression; future
prompt work (out of scope for this task) should discourage abbreviations.
A7 — Combined: pass: "false" string with score: NaN literal — full
harsh-path coverage
Hypothesis: A both-broken verdict still yields a parseable dict with coerced defaults rather than None.
Test: {"pass": "false", "score": NaN} literal — parse_verdict
returns {"pass": False, "score": 0.5, "reasons": [], "next_hint": ""}.
Verdict: PASS — both coercion paths fire; documented defaults applied.
A8 — Whitespace-only strips: newline + tab in pass value
Hypothesis: A verifier emitting pass: "\n true " (whitespace-wrapped)
should yield True after .strip().
Confirmed via test test_pass_with_surrounding_whitespace — " true "
strips cleanly. Newline/tab characters not explicitly tested but
str.strip() defaults to all whitespace; newlines strip too.
Verdict: PASS.
No BLOCKERS
A1-A8 are all documented behaviors per SPEC R1+R2+R3 + D-V1/D-V2/D-V3.
All inputs that would have caused silent corruption (the bool("false")=True
bug) or crashes (TypeError from non-numeric scores) are now handled
defensively. Recommend proceeding to doc_review.