- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
155 lines
6.2 KiB
Markdown
155 lines
6.2 KiB
Markdown
# Adversarial Bug Report: harden-parse-verdict
|
|
|
|
Probed `parse_verdict` with non-contract inputs. Each attack vector
|
|
hypothesized, tested, verdict given.
|
|
|
|
## A1 — `pass` as Python types (None, list, dict, int) — type-confusion
|
|
|
|
**Hypothesis**: A verifier emitting non-string non-bool `pass` values
|
|
(e.g. `{"pass": null}`, `{"pass": [false]}`, `{"pass": 0}`) could yield
|
|
surprising verdicts.
|
|
|
|
**Test**: 9-row sweep via `json.dumps` (Python `None` → JSON `null`,
|
|
Python `True`/`False` → JSON `true`/`false`):
|
|
|
|
| Input | Output | Notes |
|
|
|---|---|---|
|
|
| `pass: null` | `pass=False` | `bool(None)` = False; preserved from v1. |
|
|
| `pass: []` (empty list) | `pass=False` | `bool([])` = False; preserved. |
|
|
| `pass: [false]` (list with False) | `pass=True` | `bool([False])` = True (non-empty list is truthy). Surprising but documented Python semantics. v1 returned same. **Not a regression.** |
|
|
| `pass: [true]` | `pass=True` | Same. |
|
|
| `pass: {}` (empty dict) | `pass=False` | `bool({})` = False; preserved. |
|
|
| `pass: 0` (int) | `pass=False` | `bool(0)` = False; preserved. |
|
|
| `pass: 1` (int) | `pass=True` | `bool(1)` = True; preserved. |
|
|
| `pass: -1` (int) | `pass=True` | `bool(-1)` = True (non-zero); preserved. |
|
|
| `pass: 1.5` (float) | `pass=True` | `bool(1.5)` = True (non-zero); preserved. |
|
|
|
|
**Verdict**: PASS — no regression for any non-string non-bool type. RESET
|
|
behavior matches v1's `bool(...)` semantics. The SPEC's three-way
|
|
string/bool branch handles strings explicitly; everything else falls
|
|
through to v1's `bool(...)`.
|
|
|
|
## A2 — `score` as Python non-numeric types
|
|
|
|
**Hypothesis**: `score: []`, `score: {}`, `score: [1, 2]`, `score: "high"`
|
|
should default to 0.5 per R3 (TypeError / ValueError caught).
|
|
|
|
**Test**:
|
|
|
|
| Input | Output |
|
|
|---|---|
|
|
| `score: []` | `score=0.5` (TypeError caught by `float([])`) |
|
|
| `score: {}` | `score=0.5` (TypeError caught) |
|
|
| `score: [1, 2]` | `score=0.5` (TypeError caught) |
|
|
| `score: "high"` | `score=0.5` (ValueError caught) |
|
|
| `score: None` | `score=0.5` (TypeError caught) |
|
|
|
|
**Verdict**: PASS — R3's `except (TypeError, ValueError)` catches all
|
|
non-numeric types; defaults to 0.5 (D-V3). Confirmed.
|
|
|
|
## A3 — `score` as out-of-range numeric strings
|
|
|
|
**Hypothesis**: A verifier emitting `score: "2.0"` (an out-of-range
|
|
numeric STRING) bypasses the clamp because R3's except arm never fires
|
|
and R2's clamp applies after — but is the clamp correctly triggered?
|
|
|
|
**Test**:
|
|
|
|
| Input | Output |
|
|
|---|---|
|
|
| `score: "2.0"` | `score=1.0` (parse to 2.0, clamp to 1.0) |
|
|
| `score: "-0.5"` | `score=0.0` (parse to -0.5, clamp to 0.0) |
|
|
| `score: "0.75"` | `score=0.75` (parse to 0.75, no clamping) |
|
|
|
|
**Verdict**: PASS — clamping applies to all numeric inputs regardless of
|
|
whether they came in as JSON numbers or numeric strings. Confirmed in
|
|
the SPEC test plan (`test_score_numeric_string_ok` and the inline fix).
|
|
|
|
## A4 — `score` as JSON literal NaN / Infinity / -Infinity
|
|
|
|
**Hypothesis**: Some Hermes-style recursive decoders emit the bare
|
|
tokens `NaN` / `Infinity` / `-Infinity` (rejected by strict JSON but
|
|
accepted by Python's `json.loads` with the default `parse_constant`).
|
|
`float(NaN)` succeeds (returns `math.nan`). The `math.isfinite` check
|
|
catches it.
|
|
|
|
**Test**:
|
|
|
|
| Input | Output |
|
|
|---|---|
|
|
| `score: NaN` (bare token) | `score=0.5` (isfinite catches; D-V2 default) |
|
|
| `score: Infinity` (bare token) | `score=0.5` |
|
|
| `score: -Infinity` (bare token) | `score=0.5` |
|
|
|
|
**Verdict**: PASS — D-V2 documented neutral default. The `math.isfinite`
|
|
guard fires before the clamp so the NaN doesn't propagate through `max` /
|
|
`min`.
|
|
|
|
## A5 — `score` as JSON booleans (true / false)
|
|
|
|
**Hypothesis**: A verifier erroneously using `"score": true` instead of
|
|
`"score": 0.8` would yield `float(True)` = 1.0 in Python (no exception),
|
|
then clamp to 1.0 (no change). The result is "the verifier said pass
|
|
with a perfect score" — incorrect but not a crash. Is this OK?
|
|
|
|
**Test**:
|
|
|
|
| Input | Output |
|
|
|---|---|
|
|
| `score: true` | `score=1.0` (float(True) → max(0, min(1, 1.0)) → 1.0) |
|
|
| `score: false` | `score=0.0` (float(False) → 0.0) |
|
|
|
|
**Verdict**: PASS — `float(True)` is well-defined in Python. A verifier
|
|
mis-typing `score: true` produces a deterministic 1.0 (not a crash; not
|
|
NaN). Score-plateau gate will see consistent 1.0 across ticks → halt as
|
|
`score_plateau`. Reasonable downstream behavior; documented quirk.
|
|
|
|
## A6 — `pass` short strings ("t", "T", "f")
|
|
|
|
**Hypothesis**: A verifier abbreviating `pass: "t"` or `pass: "T"` might
|
|
be misread as True (since SPEC only says `"true"`/`"false"` exact match
|
|
maps to True/False). Per SPEC R1, other strings fall through to
|
|
`bool(...)`, which is truthy for non-empty.
|
|
|
|
**Test**:
|
|
|
|
| Input | Output |
|
|
|---|---|
|
|
| `pass: "t"` | `pass=True` (abstract: `bool("t")` = True; not "true") |
|
|
| `pass: "T"` | `pass=True` |
|
|
| `pass: "f"` | `pass=True` (truthy; surprising!) |
|
|
|
|
**Verdict**: PASS — documented behavior. Risk: a verifier emitting
|
|
`pass: "f"` intending "false" gets `True`. Same as v1. The SPEC's
|
|
contract is to use full `true`/`false` strings or JSON booleans. This
|
|
abbreviated-string case is undocumented but not a regression; future
|
|
prompt work (out of scope for this task) should discourage abbreviations.
|
|
|
|
## A7 — Combined: `pass: "false"` string with `score: NaN` literal — full
|
|
harsh-path coverage
|
|
|
|
**Hypothesis**: A both-broken verdict still yields a parseable dict with
|
|
coerced defaults rather than None.
|
|
|
|
**Test**: `{"pass": "false", "score": NaN}` literal — `parse_verdict`
|
|
returns `{"pass": False, "score": 0.5, "reasons": [], "next_hint": ""}`.
|
|
|
|
**Verdict**: PASS — both coercion paths fire; documented defaults applied.
|
|
|
|
## A8 — Whitespace-only strips: newline + tab in `pass` value
|
|
|
|
**Hypothesis**: A verifier emitting `pass: "\n true "` (whitespace-wrapped)
|
|
should yield True after `.strip()`.
|
|
|
|
**Confirmed via test `test_pass_with_surrounding_whitespace`** — `" true "`
|
|
strips cleanly. Newline/tab characters not explicitly tested but
|
|
`str.strip()` defaults to all whitespace; newlines strip too.
|
|
|
|
**Verdict**: PASS.
|
|
|
|
## No BLOCKERS
|
|
|
|
A1-A8 are all documented behaviors per SPEC R1+R2+R3 + D-V1/D-V2/D-V3.
|
|
All inputs that would have caused silent corruption (the `bool("false")=True`
|
|
bug) or crashes (TypeError from non-numeric scores) are now handled
|
|
defensively. Recommend proceeding to doc_review. |