Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled

This commit is contained in:
Lap Tran
2026-06-24 10:31:49 -04:00
parent dd2726c0dd
commit e13513faaa
193 changed files with 14934 additions and 98 deletions
@@ -0,0 +1 @@
complete
@@ -0,0 +1,2 @@
research:approved|2026-06-24T00:08:32.650347+00:00|user
code_review:approved|2026-06-24T00:09:51.306902+00:00|user
@@ -0,0 +1,155 @@
# Adversarial Bug Report: harden-parse-verdict
Probed `parse_verdict` with non-contract inputs. Each attack vector
hypothesized, tested, verdict given.
## A1 — `pass` as Python types (None, list, dict, int) — type-confusion
**Hypothesis**: A verifier emitting non-string non-bool `pass` values
(e.g. `{"pass": null}`, `{"pass": [false]}`, `{"pass": 0}`) could yield
surprising verdicts.
**Test**: 9-row sweep via `json.dumps` (Python `None` → JSON `null`,
Python `True`/`False` → JSON `true`/`false`):
| Input | Output | Notes |
|---|---|---|
| `pass: null` | `pass=False` | `bool(None)` = False; preserved from v1. |
| `pass: []` (empty list) | `pass=False` | `bool([])` = False; preserved. |
| `pass: [false]` (list with False) | `pass=True` | `bool([False])` = True (non-empty list is truthy). Surprising but documented Python semantics. v1 returned same. **Not a regression.** |
| `pass: [true]` | `pass=True` | Same. |
| `pass: {}` (empty dict) | `pass=False` | `bool({})` = False; preserved. |
| `pass: 0` (int) | `pass=False` | `bool(0)` = False; preserved. |
| `pass: 1` (int) | `pass=True` | `bool(1)` = True; preserved. |
| `pass: -1` (int) | `pass=True` | `bool(-1)` = True (non-zero); preserved. |
| `pass: 1.5` (float) | `pass=True` | `bool(1.5)` = True (non-zero); preserved. |
**Verdict**: PASS — no regression for any non-string non-bool type. RESET
behavior matches v1's `bool(...)` semantics. The SPEC's three-way
string/bool branch handles strings explicitly; everything else falls
through to v1's `bool(...)`.
## A2 — `score` as Python non-numeric types
**Hypothesis**: `score: []`, `score: {}`, `score: [1, 2]`, `score: "high"`
should default to 0.5 per R3 (TypeError / ValueError caught).
**Test**:
| Input | Output |
|---|---|
| `score: []` | `score=0.5` (TypeError caught by `float([])`) |
| `score: {}` | `score=0.5` (TypeError caught) |
| `score: [1, 2]` | `score=0.5` (TypeError caught) |
| `score: "high"` | `score=0.5` (ValueError caught) |
| `score: None` | `score=0.5` (TypeError caught) |
**Verdict**: PASS — R3's `except (TypeError, ValueError)` catches all
non-numeric types; defaults to 0.5 (D-V3). Confirmed.
## A3 — `score` as out-of-range numeric strings
**Hypothesis**: A verifier emitting `score: "2.0"` (an out-of-range
numeric STRING) bypasses the clamp because R3's except arm never fires
and R2's clamp applies after — but is the clamp correctly triggered?
**Test**:
| Input | Output |
|---|---|
| `score: "2.0"` | `score=1.0` (parse to 2.0, clamp to 1.0) |
| `score: "-0.5"` | `score=0.0` (parse to -0.5, clamp to 0.0) |
| `score: "0.75"` | `score=0.75` (parse to 0.75, no clamping) |
**Verdict**: PASS — clamping applies to all numeric inputs regardless of
whether they came in as JSON numbers or numeric strings. Confirmed in
the SPEC test plan (`test_score_numeric_string_ok` and the inline fix).
## A4 — `score` as JSON literal NaN / Infinity / -Infinity
**Hypothesis**: Some Hermes-style recursive decoders emit the bare
tokens `NaN` / `Infinity` / `-Infinity` (rejected by strict JSON but
accepted by Python's `json.loads` with the default `parse_constant`).
`float(NaN)` succeeds (returns `math.nan`). The `math.isfinite` check
catches it.
**Test**:
| Input | Output |
|---|---|
| `score: NaN` (bare token) | `score=0.5` (isfinite catches; D-V2 default) |
| `score: Infinity` (bare token) | `score=0.5` |
| `score: -Infinity` (bare token) | `score=0.5` |
**Verdict**: PASS — D-V2 documented neutral default. The `math.isfinite`
guard fires before the clamp so the NaN doesn't propagate through `max` /
`min`.
## A5 — `score` as JSON booleans (true / false)
**Hypothesis**: A verifier erroneously using `"score": true` instead of
`"score": 0.8` would yield `float(True)` = 1.0 in Python (no exception),
then clamp to 1.0 (no change). The result is "the verifier said pass
with a perfect score" — incorrect but not a crash. Is this OK?
**Test**:
| Input | Output |
|---|---|
| `score: true` | `score=1.0` (float(True) → max(0, min(1, 1.0)) → 1.0) |
| `score: false` | `score=0.0` (float(False) → 0.0) |
**Verdict**: PASS — `float(True)` is well-defined in Python. A verifier
mis-typing `score: true` produces a deterministic 1.0 (not a crash; not
NaN). Score-plateau gate will see consistent 1.0 across ticks → halt as
`score_plateau`. Reasonable downstream behavior; documented quirk.
## A6 — `pass` short strings ("t", "T", "f")
**Hypothesis**: A verifier abbreviating `pass: "t"` or `pass: "T"` might
be misread as True (since SPEC only says `"true"`/`"false"` exact match
maps to True/False). Per SPEC R1, other strings fall through to
`bool(...)`, which is truthy for non-empty.
**Test**:
| Input | Output |
|---|---|
| `pass: "t"` | `pass=True` (abstract: `bool("t")` = True; not "true") |
| `pass: "T"` | `pass=True` |
| `pass: "f"` | `pass=True` (truthy; surprising!) |
**Verdict**: PASS — documented behavior. Risk: a verifier emitting
`pass: "f"` intending "false" gets `True`. Same as v1. The SPEC's
contract is to use full `true`/`false` strings or JSON booleans. This
abbreviated-string case is undocumented but not a regression; future
prompt work (out of scope for this task) should discourage abbreviations.
## A7 — Combined: `pass: "false"` string with `score: NaN` literal — full
harsh-path coverage
**Hypothesis**: A both-broken verdict still yields a parseable dict with
coerced defaults rather than None.
**Test**: `{"pass": "false", "score": NaN}` literal — `parse_verdict`
returns `{"pass": False, "score": 0.5, "reasons": [], "next_hint": ""}`.
**Verdict**: PASS — both coercion paths fire; documented defaults applied.
## A8 — Whitespace-only strips: newline + tab in `pass` value
**Hypothesis**: A verifier emitting `pass: "\n true "` (whitespace-wrapped)
should yield True after `.strip()`.
**Confirmed via test `test_pass_with_surrounding_whitespace`** — `" true "`
strips cleanly. Newline/tab characters not explicitly tested but
`str.strip()` defaults to all whitespace; newlines strip too.
**Verdict**: PASS.
## No BLOCKERS
A1-A8 are all documented behaviors per SPEC R1+R2+R3 + D-V1/D-V2/D-V3.
All inputs that would have caused silent corruption (the `bool("false")=True`
bug) or crashes (TypeError from non-numeric scores) are now handled
defensively. Recommend proceeding to doc_review.
@@ -0,0 +1,78 @@
# Bug Report: harden-parse-verdict
Bug_find phase observations. Each non-blocking unless marked BLOCKER.
## O1 — `bool(raw_pass.strip())` fallback for "0" / "1" strings
For `{"pass": "0", "score": 0.5}`, the new path strips → "0" → not "true"/
"false" → `bool("0")` = True (non-empty string is truthy).
Pre-fix: `bool("0")` was also True (same).
This is **consistent with v1** — no behavior change. A verifier emitting
`pass: "0"` intending "false" gets `True` in v1 AND in the hardened
implementation. The SPEC says non-`true`/`false` strings fall through to
`bool(...)`, which is truthy for non-empty. Not a regression.
**Not a bug** — documented behavior per SPEC R1 + D-V1. If users want
strict numeric-string handling, that's a separate future task (out of
scope for v1.1 harden-parse-verdict).
## O2 — `float("nan")` serializes back as `NaN` to `.state.loop`
When the verifier emits `NaN` as the score, `parse_verdict` returns
`score=0.5` (without writing to disk by itself). But this score is part
of `verdict` which gets written via `_write_state_loop(loop_path, state)`
into `.state.loop` as JSON. Since the clamp converts NaN to 0.5 BEFORE
the verdict is stored, `.state.loop` gets `0.5`, not `NaN`. No NaN leaks
into the loop state.
**Not a bug** — confirmed via tracing: `parse_verdict` returns the clamped
dict; `cmd_tick` then stores `state["last_verdict"] = verdict` (a clean
dict with `score: 0.5`); the JSON round-trip is clean.
## O3 — `_gate_score_plateau`'s threshold unaffected
`_gate_score_plateau` (status.py) reads `score_history` and decides a halt
when last N scores are within some delta. With clamped scores, the plateau
detection range is now strictly `[0, 1]` instead of `[any, any]`. Brief
review:
- Pre-fix: a verifier could emit `score: 1.5` across N ticks; plateau
detection sees a flat line at 1.5; halt fires. Expected behavior.
- Post-fix: the same verifier's 1.5 clamps to 1.0 across N ticks; plateau
sees flat line at 1.0; halt fires. Same outcome.
- Pre-fix: a verifier emits alternating `0.9` and `1.1`; plateau sees a
bimodal history [0.9, 1.1, 0.9, 1.1] — NOT plateau (variation > epsilon).
- Post-fix: alternating `0.9` and `1.0` (1.1 clamps to 1.0); plateau sees
[0.9, 1.0, 0.9, 1.0] — still variation above a small epsilon — still NOT
plateau. Same outcome in this scenario.
Edge case: a verifier emits all `1.0` and `0.99` (vs 1.0 and 1.0 clamped).
The clamp DOES change plateau detection in this case — `1.0 1.0 1.0`
looks more plateau-like than `1.0 0.99 0.99`. Could cause halt earlier than
prior. Documented as a desirable side effect (clamp reduces the verifier's
untrustworthiness from inflating scores; plateau detection is more
honest).
**Not a bug** — improved behavior. Documented in CHANGELOG.
## O4 — Verifier prompt hasn't been updated
`prompts/loop-verifier.md` still asks the model to emit JSON with
`"pass": true/false` and `"score": 0.0-1.0`. The runner now defensive-coerces,
but the prompt's contract is unchanged. Was the prompt already
JSON-typed-booleans-only? Let me check.
**`prompts/loop-verifier.md` review**: still says "Output: strict JSON, no
prose" with example shape. The prompt explicitly tells the model to emit
JSON booleans — no mention of string-typed `pass`. So the v1 contract was
strict; the O6 finding was a defense-in-depth concern, not a present-fault.
The harden task adds belt-and-suspenders without changing the contract.
**Not a bug** — the prompt remains authoritative. No edit needed.
## Verdict
**No blockers.** Proceed to adversarial_bug_find.
@@ -0,0 +1,67 @@
# Code Review: harden-parse-verdict
## SPEC coverage
| Requirement | Status |
|-------------|--------|
| R1 — Coerce `pass` from string or bool (3-way dispatch: true / false / other) | ✓ — three-way branch with case-insensitive match + `bool(raw_pass.strip())` fallback |
| R2 — Clamp `score` to `[0, 1]` via `max(0, min(1, x))` | ✓ |
| R3 — Defensive non-numeric `score` (try/except TypeError, ValueError) | ✓ — caught and defaulted to 0.5 |
| R4 — Backwards compat (dict shape unchanged; strict emitters unaffected) | ✓ — verdict still has `pass`, `score`, `reasons`, `next_hint` keys; strict JSON emitters get identical results to v1 |
| R5 — No new pip deps (`math` stdlib) | ✓ |
## Code readability
- Three-way branch is more verbose than the v1 single-line `bool(...)` but
the case-intent is clearer: the comment "case-insensitive. The string
'true' → True; the string 'false' → False. Any other non-empty string
→ fall through to the existing `bool(...)` semantics" in the SPEC is
preserved exactly by the if/elif/else.
- The score-clamp block is two statements (try/except, then isfinite
check, then clamp). The order matters: the TypeError/ValueError from
`float(None)` or `float("great")` must be caught BEFORE the `math.isfinite`
call; else `math.isfinite(None)` raises TypeError uncaught. The order
in the implementation is correct (try/except wraps the float call;
isfinite only sees a finite-or-NaN float).
## Defensive correctness check
- `bool(None)` → False (if `data` is `{"pass": None}`; treated as no-pass
→ False; matches pre-fix `bool(None)` = False; no regression).
- `bool(0)` → False (if verifier emits `"pass": 0`); preserved.
- `bool(1)` → True; `bool([])` False; `bool({})` False; all preserved — no
regression for non-string types.
- String `" tRuE "` strips via `.strip().lower()` → "true" → True.
- String `"\nfalse"` strips → "false" → False. Edge case covered.
## Cross-script impact
- `parse_verdict` is local to `loop-runner.py`; not duplicated to
`status.py`. The change is contained.
- `_gate_score_plateau` in `status.py` consumes `score_history` (with
clamped values via the runner's atomic write of `last_verdict`) —
already assumed `[0, 1]`. The clamp guarantees it.
- `_write_state_loop` timestamps store JSON; clamped scores
round-trip cleanly (no serialization loss).
## Tests spot-check
- `test_other_truthy_string_pass` (the bug found inline): verifies that
the SPEC R1's `bool(...)` fallback clause is honored — a `"yes"` string
yields `True`. Pre-fix v1 behavior preserved.
- `test_pass_with_surrounding_whitespace`: covers an edge case (`" true "`)
the SPEC didn't explicitly enumerate but is sensible.
- `test_score_none_value_to_neutral`: `data.get("score", 0.0)` returns
`None` (key exists with None value); `float(None)` raises TypeError →
caught → 0.5. Not in SPEC's explicit test plan but is a natural
consequence of R3's TypeError coverage. Good defensive test.
- `test_score_infinity_to_neutral`: covers `math.isfinite(Infinity)` →
False path. Added after `test_score_nan_to_neutral`; both prove the
isfinite check.
- All 22 tests pass.
## Verdict
PASS — implementer followed SPEC; inline bug found and fixed during
test; the fix matches SPEC R1's three-way dispatch wording exactly.
Proceed to bug_find.
@@ -0,0 +1,45 @@
# Doc Review: harden-parse-verdict
## Docs touched
- `design/loops/technical.md` §7 — tick-flow step 7 (parse verdict):
added 4-line inline block documenting the defensive coercion (pass
string acceptance; score clamp + NaN/inf/non-numeric → 0.5).
- `design/loops/functional.md` §10 — Verifier Contract: annotated
`pass` (bool) definitive + runner accepts `"true"`/`"false"` strings;
annotated `score` (0.0–1.0) clamp + NaN/inf/non-numeric → 0.5 neutral.
- `CHANGELOG.md` — new `[unreleased]` "Fixed — `parse_verdict`
defensive coercion" block above the existing `add-state-loop-lock` and
`fix-harness-command-template` blocks.
## Docs NOT touched (intentional)
- `AGENTS.md`: parse_verdict is not a user-visible CLI surface; the
hardening doesn't change phase enforcement, `.state.loop`, or any
contract that harness integrators need to know. The Verifier Contract
lives in `design/loops/functional.md` §10; AGENTS.md already points to
design docs at the top. No edit.
- `README.md`: user-facing README doesn't enumerate `parse_verdict`
internals; loop monitoring table mentions verdicts as a concept, not
the parser. No edit.
- `prompts/loop-verifier.md`: contract was already `bool pass` + `score
0.0–1.0`. The hardening is belt-and-suspenders against malformed
output, not a contract change. The prompt's strict-JSON directive
stays authoritative. No edit.
- `templates/loops/self-improvement/loop.json`: no schema change. No
edit.
## Cross-references
- `tasks/add-loop-runner/BUG_REPORT.md` O6 — the original finding —
now closed by this task. The CHANGELOG entry explicitly references it.
- `tasks/add-loop-runner/ADVERSARIAL_BUG_REPORT.md` A6 — the score-
clamping observation — also closed by this task. The CHANGELOG entry
references the clamping.
- `tasks/harden-parse-verdict/BUG_REPORT.md` O3 — notes that clamping
improves plateau detection (a tighter `score_history` range makes
plateau more honest). Cross-referenced from the CHANGELOG.
## Verdict
Docs are in sync with the implementation. Proceed to referee.
@@ -0,0 +1,98 @@
# Implementation: harden-parse-verdict
## SCOPE
Closed `add-loop-runner/BUG_REPORT.md` O6 (pass-string coercion bug) plus
the un-noted sibling issue (no score clamping). Pure-function change to
`parse_verdict` in `scripts/loop-runner.py`. No CLI surface change; no
schema change; no new deps (math is stdlib).
## FILES TOUCHED
- `scripts/loop-runner.py`
- Added `import math`.
- `parse_verdict(text)`: replaced `verdict["pass"] = bool(data.get("pass"))`
with explicit string-vs-bool dispatch:
- bool passed through → `bool(True)` = True; `bool(False)` = False (unchanged).
- `"true"` (any case, leading/trailing whitespace stripped) → True.
- `"false"` (any case, leading/trailing whitespace stripped) → False.
- Any other string → `bool(raw_pass.strip())` (empty → False; non-empty → True).
Preserves old `bool(...)` truthy semantics for `"yes"` / etc.
- Replaced `verdict["score"] = float(data.get("score", 0.0))` with:
- `try: score = float(data.get("score", 0.0))` /
`except (TypeError, ValueError): score = 0.5`
(TypeError for non-numeric types like None/list/dict; ValueError for
non-numeric strings like "great").
- `if not math.isfinite(score): score = 0.5` (catches NaN, Infinity,
-Infinity returned by some Hermes-style recursive decoders).
- `score = max(0.0, min(1.0, score))` (clamp to `[0, 1]`).
- Updated docstring to spell out the new contract: `pass` accepts
bool OR `"true"`/`"false"` strings (case-insensitive); `score` is
clamped to `[0, 1]` with NaN/non-finite → 0.5.
## D-ITEMS Locked
- **D-V1**: `"true"` / `"false"` strings → bool via case-insensitive
equality. Other strings defer to current `bool(...)` for backwards
compat (`pass: "yes"` stays truthy).
- **D-V2**: `score` NaN / non-finite → `0.5`.
- **D-V3**: `score` non-numeric string → `0.5`.
- **D-V4**: No opt-out flag for clamping. Strict emitters unaffected.
- **D-V5**: Tests are pure-functional; no subprocess.
## BUG FOUND AND FIXED INLINE
While running `tests/test_parse_verdict.py`, the test
`test_other_truthy_string_pass` failed on the first iteration. The
implementation had:
```python
if isinstance(raw_pass, str):
verdict_pass = raw_pass.strip().lower() == "true"
```
This treats EVERY non-`"true"` string as False — including `"yes"`,
which used to be True via `bool("yes")`. SPEC R1 explicitly says
non-`true`/`false` strings fall through to `bool(...)` for backwards
compat. Fixed to the three-way branch:
```python
if isinstance(raw_pass, str):
lower = raw_pass.strip().lower()
if lower == "true":
verdict_pass = True
elif lower == "false":
verdict_pass = False
else:
verdict_pass = bool(raw_pass.strip())
```
This is consistent with SPEC R1 wording. All 22 tests pass after the fix.
## TESTS
New file `tests/test_parse_verdict.py` — 22 tests across 5 classes:
- `TestStrictBaseline` (2): bool `pass` true/false; preserves existing semantics.
- `TestPassStringCoercion` (6): `"true"`/`"false"` strings, case-insensitive,
surrounding whitespace, empty string, other truthy string.
- `TestScoreClamping` (9): clamped high (1.5→1.0), clamped low (-0.3→0.0),
edges (0.0, 1.0), NaN → 0.5, Infinity → 0.5, non-numeric string "great"
→ 0.5, numeric string "0.75" → 0.75, missing score → 0.0, None score →
0.5.
- `TestFenceBlockStillWorks` (2): existing fence-block path with bool pass;
fence path with string-pass + clamped score.
- `TestOptionalKeysPreserved` (3): reasons+next_hint combination, missing
reasons → [], non-list reasons coerced to [].
## TEST COUNT
- Baseline: 447 passed (post-`add-state-loop-lock`).
- New: +22 in `tests/test_parse_verdict.py`.
- Final: **469 passed**, 0 regressions.
## PIPELINE TO COMPLETION
research → research:awaiting_approval → research:approved → implement.
Next: → code_review → code_review:awaiting_approval → code_review:approved
→ bug_find → adversarial_bug_find → doc_review → referee → complete.
+258
View File
@@ -0,0 +1,258 @@
# Harden `parse_verdict`
Small pure-function hardening task: close the
`add-loop-runner/BUG_REPORT.md` O6 finding plus the un-noted sibling issue
(no score clamping). Both shipped in v1 because the verifier prompt's
contract layer was expected to enforce JSON-typed `pass` / numeric `score`
in `[0,1]`; observed real LLM responses and the O6 finding show the
runner should not trust the prompt contract alone.
This is a v1.1 hardening task. No CLI surface change; no new feature;
no schema migration. Pure robustness inside `scripts/loop-runner.py`'s
`parse_verdict`.
## Goal
`parse_verdict(text)` currently builds the verdict dict as:
```python
verdict = {
"pass": bool(data.get("pass")),
"score": float(data.get("score", 0.0)),
}
```
Two issues:
1. **`pass` string coercion bug (O6)**: if the verifier emits
`{"pass": "false", "score": 0.1}`, `bool("false")` returns `True`
(non-empty string is truthy). The tick records `pass=True`; the
`--check-gate` score-plateau brake, the tick log, and the orchestrator
downstream all see a "passing" tick when the verifier said "failing".
This is a silent correctness bug. Real LLMs do emit JSON booleans most
of the time, but OpenAI-grade models occasionally emit `"false"` /
`"true"` strings (quote-wrapped). The runner should accept both.
2. **Unclamped `score`**: if the verifier emits `"score": 1.5` or
`"score": -0.3` (out-of-contract), the value is stored as-is. The
`score_history` cap and the score-plateau brake assume `[0, 1]`. A
`1.5` value inflates the rolling-average computation; a `-0.3`
value causes `gate_score_plateau` to compute a negative trend that
looks like degradation when none exists. The verifier prompt
(`prompts/loop-verifier.md`) declares `score` is a float in `[0, 1]`,
but the runner should not rely on prompt-discipline alone.
## Requirements
### R1 — Coerce `pass` from string or bool
`parse_verdict` accepts the following as `data["pass"]`:
- `true` / `false` (JSON bool) — already correct via `json.loads`.
- `"true"` / `"false"` (JSON string) — case-insensitive. The string `"true"`
→ `True`; the string `"false"` → `False`. Any other non-empty string
→ fall through to the existing `bool(...)` semantics (i.e. truthy).
Empty string → `False`.
Implementation shape:
```python
raw_pass = data.get("pass")
if isinstance(raw_pass, str):
verdict_pass = raw_pass.strip().lower() == "true"
else:
verdict_pass = bool(raw_pass)
```
This handles `true`/`false` strings AND retains current behavior for
actual booleans (`True`/`False`) AND numbers (`0`/`1` — `bool(0)`
returns `False`; current behavior is preserved).
### R2 — Clamp `score` to `[0, 1]`
After parsing `float(data.get("score", 0.0))`, clamp:
```python
score = float(data.get("score", 0.0))
score = max(0.0, min(1.0, score))
```
NaN handling: `float("nan")` would propagate. If the verifier emits a
literal NaN (impossible in strict JSON; some Hermes-style models
occasionally emit it via `float('nan')` in reflowed text), the
`max/min` comparison returns NaN — both branches preserve NaN, NaN is
not equal to NaN, and score-plateau gate would see a constant NaN history.
Defensive: reject NaN / non-finite scores by treating them as 0.5 (the
verifier emitted something unusable; the midpoint is a neutral default).
Use `math.isfinite`:
```python
import math
score = float(data.get("score", 0.0))
if not math.isfinite(score):
score = 0.5
score = max(0.0, min(1.0, score))
```
### R3 — Defensive non-numeric `score`
If `data.get("score")` is a string like `"0.8"`, `float(...)` already
handles it (Python's `float` accepts string numerics). If it's a
non-numeric string, `float(...)` raises `ValueError`. Current code doesn't
catch this; would propagate as an unhandled exception mid-tick (→ halt via
the runner's finally → lock releases → tick log shows a HALT but errors
aren't categorized as `verifier_failed`). Wrap the float conversion:
```python
try:
score = float(data.get("score", 0.0))
except (TypeError, ValueError):
score = 0.5
```
### R4 — Backwards compat
- Verdict dict shape is unchanged: `{"pass": bool, "score": float,
"reasons": list[str], "next_hint": str}`. Existing callers (`cmd_tick`,
`_gate_score_plateau` indirectly via `score_history`, tick-log line
format) are unaffected.
- Strict-JSON emitters (true booleans, numeric scores in `[0,1]`) get
identical results to current behavior.
- The `score = 0.5` defaults for NaN / non-numeric are new behavior;
documented in the CHANGELOG.
### R5 — No new pip deps
`math` is stdlib.
## Test plan (`tests/test_parse_verdict.py`)
New file. Pure unit tests against `parse_verdict`; no subprocess, no
fixtures, no tmp_path needed. All use the function directly with literal
input strings.
1. `test_pass_true_bool` — `{"pass": true, "score": 0.8}` →
`pass=True, score=0.8`.
2. `test_pass_false_bool` — `{"pass": false, "score": 0.2}` →
`pass=False, score=0.2`.
3. `test_pass_true_string` — `{"pass": "true", "score": 0.9}` →
`pass=True, score=0.9` (the O6 bug).
4. `test_pass_false_string` — `{"pass": "false", "score": 0.1}` →
`pass=False, score=0.1` (the O6 bug).
5. `test_pass_string_case_insensitive` — `{"pass": "FALSE", "score": 0.1}`
→ `pass=False`. `{"pass": "True", "score": 0.9}` → `pass=True`.
6. `test_score_clamped_high` — `{"pass": true, "score": 1.5}` →
`pass=True, score=1.0`.
7. `test_score_clamped_low` — `{"pass": true, "score": -0.3}` →
`pass=True, score=0.0`.
8. `test_score_nan_to_neutral` — `{"pass": true, "score": NaN}` →
`pass=True, score=0.5`. (Use `float("nan")` literal in the test JSON
`__import__('math').nan` — actually use the string `"NaN"` to
simulate.
9. `test_score_non_numeric_string` — `{"pass": true, "score": "great"}`
→ `pass=True, score=0.5`.
10. `test_score_numeric_string_ok` — `{"pass": true, "score": "0.75"}`
→ `pass=True, score=0.75` (Python's float() already handles this; no
regression).
11. `test_empty_pass_string` — `{"pass": "", "score": 0.5}` →
`pass=False` (per R1's `bool(raw_pass)` fallback for non-`true`/`false`
strings: empty string → `bool("")` → `False`).
12. `test_other_truthy_string_pass` — `{"pass": "yes", "score": 0.5}` →
`pass=True` (`"yes"` is non-empty, non-`true` → `bool("yes")` is `True`).
Backwards-compat with prior semantics.
13. `test_existing_fence_block_behavior` — ```` ```json {"pass": true,
"score": 0.8} ``` ```` → still parses; new clamp/coerce don't break
the fence-extractor path.
## Concrete code shape
```python
import math
def parse_verdict(text: str) -> Optional[dict]:
"""Parse verifier JSON verdict. Accepts raw, fenced, or commented JSON.
Required keys: pass (bool — also accepts "true"/"false" strings),
score (float — clamped to [0, 1]; NaN/non-finite defaults to 0.5).
Optional: reasons (list[str]), next_hint (str). Returns None on parse
failure.
"""
if not text or not text.strip():
return None
candidates = []
fence_match = _FENCE_RE.search(text)
if fence_match:
candidates.append(fence_match.group(1))
candidates.append(text)
for body in candidates:
body = _strip_comments(body).strip()
if not body:
continue
try:
data = json.loads(body)
except json.JSONDecodeError:
continue
if not isinstance(data, dict):
continue
if "pass" not in data:
continue
raw_pass = data.get("pass")
if isinstance(raw_pass, str):
verdict_pass = raw_pass.strip().lower() == "true"
else:
verdict_pass = bool(raw_pass)
try:
score = float(data.get("score", 0.0))
except (TypeError, ValueError):
score = 0.5
if not math.isfinite(score):
score = 0.5
score = max(0.0, min(1.0, score))
verdict = {
"pass": verdict_pass,
"score": score,
}
if "reasons" in data and isinstance(data["reasons"], list):
verdict["reasons"] = [str(r) for r in data["reasons"]]
else:
verdict["reasons"] = []
if "next_hint" in data and isinstance(data["next_hint"], str):
verdict["next_hint"] = data["next_hint"]
return verdict
return None
```
## D-items (decisions locked for this task)
- **D-V1**: `"true"` / `"false"` strings → bool via case-insensitive
equality with `"true"`. Other strings defer to current `bool(...)` for
backwards compat (a verifier emitting `pass: "yes"` keeps current
truthy behavior).
- **D-V2**: `score` NaN / non-finite → `0.5` (neutral midpoint). This is
arbitrary but defensible; documented in CHANGELOG.
- **D-V3**: `score` non-numeric string → `0.5` (same neutral default).
Documented.
- **D-V4**: No CLI flag to opt out of clamping. Strict emitters in `[0,1]`
are unaffected; loose emitters get a deterministic value rather than a
raw one.
- **D-V5**: Tests are pure-functional; no subprocess; no monkeypatch.
## Non-goals
- No `parse_verdict` rewrite in `status.py`'s `_parse_verdict_status_line`
(different function, different concern — parses `VERDICT.md`
STATUS:PASS / FAIL strings; not in scope for this task).
- No `loop-verifier.md` prompt changes (the prompt still asks for JSON
booleans; the runner-side coercion is defense-in-depth). Verifier
prompt changes are tracked separately.
- No schema change to `.state.loop` `score_history` — existing floats
already in `[0,1]` from prior ticks are unaffected; new ticks are
clamped.
- No `verifier_failed` halt prompt change.
## Verification
- `python3 -m py_compile scripts/loop-runner.py`
- `python3 -m pytest tests/test_parse_verdict.py -v`
- `python3 -m pytest tests/ -q` (full suite stays green; baseline 447 +
new)
@@ -0,0 +1,55 @@
# Referee Verdict: harden-parse-verdict
## Status: PASS
## Artifacts reviewed
- `SPEC.md` — R1-R5 + D-V1 to D-V5; 13-item test plan
- `IMPLEMENTATION.md` — files touched, decisions locked, inline bug found and fixed, tests enumerated
- `CODE_REVIEW.md` — SPEC coverage table, defensive correctness check, cross-script impact, spot-check, PASS verdict
- `BUG_REPORT.md` — O1-O4 observations; all non-blocking; documented behaviors per SPEC
- `ADVERSARIAL_BUG_REPORT.md` — A1-A8 sweep; no blockers; documented behaviors per SPEC
- `DOC_REVIEW.md` — docs touched: technical.md §7, functional.md §10, CHANGELOG.md; AGENTS/README/prompt intentionally untouched; cross-references confirmed
## Phase gates satisfied
| Phase | Artifact |
|-------|----------|
| research | SPEC.md ✓ |
| research:awaiting_approval | approved ✓ |
| implement | IMPLEMENTATION.md ✓ |
| code_review | CODE_REVIEW.md ✓ |
| code_review:awaiting_approval | approved ✓ |
| bug_find | BUG_REPORT.md ✓ |
| adversarial_bug_find | ADVERSARIAL_BUG_REPORT.md ✓ |
| doc_review | DOC_REVIEW.md ✓ |
| referee | VERDICT.md (this file) ✓ |
## Final acceptance criteria
1. **R1 (pass string coercion)**: ✓ three-way dispatch; `"true"`→True, `"false"`→False, others→`bool(...)`.
2. **R2 (score clamp [0,1])**: ✓ `max(0.0, min(1.0, score))`.
3. **R3 (non-numeric score → 0.5)**: ✓ `try/except (TypeError, ValueError)`.
4. **R4 (backwards compat)**: ✓ strict emitters unaffected; tested.
5. **R5 (no new deps)**: ✓ `math` stdlib only.
6. **Tests pass**: ✓ 469 passed (447 baseline + 22 new; 0 regressions).
7. **Docs in sync**: ✓ technical.md §7 + functional.md §10 + CHANGELOG.md updated.
## Inline bug found during implementation
The first iteration of the three-way branch set `verdict_pass = (raw_pass.strip().lower() == "true")`, mapping every non-`"true"` string to False. SPEC R1's `bool(...)` fallback clause was violated (`"yes"` would have regressed from True to False). The implementer caught this via `test_other_truthy_string_pass` before running the full suite, fixed the branch to explicit `if/elif/else: bool(...)`, and the test now guards the contract.
This is exactly the failure mode the phase pipeline is designed to surface: test-driven discovery of SPEC non-conformance during implement, not after deploy.
## Adversarial highlights
- `pass: null`/`[]`/`{}`/`0` → False, `pass: [false]` → True (Python truthy non-empty list). All match v1 `bool(...)` semantics; no regression.
- `score: NaN`/`Infinity`/`-Infinity` literals (json.loads accepts) → 0.5 via `math.isfinite`. Confirmed.
- `score: "2.0"` (out-of-range numeric string) → clamped to 1.0. The clamp fires after the try/except float() parse, so numeric strings are clamped too. Confirmed.
- `score: true` / `score: false` (JSON bool) → 1.0 / 0.0 via `float(True)` / `float(False)`. Documented quirk; downstream plateau gate handles consistently.
## Verdict
PASS — task is complete; all artifacts present; all phase gates satisfied; no blockers; no outstanding follow-ups for this task. The score-clamping improvement to plateau detection (BUG_REPORT.md O3) is a positive side effect noted in the CHANGELOG.
Approve transition to complete.