258 lines
9.7 KiB
Markdown
258 lines
9.7 KiB
Markdown
# Harden `parse_verdict`
|
|
|
|
Small pure-function hardening task: close the
|
|
`add-loop-runner/BUG_REPORT.md` O6 finding plus the un-noted sibling issue
|
|
(no score clamping). Both shipped in v1 because the verifier prompt's
|
|
contract layer was expected to enforce JSON-typed `pass` / numeric `score`
|
|
in `[0,1]`; observed real LLM responses and the O6 finding show the
|
|
runner should not trust the prompt contract alone.
|
|
|
|
This is a v1.1 hardening task. No CLI surface change; no new feature;
|
|
no schema migration. Pure robustness inside `scripts/loop-runner.py`'s
|
|
`parse_verdict`.
|
|
|
|
## Goal
|
|
|
|
`parse_verdict(text)` currently builds the verdict dict as:
|
|
|
|
```python
|
|
verdict = {
|
|
"pass": bool(data.get("pass")),
|
|
"score": float(data.get("score", 0.0)),
|
|
}
|
|
```
|
|
|
|
Two issues:
|
|
|
|
1. **`pass` string coercion bug (O6)**: if the verifier emits
|
|
`{"pass": "false", "score": 0.1}`, `bool("false")` returns `True`
|
|
(non-empty string is truthy). The tick records `pass=True`; the
|
|
`--check-gate` score-plateau brake, the tick log, and the orchestrator
|
|
downstream all see a "passing" tick when the verifier said "failing".
|
|
This is a silent correctness bug. Real LLMs do emit JSON booleans most
|
|
of the time, but OpenAI-grade models occasionally emit `"false"` /
|
|
`"true"` strings (quote-wrapped). The runner should accept both.
|
|
|
|
2. **Unclamped `score`**: if the verifier emits `"score": 1.5` or
|
|
`"score": -0.3` (out-of-contract), the value is stored as-is. The
|
|
`score_history` cap and the score-plateau brake assume `[0, 1]`. A
|
|
`1.5` value inflates the rolling-average computation; a `-0.3`
|
|
value causes `gate_score_plateau` to compute a negative trend that
|
|
looks like degradation when none exists. The verifier prompt
|
|
(`prompts/loop-verifier.md`) declares `score` is a float in `[0, 1]`,
|
|
but the runner should not rely on prompt-discipline alone.
|
|
|
|
## Requirements
|
|
|
|
### R1 — Coerce `pass` from string or bool
|
|
|
|
`parse_verdict` accepts the following as `data["pass"]`:
|
|
|
|
- `true` / `false` (JSON bool) — already correct via `json.loads`.
|
|
- `"true"` / `"false"` (JSON string) — case-insensitive. The string `"true"`
|
|
→ `True`; the string `"false"` → `False`. Any other non-empty string
|
|
→ fall through to the existing `bool(...)` semantics (i.e. truthy).
|
|
Empty string → `False`.
|
|
|
|
Implementation shape:
|
|
|
|
```python
|
|
raw_pass = data.get("pass")
|
|
if isinstance(raw_pass, str):
|
|
verdict_pass = raw_pass.strip().lower() == "true"
|
|
else:
|
|
verdict_pass = bool(raw_pass)
|
|
```
|
|
|
|
This handles `true`/`false` strings AND retains current behavior for
|
|
actual booleans (`True`/`False`) AND numbers (`0`/`1` — `bool(0)`
|
|
returns `False`; current behavior is preserved).
|
|
|
|
### R2 — Clamp `score` to `[0, 1]`
|
|
|
|
After parsing `float(data.get("score", 0.0))`, clamp:
|
|
|
|
```python
|
|
score = float(data.get("score", 0.0))
|
|
score = max(0.0, min(1.0, score))
|
|
```
|
|
|
|
NaN handling: `float("nan")` would propagate. If the verifier emits a
|
|
literal NaN (impossible in strict JSON; some Hermes-style models
|
|
occasionally emit it via `float('nan')` in reflowed text), the
|
|
`max/min` comparison returns NaN — both branches preserve NaN, NaN is
|
|
not equal to NaN, and score-plateau gate would see a constant NaN history.
|
|
Defensive: reject NaN / non-finite scores by treating them as 0.5 (the
|
|
verifier emitted something unusable; the midpoint is a neutral default).
|
|
Use `math.isfinite`:
|
|
|
|
```python
|
|
import math
|
|
score = float(data.get("score", 0.0))
|
|
if not math.isfinite(score):
|
|
score = 0.5
|
|
score = max(0.0, min(1.0, score))
|
|
```
|
|
|
|
### R3 — Defensive non-numeric `score`
|
|
|
|
If `data.get("score")` is a string like `"0.8"`, `float(...)` already
|
|
handles it (Python's `float` accepts string numerics). If it's a
|
|
non-numeric string, `float(...)` raises `ValueError`. Current code doesn't
|
|
catch this; would propagate as an unhandled exception mid-tick (→ halt via
|
|
the runner's finally → lock releases → tick log shows a HALT but errors
|
|
aren't categorized as `verifier_failed`). Wrap the float conversion:
|
|
|
|
```python
|
|
try:
|
|
score = float(data.get("score", 0.0))
|
|
except (TypeError, ValueError):
|
|
score = 0.5
|
|
```
|
|
|
|
### R4 — Backwards compat
|
|
|
|
- Verdict dict shape is unchanged: `{"pass": bool, "score": float,
|
|
"reasons": list[str], "next_hint": str}`. Existing callers (`cmd_tick`,
|
|
`_gate_score_plateau` indirectly via `score_history`, tick-log line
|
|
format) are unaffected.
|
|
- Strict-JSON emitters (true booleans, numeric scores in `[0,1]`) get
|
|
identical results to current behavior.
|
|
- The `score = 0.5` defaults for NaN / non-numeric are new behavior;
|
|
documented in the CHANGELOG.
|
|
|
|
### R5 — No new pip deps
|
|
|
|
`math` is stdlib.
|
|
|
|
## Test plan (`tests/test_parse_verdict.py`)
|
|
|
|
New file. Pure unit tests against `parse_verdict`; no subprocess, no
|
|
fixtures, no tmp_path needed. All use the function directly with literal
|
|
input strings.
|
|
|
|
1. `test_pass_true_bool` — `{"pass": true, "score": 0.8}` →
|
|
`pass=True, score=0.8`.
|
|
2. `test_pass_false_bool` — `{"pass": false, "score": 0.2}` →
|
|
`pass=False, score=0.2`.
|
|
3. `test_pass_true_string` — `{"pass": "true", "score": 0.9}` →
|
|
`pass=True, score=0.9` (the O6 bug).
|
|
4. `test_pass_false_string` — `{"pass": "false", "score": 0.1}` →
|
|
`pass=False, score=0.1` (the O6 bug).
|
|
5. `test_pass_string_case_insensitive` — `{"pass": "FALSE", "score": 0.1}`
|
|
→ `pass=False`. `{"pass": "True", "score": 0.9}` → `pass=True`.
|
|
6. `test_score_clamped_high` — `{"pass": true, "score": 1.5}` →
|
|
`pass=True, score=1.0`.
|
|
7. `test_score_clamped_low` — `{"pass": true, "score": -0.3}` →
|
|
`pass=True, score=0.0`.
|
|
8. `test_score_nan_to_neutral` — `{"pass": true, "score": NaN}` →
|
|
`pass=True, score=0.5`. (Use `float("nan")` literal in the test JSON
|
|
`__import__('math').nan` — actually use the string `"NaN"` to
|
|
simulate.
|
|
9. `test_score_non_numeric_string` — `{"pass": true, "score": "great"}`
|
|
→ `pass=True, score=0.5`.
|
|
10. `test_score_numeric_string_ok` — `{"pass": true, "score": "0.75"}`
|
|
→ `pass=True, score=0.75` (Python's float() already handles this; no
|
|
regression).
|
|
11. `test_empty_pass_string` — `{"pass": "", "score": 0.5}` →
|
|
`pass=False` (per R1's `bool(raw_pass)` fallback for non-`true`/`false`
|
|
strings: empty string → `bool("")` → `False`).
|
|
12. `test_other_truthy_string_pass` — `{"pass": "yes", "score": 0.5}` →
|
|
`pass=True` (`"yes"` is non-empty, non-`true` → `bool("yes")` is `True`).
|
|
Backwards-compat with prior semantics.
|
|
13. `test_existing_fence_block_behavior` — ```` ```json {"pass": true,
|
|
"score": 0.8} ``` ```` → still parses; new clamp/coerce don't break
|
|
the fence-extractor path.
|
|
|
|
## Concrete code shape
|
|
|
|
```python
|
|
import math
|
|
|
|
def parse_verdict(text: str) -> Optional[dict]:
|
|
"""Parse verifier JSON verdict. Accepts raw, fenced, or commented JSON.
|
|
|
|
Required keys: pass (bool — also accepts "true"/"false" strings),
|
|
score (float — clamped to [0, 1]; NaN/non-finite defaults to 0.5).
|
|
Optional: reasons (list[str]), next_hint (str). Returns None on parse
|
|
failure.
|
|
"""
|
|
if not text or not text.strip():
|
|
return None
|
|
candidates = []
|
|
fence_match = _FENCE_RE.search(text)
|
|
if fence_match:
|
|
candidates.append(fence_match.group(1))
|
|
candidates.append(text)
|
|
for body in candidates:
|
|
body = _strip_comments(body).strip()
|
|
if not body:
|
|
continue
|
|
try:
|
|
data = json.loads(body)
|
|
except json.JSONDecodeError:
|
|
continue
|
|
if not isinstance(data, dict):
|
|
continue
|
|
if "pass" not in data:
|
|
continue
|
|
raw_pass = data.get("pass")
|
|
if isinstance(raw_pass, str):
|
|
verdict_pass = raw_pass.strip().lower() == "true"
|
|
else:
|
|
verdict_pass = bool(raw_pass)
|
|
try:
|
|
score = float(data.get("score", 0.0))
|
|
except (TypeError, ValueError):
|
|
score = 0.5
|
|
if not math.isfinite(score):
|
|
score = 0.5
|
|
score = max(0.0, min(1.0, score))
|
|
verdict = {
|
|
"pass": verdict_pass,
|
|
"score": score,
|
|
}
|
|
if "reasons" in data and isinstance(data["reasons"], list):
|
|
verdict["reasons"] = [str(r) for r in data["reasons"]]
|
|
else:
|
|
verdict["reasons"] = []
|
|
if "next_hint" in data and isinstance(data["next_hint"], str):
|
|
verdict["next_hint"] = data["next_hint"]
|
|
return verdict
|
|
return None
|
|
```
|
|
|
|
## D-items (decisions locked for this task)
|
|
|
|
- **D-V1**: `"true"` / `"false"` strings → bool via case-insensitive
|
|
equality with `"true"`. Other strings defer to current `bool(...)` for
|
|
backwards compat (a verifier emitting `pass: "yes"` keeps current
|
|
truthy behavior).
|
|
- **D-V2**: `score` NaN / non-finite → `0.5` (neutral midpoint). This is
|
|
arbitrary but defensible; documented in CHANGELOG.
|
|
- **D-V3**: `score` non-numeric string → `0.5` (same neutral default).
|
|
Documented.
|
|
- **D-V4**: No CLI flag to opt out of clamping. Strict emitters in `[0,1]`
|
|
are unaffected; loose emitters get a deterministic value rather than a
|
|
raw one.
|
|
- **D-V5**: Tests are pure-functional; no subprocess; no monkeypatch.
|
|
|
|
## Non-goals
|
|
|
|
- No `parse_verdict` rewrite in `status.py`'s `_parse_verdict_status_line`
|
|
(different function, different concern — parses `VERDICT.md`
|
|
STATUS:PASS / FAIL strings; not in scope for this task).
|
|
- No `loop-verifier.md` prompt changes (the prompt still asks for JSON
|
|
booleans; the runner-side coercion is defense-in-depth). Verifier
|
|
prompt changes are tracked separately.
|
|
- No schema change to `.state.loop` `score_history` — existing floats
|
|
already in `[0,1]` from prior ticks are unaffected; new ticks are
|
|
clamped.
|
|
- No `verifier_failed` halt prompt change.
|
|
|
|
## Verification
|
|
|
|
- `python3 -m py_compile scripts/loop-runner.py`
|
|
- `python3 -m pytest tests/test_parse_verdict.py -v`
|
|
- `python3 -m pytest tests/ -q` (full suite stays green; baseline 447 +
|
|
new) |