Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
CI / build (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1 @@
|
||||
complete
|
||||
@@ -0,0 +1,2 @@
|
||||
research:approved|2026-06-24T00:08:32.650347+00:00|user
|
||||
code_review:approved|2026-06-24T00:09:51.306902+00:00|user
|
||||
@@ -0,0 +1,155 @@
|
||||
# Adversarial Bug Report: harden-parse-verdict
|
||||
|
||||
Probed `parse_verdict` with non-contract inputs. Each attack vector
|
||||
hypothesized, tested, verdict given.
|
||||
|
||||
## A1 — `pass` as Python types (None, list, dict, int) — type-confusion
|
||||
|
||||
**Hypothesis**: A verifier emitting non-string non-bool `pass` values
|
||||
(e.g. `{"pass": null}`, `{"pass": [false]}`, `{"pass": 0}`) could yield
|
||||
surprising verdicts.
|
||||
|
||||
**Test**: 9-row sweep via `json.dumps` (Python `None` → JSON `null`,
|
||||
Python `True`/`False` → JSON `true`/`false`):
|
||||
|
||||
| Input | Output | Notes |
|
||||
|---|---|---|
|
||||
| `pass: null` | `pass=False` | `bool(None)` = False; preserved from v1. |
|
||||
| `pass: []` (empty list) | `pass=False` | `bool([])` = False; preserved. |
|
||||
| `pass: [false]` (list with False) | `pass=True` | `bool([False])` = True (non-empty list is truthy). Surprising but documented Python semantics. v1 returned same. **Not a regression.** |
|
||||
| `pass: [true]` | `pass=True` | Same. |
|
||||
| `pass: {}` (empty dict) | `pass=False` | `bool({})` = False; preserved. |
|
||||
| `pass: 0` (int) | `pass=False` | `bool(0)` = False; preserved. |
|
||||
| `pass: 1` (int) | `pass=True` | `bool(1)` = True; preserved. |
|
||||
| `pass: -1` (int) | `pass=True` | `bool(-1)` = True (non-zero); preserved. |
|
||||
| `pass: 1.5` (float) | `pass=True` | `bool(1.5)` = True (non-zero); preserved. |
|
||||
|
||||
**Verdict**: PASS — no regression for any non-string non-bool type. RESET
|
||||
behavior matches v1's `bool(...)` semantics. The SPEC's three-way
|
||||
string/bool branch handles strings explicitly; everything else falls
|
||||
through to v1's `bool(...)`.
|
||||
|
||||
## A2 — `score` as Python non-numeric types
|
||||
|
||||
**Hypothesis**: `score: []`, `score: {}`, `score: [1, 2]`, `score: "high"`
|
||||
should default to 0.5 per R3 (TypeError / ValueError caught).
|
||||
|
||||
**Test**:
|
||||
|
||||
| Input | Output |
|
||||
|---|---|
|
||||
| `score: []` | `score=0.5` (TypeError caught by `float([])`) |
|
||||
| `score: {}` | `score=0.5` (TypeError caught) |
|
||||
| `score: [1, 2]` | `score=0.5` (TypeError caught) |
|
||||
| `score: "high"` | `score=0.5` (ValueError caught) |
|
||||
| `score: None` | `score=0.5` (TypeError caught) |
|
||||
|
||||
**Verdict**: PASS — R3's `except (TypeError, ValueError)` catches all
|
||||
non-numeric types; defaults to 0.5 (D-V3). Confirmed.
|
||||
|
||||
## A3 — `score` as out-of-range numeric strings
|
||||
|
||||
**Hypothesis**: A verifier emitting `score: "2.0"` (an out-of-range
|
||||
numeric STRING) bypasses the clamp because R3's except arm never fires
|
||||
and R2's clamp applies after — but is the clamp correctly triggered?
|
||||
|
||||
**Test**:
|
||||
|
||||
| Input | Output |
|
||||
|---|---|
|
||||
| `score: "2.0"` | `score=1.0` (parse to 2.0, clamp to 1.0) |
|
||||
| `score: "-0.5"` | `score=0.0` (parse to -0.5, clamp to 0.0) |
|
||||
| `score: "0.75"` | `score=0.75` (parse to 0.75, no clamping) |
|
||||
|
||||
**Verdict**: PASS — clamping applies to all numeric inputs regardless of
|
||||
whether they came in as JSON numbers or numeric strings. Confirmed in
|
||||
the SPEC test plan (`test_score_numeric_string_ok` and the inline fix).
|
||||
|
||||
## A4 — `score` as JSON literal NaN / Infinity / -Infinity
|
||||
|
||||
**Hypothesis**: Some Hermes-style recursive decoders emit the bare
|
||||
tokens `NaN` / `Infinity` / `-Infinity` (rejected by strict JSON but
|
||||
accepted by Python's `json.loads` with the default `parse_constant`).
|
||||
`float(NaN)` succeeds (returns `math.nan`). The `math.isfinite` check
|
||||
catches it.
|
||||
|
||||
**Test**:
|
||||
|
||||
| Input | Output |
|
||||
|---|---|
|
||||
| `score: NaN` (bare token) | `score=0.5` (isfinite catches; D-V2 default) |
|
||||
| `score: Infinity` (bare token) | `score=0.5` |
|
||||
| `score: -Infinity` (bare token) | `score=0.5` |
|
||||
|
||||
**Verdict**: PASS — D-V2 documented neutral default. The `math.isfinite`
|
||||
guard fires before the clamp so the NaN doesn't propagate through `max` /
|
||||
`min`.
|
||||
|
||||
## A5 — `score` as JSON booleans (true / false)
|
||||
|
||||
**Hypothesis**: A verifier erroneously using `"score": true` instead of
|
||||
`"score": 0.8` would yield `float(True)` = 1.0 in Python (no exception),
|
||||
then clamp to 1.0 (no change). The result is "the verifier said pass
|
||||
with a perfect score" — incorrect but not a crash. Is this OK?
|
||||
|
||||
**Test**:
|
||||
|
||||
| Input | Output |
|
||||
|---|---|
|
||||
| `score: true` | `score=1.0` (float(True) → max(0, min(1, 1.0)) → 1.0) |
|
||||
| `score: false` | `score=0.0` (float(False) → 0.0) |
|
||||
|
||||
**Verdict**: PASS — `float(True)` is well-defined in Python. A verifier
|
||||
mis-typing `score: true` produces a deterministic 1.0 (not a crash; not
|
||||
NaN). Score-plateau gate will see consistent 1.0 across ticks → halt as
|
||||
`score_plateau`. Reasonable downstream behavior; documented quirk.
|
||||
|
||||
## A6 — `pass` short strings ("t", "T", "f")
|
||||
|
||||
**Hypothesis**: A verifier abbreviating `pass: "t"` or `pass: "T"` might
|
||||
be misread as True (since SPEC only says `"true"`/`"false"` exact match
|
||||
maps to True/False). Per SPEC R1, other strings fall through to
|
||||
`bool(...)`, which is truthy for non-empty.
|
||||
|
||||
**Test**:
|
||||
|
||||
| Input | Output |
|
||||
|---|---|
|
||||
| `pass: "t"` | `pass=True` (abstract: `bool("t")` = True; not "true") |
|
||||
| `pass: "T"` | `pass=True` |
|
||||
| `pass: "f"` | `pass=True` (truthy; surprising!) |
|
||||
|
||||
**Verdict**: PASS — documented behavior. Risk: a verifier emitting
|
||||
`pass: "f"` intending "false" gets `True`. Same as v1. The SPEC's
|
||||
contract is to use full `true`/`false` strings or JSON booleans. This
|
||||
abbreviated-string case is undocumented but not a regression; future
|
||||
prompt work (out of scope for this task) should discourage abbreviations.
|
||||
|
||||
## A7 — Combined: `pass: "false"` string with `score: NaN` literal — full
|
||||
harsh-path coverage
|
||||
|
||||
**Hypothesis**: A both-broken verdict still yields a parseable dict with
|
||||
coerced defaults rather than None.
|
||||
|
||||
**Test**: `{"pass": "false", "score": NaN}` literal — `parse_verdict`
|
||||
returns `{"pass": False, "score": 0.5, "reasons": [], "next_hint": ""}`.
|
||||
|
||||
**Verdict**: PASS — both coercion paths fire; documented defaults applied.
|
||||
|
||||
## A8 — Whitespace-only strips: newline + tab in `pass` value
|
||||
|
||||
**Hypothesis**: A verifier emitting `pass: "\n true "` (whitespace-wrapped)
|
||||
should yield True after `.strip()`.
|
||||
|
||||
**Confirmed via test `test_pass_with_surrounding_whitespace`** — `" true "`
|
||||
strips cleanly. Newline/tab characters not explicitly tested but
|
||||
`str.strip()` defaults to all whitespace; newlines strip too.
|
||||
|
||||
**Verdict**: PASS.
|
||||
|
||||
## No BLOCKERS
|
||||
|
||||
A1-A8 are all documented behaviors per SPEC R1+R2+R3 + D-V1/D-V2/D-V3.
|
||||
All inputs that would have caused silent corruption (the `bool("false")=True`
|
||||
bug) or crashes (TypeError from non-numeric scores) are now handled
|
||||
defensively. Recommend proceeding to doc_review.
|
||||
@@ -0,0 +1,78 @@
|
||||
# Bug Report: harden-parse-verdict
|
||||
|
||||
Bug_find phase observations. Each non-blocking unless marked BLOCKER.
|
||||
|
||||
## O1 — `bool(raw_pass.strip())` fallback for "0" / "1" strings
|
||||
|
||||
For `{"pass": "0", "score": 0.5}`, the new path strips → "0" → not "true"/
|
||||
"false" → `bool("0")` = True (non-empty string is truthy).
|
||||
|
||||
Pre-fix: `bool("0")` was also True (same).
|
||||
|
||||
This is **consistent with v1** — no behavior change. A verifier emitting
|
||||
`pass: "0"` intending "false" gets `True` in v1 AND in the hardened
|
||||
implementation. The SPEC says non-`true`/`false` strings fall through to
|
||||
`bool(...)`, which is truthy for non-empty. Not a regression.
|
||||
|
||||
**Not a bug** — documented behavior per SPEC R1 + D-V1. If users want
|
||||
strict numeric-string handling, that's a separate future task (out of
|
||||
scope for v1.1 harden-parse-verdict).
|
||||
|
||||
## O2 — `float("nan")` serializes back as `NaN` to `.state.loop`
|
||||
|
||||
When the verifier emits `NaN` as the score, `parse_verdict` returns
|
||||
`score=0.5` (without writing to disk by itself). But this score is part
|
||||
of `verdict` which gets written via `_write_state_loop(loop_path, state)`
|
||||
into `.state.loop` as JSON. Since the clamp converts NaN to 0.5 BEFORE
|
||||
the verdict is stored, `.state.loop` gets `0.5`, not `NaN`. No NaN leaks
|
||||
into the loop state.
|
||||
|
||||
**Not a bug** — confirmed via tracing: `parse_verdict` returns the clamped
|
||||
dict; `cmd_tick` then stores `state["last_verdict"] = verdict` (a clean
|
||||
dict with `score: 0.5`); the JSON round-trip is clean.
|
||||
|
||||
## O3 — `_gate_score_plateau`'s threshold unaffected
|
||||
|
||||
`_gate_score_plateau` (status.py) reads `score_history` and decides a halt
|
||||
when last N scores are within some delta. With clamped scores, the plateau
|
||||
detection range is now strictly `[0, 1]` instead of `[any, any]`. Brief
|
||||
review:
|
||||
|
||||
- Pre-fix: a verifier could emit `score: 1.5` across N ticks; plateau
|
||||
detection sees a flat line at 1.5; halt fires. Expected behavior.
|
||||
- Post-fix: the same verifier's 1.5 clamps to 1.0 across N ticks; plateau
|
||||
sees flat line at 1.0; halt fires. Same outcome.
|
||||
|
||||
- Pre-fix: a verifier emits alternating `0.9` and `1.1`; plateau sees a
|
||||
bimodal history [0.9, 1.1, 0.9, 1.1] — NOT plateau (variation > epsilon).
|
||||
- Post-fix: alternating `0.9` and `1.0` (1.1 clamps to 1.0); plateau sees
|
||||
[0.9, 1.0, 0.9, 1.0] — still variation above a small epsilon — still NOT
|
||||
plateau. Same outcome in this scenario.
|
||||
|
||||
Edge case: a verifier emits all `1.0` and `0.99` (vs 1.0 and 1.0 clamped).
|
||||
The clamp DOES change plateau detection in this case — `1.0 1.0 1.0`
|
||||
looks more plateau-like than `1.0 0.99 0.99`. Could cause halt earlier than
|
||||
prior. Documented as a desirable side effect (clamp reduces the verifier's
|
||||
untrustworthiness from inflating scores; plateau detection is more
|
||||
honest).
|
||||
|
||||
**Not a bug** — improved behavior. Documented in CHANGELOG.
|
||||
|
||||
## O4 — Verifier prompt hasn't been updated
|
||||
|
||||
`prompts/loop-verifier.md` still asks the model to emit JSON with
|
||||
`"pass": true/false` and `"score": 0.0-1.0`. The runner now defensive-coerces,
|
||||
but the prompt's contract is unchanged. Was the prompt already
|
||||
JSON-typed-booleans-only? Let me check.
|
||||
|
||||
**`prompts/loop-verifier.md` review**: still says "Output: strict JSON, no
|
||||
prose" with example shape. The prompt explicitly tells the model to emit
|
||||
JSON booleans — no mention of string-typed `pass`. So the v1 contract was
|
||||
strict; the O6 finding was a defense-in-depth concern, not a present-fault.
|
||||
The harden task adds belt-and-suspenders without changing the contract.
|
||||
|
||||
**Not a bug** — the prompt remains authoritative. No edit needed.
|
||||
|
||||
## Verdict
|
||||
|
||||
**No blockers.** Proceed to adversarial_bug_find.
|
||||
@@ -0,0 +1,67 @@
|
||||
# Code Review: harden-parse-verdict
|
||||
|
||||
## SPEC coverage
|
||||
|
||||
| Requirement | Status |
|
||||
|-------------|--------|
|
||||
| R1 — Coerce `pass` from string or bool (3-way dispatch: true / false / other) | ✓ — three-way branch with case-insensitive match + `bool(raw_pass.strip())` fallback |
|
||||
| R2 — Clamp `score` to `[0, 1]` via `max(0, min(1, x))` | ✓ |
|
||||
| R3 — Defensive non-numeric `score` (try/except TypeError, ValueError) | ✓ — caught and defaulted to 0.5 |
|
||||
| R4 — Backwards compat (dict shape unchanged; strict emitters unaffected) | ✓ — verdict still has `pass`, `score`, `reasons`, `next_hint` keys; strict JSON emitters get identical results to v1 |
|
||||
| R5 — No new pip deps (`math` stdlib) | ✓ |
|
||||
|
||||
## Code readability
|
||||
|
||||
- Three-way branch is more verbose than the v1 single-line `bool(...)` but
|
||||
the case-intent is clearer: the comment "case-insensitive. The string
|
||||
'true' → True; the string 'false' → False. Any other non-empty string
|
||||
→ fall through to the existing `bool(...)` semantics" in the SPEC is
|
||||
preserved exactly by the if/elif/else.
|
||||
- The score-clamp block is two statements (try/except, then isfinite
|
||||
check, then clamp). The order matters: the TypeError/ValueError from
|
||||
`float(None)` or `float("great")` must be caught BEFORE the `math.isfinite`
|
||||
call; else `math.isfinite(None)` raises TypeError uncaught. The order
|
||||
in the implementation is correct (try/except wraps the float call;
|
||||
isfinite only sees a finite-or-NaN float).
|
||||
|
||||
## Defensive correctness check
|
||||
|
||||
- `bool(None)` → False (if `data` is `{"pass": None}`; treated as no-pass
|
||||
→ False; matches pre-fix `bool(None)` = False; no regression).
|
||||
- `bool(0)` → False (if verifier emits `"pass": 0`); preserved.
|
||||
- `bool(1)` → True; `bool([])` False; `bool({})` False; all preserved — no
|
||||
regression for non-string types.
|
||||
- String `" tRuE "` strips via `.strip().lower()` → "true" → True.
|
||||
- String `"\nfalse"` strips → "false" → False. Edge case covered.
|
||||
|
||||
## Cross-script impact
|
||||
|
||||
- `parse_verdict` is local to `loop-runner.py`; not duplicated to
|
||||
`status.py`. The change is contained.
|
||||
- `_gate_score_plateau` in `status.py` consumes `score_history` (with
|
||||
clamped values via the runner's atomic write of `last_verdict`) —
|
||||
already assumed `[0, 1]`. The clamp guarantees it.
|
||||
- `_write_state_loop` timestamps store JSON; clamped scores
|
||||
round-trip cleanly (no serialization loss).
|
||||
|
||||
## Tests spot-check
|
||||
|
||||
- `test_other_truthy_string_pass` (the bug found inline): verifies that
|
||||
the SPEC R1's `bool(...)` fallback clause is honored — a `"yes"` string
|
||||
yields `True`. Pre-fix v1 behavior preserved.
|
||||
- `test_pass_with_surrounding_whitespace`: covers an edge case (`" true "`)
|
||||
the SPEC didn't explicitly enumerate but is sensible.
|
||||
- `test_score_none_value_to_neutral`: `data.get("score", 0.0)` returns
|
||||
`None` (key exists with None value); `float(None)` raises TypeError →
|
||||
caught → 0.5. Not in SPEC's explicit test plan but is a natural
|
||||
consequence of R3's TypeError coverage. Good defensive test.
|
||||
- `test_score_infinity_to_neutral`: covers `math.isfinite(Infinity)` →
|
||||
False path. Added after `test_score_nan_to_neutral`; both prove the
|
||||
isfinite check.
|
||||
- All 22 tests pass.
|
||||
|
||||
## Verdict
|
||||
|
||||
PASS — implementer followed SPEC; inline bug found and fixed during
|
||||
test; the fix matches SPEC R1's three-way dispatch wording exactly.
|
||||
Proceed to bug_find.
|
||||
@@ -0,0 +1,45 @@
|
||||
# Doc Review: harden-parse-verdict
|
||||
|
||||
## Docs touched
|
||||
|
||||
- `design/loops/technical.md` §7 — tick-flow step 7 (parse verdict):
|
||||
added 4-line inline block documenting the defensive coercion (pass
|
||||
string acceptance; score clamp + NaN/inf/non-numeric → 0.5).
|
||||
- `design/loops/functional.md` §10 — Verifier Contract: annotated
|
||||
`pass` (bool) definitive + runner accepts `"true"`/`"false"` strings;
|
||||
annotated `score` (0.0–1.0) clamp + NaN/inf/non-numeric → 0.5 neutral.
|
||||
- `CHANGELOG.md` — new `[unreleased]` "Fixed — `parse_verdict`
|
||||
defensive coercion" block above the existing `add-state-loop-lock` and
|
||||
`fix-harness-command-template` blocks.
|
||||
|
||||
## Docs NOT touched (intentional)
|
||||
|
||||
- `AGENTS.md`: parse_verdict is not a user-visible CLI surface; the
|
||||
hardening doesn't change phase enforcement, `.state.loop`, or any
|
||||
contract that harness integrators need to know. The Verifier Contract
|
||||
lives in `design/loops/functional.md` §10; AGENTS.md already points to
|
||||
design docs at the top. No edit.
|
||||
- `README.md`: user-facing README doesn't enumerate `parse_verdict`
|
||||
internals; loop monitoring table mentions verdicts as a concept, not
|
||||
the parser. No edit.
|
||||
- `prompts/loop-verifier.md`: contract was already `bool pass` + `score
|
||||
0.0–1.0`. The hardening is belt-and-suspenders against malformed
|
||||
output, not a contract change. The prompt's strict-JSON directive
|
||||
stays authoritative. No edit.
|
||||
- `templates/loops/self-improvement/loop.json`: no schema change. No
|
||||
edit.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- `tasks/add-loop-runner/BUG_REPORT.md` O6 — the original finding —
|
||||
now closed by this task. The CHANGELOG entry explicitly references it.
|
||||
- `tasks/add-loop-runner/ADVERSARIAL_BUG_REPORT.md` A6 — the score-
|
||||
clamping observation — also closed by this task. The CHANGELOG entry
|
||||
references the clamping.
|
||||
- `tasks/harden-parse-verdict/BUG_REPORT.md` O3 — notes that clamping
|
||||
improves plateau detection (a tighter `score_history` range makes
|
||||
plateau more honest). Cross-referenced from the CHANGELOG.
|
||||
|
||||
## Verdict
|
||||
|
||||
Docs are in sync with the implementation. Proceed to referee.
|
||||
@@ -0,0 +1,98 @@
|
||||
# Implementation: harden-parse-verdict
|
||||
|
||||
## SCOPE
|
||||
|
||||
Closed `add-loop-runner/BUG_REPORT.md` O6 (pass-string coercion bug) plus
|
||||
the un-noted sibling issue (no score clamping). Pure-function change to
|
||||
`parse_verdict` in `scripts/loop-runner.py`. No CLI surface change; no
|
||||
schema change; no new deps (math is stdlib).
|
||||
|
||||
## FILES TOUCHED
|
||||
|
||||
- `scripts/loop-runner.py`
|
||||
- Added `import math`.
|
||||
- `parse_verdict(text)`: replaced `verdict["pass"] = bool(data.get("pass"))`
|
||||
with explicit string-vs-bool dispatch:
|
||||
- bool passed through → `bool(True)` = True; `bool(False)` = False (unchanged).
|
||||
- `"true"` (any case, leading/trailing whitespace stripped) → True.
|
||||
- `"false"` (any case, leading/trailing whitespace stripped) → False.
|
||||
- Any other string → `bool(raw_pass.strip())` (empty → False; non-empty → True).
|
||||
Preserves old `bool(...)` truthy semantics for `"yes"` / etc.
|
||||
- Replaced `verdict["score"] = float(data.get("score", 0.0))` with:
|
||||
- `try: score = float(data.get("score", 0.0))` /
|
||||
`except (TypeError, ValueError): score = 0.5`
|
||||
(TypeError for non-numeric types like None/list/dict; ValueError for
|
||||
non-numeric strings like "great").
|
||||
- `if not math.isfinite(score): score = 0.5` (catches NaN, Infinity,
|
||||
-Infinity returned by some Hermes-style recursive decoders).
|
||||
- `score = max(0.0, min(1.0, score))` (clamp to `[0, 1]`).
|
||||
- Updated docstring to spell out the new contract: `pass` accepts
|
||||
bool OR `"true"`/`"false"` strings (case-insensitive); `score` is
|
||||
clamped to `[0, 1]` with NaN/non-finite → 0.5.
|
||||
|
||||
## D-ITEMS Locked
|
||||
|
||||
- **D-V1**: `"true"` / `"false"` strings → bool via case-insensitive
|
||||
equality. Other strings defer to current `bool(...)` for backwards
|
||||
compat (`pass: "yes"` stays truthy).
|
||||
- **D-V2**: `score` NaN / non-finite → `0.5`.
|
||||
- **D-V3**: `score` non-numeric string → `0.5`.
|
||||
- **D-V4**: No opt-out flag for clamping. Strict emitters unaffected.
|
||||
- **D-V5**: Tests are pure-functional; no subprocess.
|
||||
|
||||
## BUG FOUND AND FIXED INLINE
|
||||
|
||||
While running `tests/test_parse_verdict.py`, the test
|
||||
`test_other_truthy_string_pass` failed on the first iteration. The
|
||||
implementation had:
|
||||
|
||||
```python
|
||||
if isinstance(raw_pass, str):
|
||||
verdict_pass = raw_pass.strip().lower() == "true"
|
||||
```
|
||||
|
||||
This treats EVERY non-`"true"` string as False — including `"yes"`,
|
||||
which used to be True via `bool("yes")`. SPEC R1 explicitly says
|
||||
non-`true`/`false` strings fall through to `bool(...)` for backwards
|
||||
compat. Fixed to the three-way branch:
|
||||
|
||||
```python
|
||||
if isinstance(raw_pass, str):
|
||||
lower = raw_pass.strip().lower()
|
||||
if lower == "true":
|
||||
verdict_pass = True
|
||||
elif lower == "false":
|
||||
verdict_pass = False
|
||||
else:
|
||||
verdict_pass = bool(raw_pass.strip())
|
||||
```
|
||||
|
||||
This is consistent with SPEC R1 wording. All 22 tests pass after the fix.
|
||||
|
||||
## TESTS
|
||||
|
||||
New file `tests/test_parse_verdict.py` — 22 tests across 5 classes:
|
||||
|
||||
- `TestStrictBaseline` (2): bool `pass` true/false; preserves existing semantics.
|
||||
- `TestPassStringCoercion` (6): `"true"`/`"false"` strings, case-insensitive,
|
||||
surrounding whitespace, empty string, other truthy string.
|
||||
- `TestScoreClamping` (9): clamped high (1.5→1.0), clamped low (-0.3→0.0),
|
||||
edges (0.0, 1.0), NaN → 0.5, Infinity → 0.5, non-numeric string "great"
|
||||
→ 0.5, numeric string "0.75" → 0.75, missing score → 0.0, None score →
|
||||
0.5.
|
||||
- `TestFenceBlockStillWorks` (2): existing fence-block path with bool pass;
|
||||
fence path with string-pass + clamped score.
|
||||
- `TestOptionalKeysPreserved` (3): reasons+next_hint combination, missing
|
||||
reasons → [], non-list reasons coerced to [].
|
||||
|
||||
## TEST COUNT
|
||||
|
||||
- Baseline: 447 passed (post-`add-state-loop-lock`).
|
||||
- New: +22 in `tests/test_parse_verdict.py`.
|
||||
- Final: **469 passed**, 0 regressions.
|
||||
|
||||
## PIPELINE TO COMPLETION
|
||||
|
||||
research → research:awaiting_approval → research:approved → implement.
|
||||
Next: → code_review → code_review:awaiting_approval → code_review:approved
|
||||
→ bug_find → adversarial_bug_find → doc_review → referee → complete.
|
||||
@@ -0,0 +1,258 @@
|
||||
# Harden `parse_verdict`
|
||||
|
||||
Small pure-function hardening task: close the
|
||||
`add-loop-runner/BUG_REPORT.md` O6 finding plus the un-noted sibling issue
|
||||
(no score clamping). Both shipped in v1 because the verifier prompt's
|
||||
contract layer was expected to enforce JSON-typed `pass` / numeric `score`
|
||||
in `[0,1]`; observed real LLM responses and the O6 finding show the
|
||||
runner should not trust the prompt contract alone.
|
||||
|
||||
This is a v1.1 hardening task. No CLI surface change; no new feature;
|
||||
no schema migration. Pure robustness inside `scripts/loop-runner.py`'s
|
||||
`parse_verdict`.
|
||||
|
||||
## Goal
|
||||
|
||||
`parse_verdict(text)` currently builds the verdict dict as:
|
||||
|
||||
```python
|
||||
verdict = {
|
||||
"pass": bool(data.get("pass")),
|
||||
"score": float(data.get("score", 0.0)),
|
||||
}
|
||||
```
|
||||
|
||||
Two issues:
|
||||
|
||||
1. **`pass` string coercion bug (O6)**: if the verifier emits
|
||||
`{"pass": "false", "score": 0.1}`, `bool("false")` returns `True`
|
||||
(non-empty string is truthy). The tick records `pass=True`; the
|
||||
`--check-gate` score-plateau brake, the tick log, and the orchestrator
|
||||
downstream all see a "passing" tick when the verifier said "failing".
|
||||
This is a silent correctness bug. Real LLMs do emit JSON booleans most
|
||||
of the time, but OpenAI-grade models occasionally emit `"false"` /
|
||||
`"true"` strings (quote-wrapped). The runner should accept both.
|
||||
|
||||
2. **Unclamped `score`**: if the verifier emits `"score": 1.5` or
|
||||
`"score": -0.3` (out-of-contract), the value is stored as-is. The
|
||||
`score_history` cap and the score-plateau brake assume `[0, 1]`. A
|
||||
`1.5` value inflates the rolling-average computation; a `-0.3`
|
||||
value causes `gate_score_plateau` to compute a negative trend that
|
||||
looks like degradation when none exists. The verifier prompt
|
||||
(`prompts/loop-verifier.md`) declares `score` is a float in `[0, 1]`,
|
||||
but the runner should not rely on prompt-discipline alone.
|
||||
|
||||
## Requirements
|
||||
|
||||
### R1 — Coerce `pass` from string or bool
|
||||
|
||||
`parse_verdict` accepts the following as `data["pass"]`:
|
||||
|
||||
- `true` / `false` (JSON bool) — already correct via `json.loads`.
|
||||
- `"true"` / `"false"` (JSON string) — case-insensitive. The string `"true"`
|
||||
→ `True`; the string `"false"` → `False`. Any other non-empty string
|
||||
→ fall through to the existing `bool(...)` semantics (i.e. truthy).
|
||||
Empty string → `False`.
|
||||
|
||||
Implementation shape:
|
||||
|
||||
```python
|
||||
raw_pass = data.get("pass")
|
||||
if isinstance(raw_pass, str):
|
||||
verdict_pass = raw_pass.strip().lower() == "true"
|
||||
else:
|
||||
verdict_pass = bool(raw_pass)
|
||||
```
|
||||
|
||||
This handles `true`/`false` strings AND retains current behavior for
|
||||
actual booleans (`True`/`False`) AND numbers (`0`/`1` — `bool(0)`
|
||||
returns `False`; current behavior is preserved).
|
||||
|
||||
### R2 — Clamp `score` to `[0, 1]`
|
||||
|
||||
After parsing `float(data.get("score", 0.0))`, clamp:
|
||||
|
||||
```python
|
||||
score = float(data.get("score", 0.0))
|
||||
score = max(0.0, min(1.0, score))
|
||||
```
|
||||
|
||||
NaN handling: `float("nan")` would propagate. If the verifier emits a
|
||||
literal NaN (impossible in strict JSON; some Hermes-style models
|
||||
occasionally emit it via `float('nan')` in reflowed text), the
|
||||
`max/min` comparison returns NaN — both branches preserve NaN, NaN is
|
||||
not equal to NaN, and score-plateau gate would see a constant NaN history.
|
||||
Defensive: reject NaN / non-finite scores by treating them as 0.5 (the
|
||||
verifier emitted something unusable; the midpoint is a neutral default).
|
||||
Use `math.isfinite`:
|
||||
|
||||
```python
|
||||
import math
|
||||
score = float(data.get("score", 0.0))
|
||||
if not math.isfinite(score):
|
||||
score = 0.5
|
||||
score = max(0.0, min(1.0, score))
|
||||
```
|
||||
|
||||
### R3 — Defensive non-numeric `score`
|
||||
|
||||
If `data.get("score")` is a string like `"0.8"`, `float(...)` already
|
||||
handles it (Python's `float` accepts string numerics). If it's a
|
||||
non-numeric string, `float(...)` raises `ValueError`. Current code doesn't
|
||||
catch this; would propagate as an unhandled exception mid-tick (→ halt via
|
||||
the runner's finally → lock releases → tick log shows a HALT but errors
|
||||
aren't categorized as `verifier_failed`). Wrap the float conversion:
|
||||
|
||||
```python
|
||||
try:
|
||||
score = float(data.get("score", 0.0))
|
||||
except (TypeError, ValueError):
|
||||
score = 0.5
|
||||
```
|
||||
|
||||
### R4 — Backwards compat
|
||||
|
||||
- Verdict dict shape is unchanged: `{"pass": bool, "score": float,
|
||||
"reasons": list[str], "next_hint": str}`. Existing callers (`cmd_tick`,
|
||||
`_gate_score_plateau` indirectly via `score_history`, tick-log line
|
||||
format) are unaffected.
|
||||
- Strict-JSON emitters (true booleans, numeric scores in `[0,1]`) get
|
||||
identical results to current behavior.
|
||||
- The `score = 0.5` defaults for NaN / non-numeric are new behavior;
|
||||
documented in the CHANGELOG.
|
||||
|
||||
### R5 — No new pip deps
|
||||
|
||||
`math` is stdlib.
|
||||
|
||||
## Test plan (`tests/test_parse_verdict.py`)
|
||||
|
||||
New file. Pure unit tests against `parse_verdict`; no subprocess, no
|
||||
fixtures, no tmp_path needed. All use the function directly with literal
|
||||
input strings.
|
||||
|
||||
1. `test_pass_true_bool` — `{"pass": true, "score": 0.8}` →
|
||||
`pass=True, score=0.8`.
|
||||
2. `test_pass_false_bool` — `{"pass": false, "score": 0.2}` →
|
||||
`pass=False, score=0.2`.
|
||||
3. `test_pass_true_string` — `{"pass": "true", "score": 0.9}` →
|
||||
`pass=True, score=0.9` (the O6 bug).
|
||||
4. `test_pass_false_string` — `{"pass": "false", "score": 0.1}` →
|
||||
`pass=False, score=0.1` (the O6 bug).
|
||||
5. `test_pass_string_case_insensitive` — `{"pass": "FALSE", "score": 0.1}`
|
||||
→ `pass=False`. `{"pass": "True", "score": 0.9}` → `pass=True`.
|
||||
6. `test_score_clamped_high` — `{"pass": true, "score": 1.5}` →
|
||||
`pass=True, score=1.0`.
|
||||
7. `test_score_clamped_low` — `{"pass": true, "score": -0.3}` →
|
||||
`pass=True, score=0.0`.
|
||||
8. `test_score_nan_to_neutral` — `{"pass": true, "score": NaN}` →
|
||||
`pass=True, score=0.5`. (Use `float("nan")` literal in the test JSON
|
||||
`__import__('math').nan` — actually use the string `"NaN"` to
|
||||
simulate.
|
||||
9. `test_score_non_numeric_string` — `{"pass": true, "score": "great"}`
|
||||
→ `pass=True, score=0.5`.
|
||||
10. `test_score_numeric_string_ok` — `{"pass": true, "score": "0.75"}`
|
||||
→ `pass=True, score=0.75` (Python's float() already handles this; no
|
||||
regression).
|
||||
11. `test_empty_pass_string` — `{"pass": "", "score": 0.5}` →
|
||||
`pass=False` (per R1's `bool(raw_pass)` fallback for non-`true`/`false`
|
||||
strings: empty string → `bool("")` → `False`).
|
||||
12. `test_other_truthy_string_pass` — `{"pass": "yes", "score": 0.5}` →
|
||||
`pass=True` (`"yes"` is non-empty, non-`true` → `bool("yes")` is `True`).
|
||||
Backwards-compat with prior semantics.
|
||||
13. `test_existing_fence_block_behavior` — ```` ```json {"pass": true,
|
||||
"score": 0.8} ``` ```` → still parses; new clamp/coerce don't break
|
||||
the fence-extractor path.
|
||||
|
||||
## Concrete code shape
|
||||
|
||||
```python
|
||||
import math
|
||||
|
||||
def parse_verdict(text: str) -> Optional[dict]:
|
||||
"""Parse verifier JSON verdict. Accepts raw, fenced, or commented JSON.
|
||||
|
||||
Required keys: pass (bool — also accepts "true"/"false" strings),
|
||||
score (float — clamped to [0, 1]; NaN/non-finite defaults to 0.5).
|
||||
Optional: reasons (list[str]), next_hint (str). Returns None on parse
|
||||
failure.
|
||||
"""
|
||||
if not text or not text.strip():
|
||||
return None
|
||||
candidates = []
|
||||
fence_match = _FENCE_RE.search(text)
|
||||
if fence_match:
|
||||
candidates.append(fence_match.group(1))
|
||||
candidates.append(text)
|
||||
for body in candidates:
|
||||
body = _strip_comments(body).strip()
|
||||
if not body:
|
||||
continue
|
||||
try:
|
||||
data = json.loads(body)
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
if not isinstance(data, dict):
|
||||
continue
|
||||
if "pass" not in data:
|
||||
continue
|
||||
raw_pass = data.get("pass")
|
||||
if isinstance(raw_pass, str):
|
||||
verdict_pass = raw_pass.strip().lower() == "true"
|
||||
else:
|
||||
verdict_pass = bool(raw_pass)
|
||||
try:
|
||||
score = float(data.get("score", 0.0))
|
||||
except (TypeError, ValueError):
|
||||
score = 0.5
|
||||
if not math.isfinite(score):
|
||||
score = 0.5
|
||||
score = max(0.0, min(1.0, score))
|
||||
verdict = {
|
||||
"pass": verdict_pass,
|
||||
"score": score,
|
||||
}
|
||||
if "reasons" in data and isinstance(data["reasons"], list):
|
||||
verdict["reasons"] = [str(r) for r in data["reasons"]]
|
||||
else:
|
||||
verdict["reasons"] = []
|
||||
if "next_hint" in data and isinstance(data["next_hint"], str):
|
||||
verdict["next_hint"] = data["next_hint"]
|
||||
return verdict
|
||||
return None
|
||||
```
|
||||
|
||||
## D-items (decisions locked for this task)
|
||||
|
||||
- **D-V1**: `"true"` / `"false"` strings → bool via case-insensitive
|
||||
equality with `"true"`. Other strings defer to current `bool(...)` for
|
||||
backwards compat (a verifier emitting `pass: "yes"` keeps current
|
||||
truthy behavior).
|
||||
- **D-V2**: `score` NaN / non-finite → `0.5` (neutral midpoint). This is
|
||||
arbitrary but defensible; documented in CHANGELOG.
|
||||
- **D-V3**: `score` non-numeric string → `0.5` (same neutral default).
|
||||
Documented.
|
||||
- **D-V4**: No CLI flag to opt out of clamping. Strict emitters in `[0,1]`
|
||||
are unaffected; loose emitters get a deterministic value rather than a
|
||||
raw one.
|
||||
- **D-V5**: Tests are pure-functional; no subprocess; no monkeypatch.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No `parse_verdict` rewrite in `status.py`'s `_parse_verdict_status_line`
|
||||
(different function, different concern — parses `VERDICT.md`
|
||||
STATUS:PASS / FAIL strings; not in scope for this task).
|
||||
- No `loop-verifier.md` prompt changes (the prompt still asks for JSON
|
||||
booleans; the runner-side coercion is defense-in-depth). Verifier
|
||||
prompt changes are tracked separately.
|
||||
- No schema change to `.state.loop` `score_history` — existing floats
|
||||
already in `[0,1]` from prior ticks are unaffected; new ticks are
|
||||
clamped.
|
||||
- No `verifier_failed` halt prompt change.
|
||||
|
||||
## Verification
|
||||
|
||||
- `python3 -m py_compile scripts/loop-runner.py`
|
||||
- `python3 -m pytest tests/test_parse_verdict.py -v`
|
||||
- `python3 -m pytest tests/ -q` (full suite stays green; baseline 447 +
|
||||
new)
|
||||
@@ -0,0 +1,55 @@
|
||||
# Referee Verdict: harden-parse-verdict
|
||||
|
||||
## Status: PASS
|
||||
|
||||
## Artifacts reviewed
|
||||
|
||||
- `SPEC.md` — R1-R5 + D-V1 to D-V5; 13-item test plan
|
||||
- `IMPLEMENTATION.md` — files touched, decisions locked, inline bug found and fixed, tests enumerated
|
||||
- `CODE_REVIEW.md` — SPEC coverage table, defensive correctness check, cross-script impact, spot-check, PASS verdict
|
||||
- `BUG_REPORT.md` — O1-O4 observations; all non-blocking; documented behaviors per SPEC
|
||||
- `ADVERSARIAL_BUG_REPORT.md` — A1-A8 sweep; no blockers; documented behaviors per SPEC
|
||||
- `DOC_REVIEW.md` — docs touched: technical.md §7, functional.md §10, CHANGELOG.md; AGENTS/README/prompt intentionally untouched; cross-references confirmed
|
||||
|
||||
## Phase gates satisfied
|
||||
|
||||
| Phase | Artifact |
|
||||
|-------|----------|
|
||||
| research | SPEC.md ✓ |
|
||||
| research:awaiting_approval | approved ✓ |
|
||||
| implement | IMPLEMENTATION.md ✓ |
|
||||
| code_review | CODE_REVIEW.md ✓ |
|
||||
| code_review:awaiting_approval | approved ✓ |
|
||||
| bug_find | BUG_REPORT.md ✓ |
|
||||
| adversarial_bug_find | ADVERSARIAL_BUG_REPORT.md ✓ |
|
||||
| doc_review | DOC_REVIEW.md ✓ |
|
||||
| referee | VERDICT.md (this file) ✓ |
|
||||
|
||||
## Final acceptance criteria
|
||||
|
||||
1. **R1 (pass string coercion)**: ✓ three-way dispatch; `"true"`→True, `"false"`→False, others→`bool(...)`.
|
||||
2. **R2 (score clamp [0,1])**: ✓ `max(0.0, min(1.0, score))`.
|
||||
3. **R3 (non-numeric score → 0.5)**: ✓ `try/except (TypeError, ValueError)`.
|
||||
4. **R4 (backwards compat)**: ✓ strict emitters unaffected; tested.
|
||||
5. **R5 (no new deps)**: ✓ `math` stdlib only.
|
||||
6. **Tests pass**: ✓ 469 passed (447 baseline + 22 new; 0 regressions).
|
||||
7. **Docs in sync**: ✓ technical.md §7 + functional.md §10 + CHANGELOG.md updated.
|
||||
|
||||
## Inline bug found during implementation
|
||||
|
||||
The first iteration of the three-way branch set `verdict_pass = (raw_pass.strip().lower() == "true")`, mapping every non-`"true"` string to False. SPEC R1's `bool(...)` fallback clause was violated (`"yes"` would have regressed from True to False). The implementer caught this via `test_other_truthy_string_pass` before running the full suite, fixed the branch to explicit `if/elif/else: bool(...)`, and the test now guards the contract.
|
||||
|
||||
This is exactly the failure mode the phase pipeline is designed to surface: test-driven discovery of SPEC non-conformance during implement, not after deploy.
|
||||
|
||||
## Adversarial highlights
|
||||
|
||||
- `pass: null`/`[]`/`{}`/`0` → False, `pass: [false]` → True (Python truthy non-empty list). All match v1 `bool(...)` semantics; no regression.
|
||||
- `score: NaN`/`Infinity`/`-Infinity` literals (json.loads accepts) → 0.5 via `math.isfinite`. Confirmed.
|
||||
- `score: "2.0"` (out-of-range numeric string) → clamped to 1.0. The clamp fires after the try/except float() parse, so numeric strings are clamped too. Confirmed.
|
||||
- `score: true` / `score: false` (JSON bool) → 1.0 / 0.0 via `float(True)` / `float(False)`. Documented quirk; downstream plateau gate handles consistently.
|
||||
|
||||
## Verdict
|
||||
|
||||
PASS — task is complete; all artifacts present; all phase gates satisfied; no blockers; no outstanding follow-ups for this task. The score-clamping improvement to plateau detection (BUG_REPORT.md O3) is a positive side effect noted in the CHANGELOG.
|
||||
|
||||
Approve transition to complete.
|
||||
Reference in New Issue
Block a user