Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
This commit is contained in:
@@ -0,0 +1,258 @@
|
||||
# Harden `parse_verdict`
|
||||
|
||||
Small pure-function hardening task: close the
|
||||
`add-loop-runner/BUG_REPORT.md` O6 finding plus the un-noted sibling issue
|
||||
(no score clamping). Both shipped in v1 because the verifier prompt's
|
||||
contract layer was expected to enforce JSON-typed `pass` / numeric `score`
|
||||
in `[0,1]`; observed real LLM responses and the O6 finding show the
|
||||
runner should not trust the prompt contract alone.
|
||||
|
||||
This is a v1.1 hardening task. No CLI surface change; no new feature;
|
||||
no schema migration. Pure robustness inside `scripts/loop-runner.py`'s
|
||||
`parse_verdict`.
|
||||
|
||||
## Goal
|
||||
|
||||
`parse_verdict(text)` currently builds the verdict dict as:
|
||||
|
||||
```python
|
||||
verdict = {
|
||||
"pass": bool(data.get("pass")),
|
||||
"score": float(data.get("score", 0.0)),
|
||||
}
|
||||
```
|
||||
|
||||
Two issues:
|
||||
|
||||
1. **`pass` string coercion bug (O6)**: if the verifier emits
|
||||
`{"pass": "false", "score": 0.1}`, `bool("false")` returns `True`
|
||||
(non-empty string is truthy). The tick records `pass=True`; the
|
||||
`--check-gate` score-plateau brake, the tick log, and the orchestrator
|
||||
downstream all see a "passing" tick when the verifier said "failing".
|
||||
This is a silent correctness bug. Real LLMs do emit JSON booleans most
|
||||
of the time, but OpenAI-grade models occasionally emit `"false"` /
|
||||
`"true"` strings (quote-wrapped). The runner should accept both.
|
||||
|
||||
2. **Unclamped `score`**: if the verifier emits `"score": 1.5` or
|
||||
`"score": -0.3` (out-of-contract), the value is stored as-is. The
|
||||
`score_history` cap and the score-plateau brake assume `[0, 1]`. A
|
||||
`1.5` value inflates the rolling-average computation; a `-0.3`
|
||||
value causes `gate_score_plateau` to compute a negative trend that
|
||||
looks like degradation when none exists. The verifier prompt
|
||||
(`prompts/loop-verifier.md`) declares `score` is a float in `[0, 1]`,
|
||||
but the runner should not rely on prompt-discipline alone.
|
||||
|
||||
## Requirements
|
||||
|
||||
### R1 — Coerce `pass` from string or bool
|
||||
|
||||
`parse_verdict` accepts the following as `data["pass"]`:
|
||||
|
||||
- `true` / `false` (JSON bool) — already correct via `json.loads`.
|
||||
- `"true"` / `"false"` (JSON string) — case-insensitive. The string `"true"`
|
||||
→ `True`; the string `"false"` → `False`. Any other non-empty string
|
||||
→ fall through to the existing `bool(...)` semantics (i.e. truthy).
|
||||
Empty string → `False`.
|
||||
|
||||
Implementation shape:
|
||||
|
||||
```python
|
||||
raw_pass = data.get("pass")
|
||||
if isinstance(raw_pass, str):
|
||||
verdict_pass = raw_pass.strip().lower() == "true"
|
||||
else:
|
||||
verdict_pass = bool(raw_pass)
|
||||
```
|
||||
|
||||
This handles `true`/`false` strings AND retains current behavior for
|
||||
actual booleans (`True`/`False`) AND numbers (`0`/`1` — `bool(0)`
|
||||
returns `False`; current behavior is preserved).
|
||||
|
||||
### R2 — Clamp `score` to `[0, 1]`
|
||||
|
||||
After parsing `float(data.get("score", 0.0))`, clamp:
|
||||
|
||||
```python
|
||||
score = float(data.get("score", 0.0))
|
||||
score = max(0.0, min(1.0, score))
|
||||
```
|
||||
|
||||
NaN handling: `float("nan")` would propagate. If the verifier emits a
|
||||
literal NaN (impossible in strict JSON; some Hermes-style models
|
||||
occasionally emit it via `float('nan')` in reflowed text), the
|
||||
`max/min` comparison returns NaN — both branches preserve NaN, NaN is
|
||||
not equal to NaN, and score-plateau gate would see a constant NaN history.
|
||||
Defensive: reject NaN / non-finite scores by treating them as 0.5 (the
|
||||
verifier emitted something unusable; the midpoint is a neutral default).
|
||||
Use `math.isfinite`:
|
||||
|
||||
```python
|
||||
import math
|
||||
score = float(data.get("score", 0.0))
|
||||
if not math.isfinite(score):
|
||||
score = 0.5
|
||||
score = max(0.0, min(1.0, score))
|
||||
```
|
||||
|
||||
### R3 — Defensive non-numeric `score`
|
||||
|
||||
If `data.get("score")` is a string like `"0.8"`, `float(...)` already
|
||||
handles it (Python's `float` accepts string numerics). If it's a
|
||||
non-numeric string, `float(...)` raises `ValueError`. Current code doesn't
|
||||
catch this; would propagate as an unhandled exception mid-tick (→ halt via
|
||||
the runner's finally → lock releases → tick log shows a HALT but errors
|
||||
aren't categorized as `verifier_failed`). Wrap the float conversion:
|
||||
|
||||
```python
|
||||
try:
|
||||
score = float(data.get("score", 0.0))
|
||||
except (TypeError, ValueError):
|
||||
score = 0.5
|
||||
```
|
||||
|
||||
### R4 — Backwards compat
|
||||
|
||||
- Verdict dict shape is unchanged: `{"pass": bool, "score": float,
|
||||
"reasons": list[str], "next_hint": str}`. Existing callers (`cmd_tick`,
|
||||
`_gate_score_plateau` indirectly via `score_history`, tick-log line
|
||||
format) are unaffected.
|
||||
- Strict-JSON emitters (true booleans, numeric scores in `[0,1]`) get
|
||||
identical results to current behavior.
|
||||
- The `score = 0.5` defaults for NaN / non-numeric are new behavior;
|
||||
documented in the CHANGELOG.
|
||||
|
||||
### R5 — No new pip deps
|
||||
|
||||
`math` is stdlib.
|
||||
|
||||
## Test plan (`tests/test_parse_verdict.py`)
|
||||
|
||||
New file. Pure unit tests against `parse_verdict`; no subprocess, no
|
||||
fixtures, no tmp_path needed. All use the function directly with literal
|
||||
input strings.
|
||||
|
||||
1. `test_pass_true_bool` — `{"pass": true, "score": 0.8}` →
|
||||
`pass=True, score=0.8`.
|
||||
2. `test_pass_false_bool` — `{"pass": false, "score": 0.2}` →
|
||||
`pass=False, score=0.2`.
|
||||
3. `test_pass_true_string` — `{"pass": "true", "score": 0.9}` →
|
||||
`pass=True, score=0.9` (the O6 bug).
|
||||
4. `test_pass_false_string` — `{"pass": "false", "score": 0.1}` →
|
||||
`pass=False, score=0.1` (the O6 bug).
|
||||
5. `test_pass_string_case_insensitive` — `{"pass": "FALSE", "score": 0.1}`
|
||||
→ `pass=False`. `{"pass": "True", "score": 0.9}` → `pass=True`.
|
||||
6. `test_score_clamped_high` — `{"pass": true, "score": 1.5}` →
|
||||
`pass=True, score=1.0`.
|
||||
7. `test_score_clamped_low` — `{"pass": true, "score": -0.3}` →
|
||||
`pass=True, score=0.0`.
|
||||
8. `test_score_nan_to_neutral` — `{"pass": true, "score": NaN}` →
|
||||
`pass=True, score=0.5`. (Use `float("nan")` literal in the test JSON
|
||||
`__import__('math').nan` — actually use the string `"NaN"` to
|
||||
simulate.
|
||||
9. `test_score_non_numeric_string` — `{"pass": true, "score": "great"}`
|
||||
→ `pass=True, score=0.5`.
|
||||
10. `test_score_numeric_string_ok` — `{"pass": true, "score": "0.75"}`
|
||||
→ `pass=True, score=0.75` (Python's float() already handles this; no
|
||||
regression).
|
||||
11. `test_empty_pass_string` — `{"pass": "", "score": 0.5}` →
|
||||
`pass=False` (per R1's `bool(raw_pass)` fallback for non-`true`/`false`
|
||||
strings: empty string → `bool("")` → `False`).
|
||||
12. `test_other_truthy_string_pass` — `{"pass": "yes", "score": 0.5}` →
|
||||
`pass=True` (`"yes"` is non-empty, non-`true` → `bool("yes")` is `True`).
|
||||
Backwards-compat with prior semantics.
|
||||
13. `test_existing_fence_block_behavior` — ```` ```json {"pass": true,
|
||||
"score": 0.8} ``` ```` → still parses; new clamp/coerce don't break
|
||||
the fence-extractor path.
|
||||
|
||||
## Concrete code shape
|
||||
|
||||
```python
|
||||
import math
|
||||
|
||||
def parse_verdict(text: str) -> Optional[dict]:
|
||||
"""Parse verifier JSON verdict. Accepts raw, fenced, or commented JSON.
|
||||
|
||||
Required keys: pass (bool — also accepts "true"/"false" strings),
|
||||
score (float — clamped to [0, 1]; NaN/non-finite defaults to 0.5).
|
||||
Optional: reasons (list[str]), next_hint (str). Returns None on parse
|
||||
failure.
|
||||
"""
|
||||
if not text or not text.strip():
|
||||
return None
|
||||
candidates = []
|
||||
fence_match = _FENCE_RE.search(text)
|
||||
if fence_match:
|
||||
candidates.append(fence_match.group(1))
|
||||
candidates.append(text)
|
||||
for body in candidates:
|
||||
body = _strip_comments(body).strip()
|
||||
if not body:
|
||||
continue
|
||||
try:
|
||||
data = json.loads(body)
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
if not isinstance(data, dict):
|
||||
continue
|
||||
if "pass" not in data:
|
||||
continue
|
||||
raw_pass = data.get("pass")
|
||||
if isinstance(raw_pass, str):
|
||||
verdict_pass = raw_pass.strip().lower() == "true"
|
||||
else:
|
||||
verdict_pass = bool(raw_pass)
|
||||
try:
|
||||
score = float(data.get("score", 0.0))
|
||||
except (TypeError, ValueError):
|
||||
score = 0.5
|
||||
if not math.isfinite(score):
|
||||
score = 0.5
|
||||
score = max(0.0, min(1.0, score))
|
||||
verdict = {
|
||||
"pass": verdict_pass,
|
||||
"score": score,
|
||||
}
|
||||
if "reasons" in data and isinstance(data["reasons"], list):
|
||||
verdict["reasons"] = [str(r) for r in data["reasons"]]
|
||||
else:
|
||||
verdict["reasons"] = []
|
||||
if "next_hint" in data and isinstance(data["next_hint"], str):
|
||||
verdict["next_hint"] = data["next_hint"]
|
||||
return verdict
|
||||
return None
|
||||
```
|
||||
|
||||
## D-items (decisions locked for this task)
|
||||
|
||||
- **D-V1**: `"true"` / `"false"` strings → bool via case-insensitive
|
||||
equality with `"true"`. Other strings defer to current `bool(...)` for
|
||||
backwards compat (a verifier emitting `pass: "yes"` keeps current
|
||||
truthy behavior).
|
||||
- **D-V2**: `score` NaN / non-finite → `0.5` (neutral midpoint). This is
|
||||
arbitrary but defensible; documented in CHANGELOG.
|
||||
- **D-V3**: `score` non-numeric string → `0.5` (same neutral default).
|
||||
Documented.
|
||||
- **D-V4**: No CLI flag to opt out of clamping. Strict emitters in `[0,1]`
|
||||
are unaffected; loose emitters get a deterministic value rather than a
|
||||
raw one.
|
||||
- **D-V5**: Tests are pure-functional; no subprocess; no monkeypatch.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No `parse_verdict` rewrite in `status.py`'s `_parse_verdict_status_line`
|
||||
(different function, different concern — parses `VERDICT.md`
|
||||
STATUS:PASS / FAIL strings; not in scope for this task).
|
||||
- No `loop-verifier.md` prompt changes (the prompt still asks for JSON
|
||||
booleans; the runner-side coercion is defense-in-depth). Verifier
|
||||
prompt changes are tracked separately.
|
||||
- No schema change to `.state.loop` `score_history` — existing floats
|
||||
already in `[0,1]` from prior ticks are unaffected; new ticks are
|
||||
clamped.
|
||||
- No `verifier_failed` halt prompt change.
|
||||
|
||||
## Verification
|
||||
|
||||
- `python3 -m py_compile scripts/loop-runner.py`
|
||||
- `python3 -m pytest tests/test_parse_verdict.py -v`
|
||||
- `python3 -m pytest tests/ -q` (full suite stays green; baseline 447 +
|
||||
new)
|
||||
Reference in New Issue
Block a user