- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
67 lines
3.3 KiB
Markdown
67 lines
3.3 KiB
Markdown
# Code Review: harden-parse-verdict
|
|
|
|
## SPEC coverage
|
|
|
|
| Requirement | Status |
|
|
|-------------|--------|
|
|
| R1 — Coerce `pass` from string or bool (3-way dispatch: true / false / other) | ✓ — three-way branch with case-insensitive match + `bool(raw_pass.strip())` fallback |
|
|
| R2 — Clamp `score` to `[0, 1]` via `max(0, min(1, x))` | ✓ |
|
|
| R3 — Defensive non-numeric `score` (try/except TypeError, ValueError) | ✓ — caught and defaulted to 0.5 |
|
|
| R4 — Backwards compat (dict shape unchanged; strict emitters unaffected) | ✓ — verdict still has `pass`, `score`, `reasons`, `next_hint` keys; strict JSON emitters get identical results to v1 |
|
|
| R5 — No new pip deps (`math` stdlib) | ✓ |
|
|
|
|
## Code readability
|
|
|
|
- Three-way branch is more verbose than the v1 single-line `bool(...)` but
|
|
the case-intent is clearer: the comment "case-insensitive. The string
|
|
'true' → True; the string 'false' → False. Any other non-empty string
|
|
→ fall through to the existing `bool(...)` semantics" in the SPEC is
|
|
preserved exactly by the if/elif/else.
|
|
- The score-clamp block is two statements (try/except, then isfinite
|
|
check, then clamp). The order matters: the TypeError/ValueError from
|
|
`float(None)` or `float("great")` must be caught BEFORE the `math.isfinite`
|
|
call; else `math.isfinite(None)` raises TypeError uncaught. The order
|
|
in the implementation is correct (try/except wraps the float call;
|
|
isfinite only sees a finite-or-NaN float).
|
|
|
|
## Defensive correctness check
|
|
|
|
- `bool(None)` → False (if `data` is `{"pass": None}`; treated as no-pass
|
|
→ False; matches pre-fix `bool(None)` = False; no regression).
|
|
- `bool(0)` → False (if verifier emits `"pass": 0`); preserved.
|
|
- `bool(1)` → True; `bool([])` False; `bool({})` False; all preserved — no
|
|
regression for non-string types.
|
|
- String `" tRuE "` strips via `.strip().lower()` → "true" → True.
|
|
- String `"\nfalse"` strips → "false" → False. Edge case covered.
|
|
|
|
## Cross-script impact
|
|
|
|
- `parse_verdict` is local to `loop-runner.py`; not duplicated to
|
|
`status.py`. The change is contained.
|
|
- `_gate_score_plateau` in `status.py` consumes `score_history` (with
|
|
clamped values via the runner's atomic write of `last_verdict`) —
|
|
already assumed `[0, 1]`. The clamp guarantees it.
|
|
- `_write_state_loop` timestamps store JSON; clamped scores
|
|
round-trip cleanly (no serialization loss).
|
|
|
|
## Tests spot-check
|
|
|
|
- `test_other_truthy_string_pass` (the bug found inline): verifies that
|
|
the SPEC R1's `bool(...)` fallback clause is honored — a `"yes"` string
|
|
yields `True`. Pre-fix v1 behavior preserved.
|
|
- `test_pass_with_surrounding_whitespace`: covers an edge case (`" true "`)
|
|
the SPEC didn't explicitly enumerate but is sensible.
|
|
- `test_score_none_value_to_neutral`: `data.get("score", 0.0)` returns
|
|
`None` (key exists with None value); `float(None)` raises TypeError →
|
|
caught → 0.5. Not in SPEC's explicit test plan but is a natural
|
|
consequence of R3's TypeError coverage. Good defensive test.
|
|
- `test_score_infinity_to_neutral`: covers `math.isfinite(Infinity)` →
|
|
False path. Added after `test_score_nan_to_neutral`; both prove the
|
|
isfinite check.
|
|
- All 22 tests pass.
|
|
|
|
## Verdict
|
|
|
|
PASS — implementer followed SPEC; inline bug found and fixed during
|
|
test; the fix matches SPEC R1's three-way dispatch wording exactly.
|
|
Proceed to bug_find. |