Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test

- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
This commit is contained in:
Lap Tran
2026-06-26 10:05:18 -04:00
parent fe43b9e1fc
commit bc7daf8590
666 changed files with 15994 additions and 69 deletions
+258
View File
@@ -0,0 +1,258 @@
# Harden `parse_verdict`
Small pure-function hardening task: close the
`add-loop-runner/BUG_REPORT.md` O6 finding plus the un-noted sibling issue
(no score clamping). Both shipped in v1 because the verifier prompt's
contract layer was expected to enforce JSON-typed `pass` / numeric `score`
in `[0,1]`; observed real LLM responses and the O6 finding show the
runner should not trust the prompt contract alone.
This is a v1.1 hardening task. No CLI surface change; no new feature;
no schema migration. Pure robustness inside `scripts/loop-runner.py`'s
`parse_verdict`.
## Goal
`parse_verdict(text)` currently builds the verdict dict as:
```python
verdict = {
"pass": bool(data.get("pass")),
"score": float(data.get("score", 0.0)),
}
```
Two issues:
1. **`pass` string coercion bug (O6)**: if the verifier emits
`{"pass": "false", "score": 0.1}`, `bool("false")` returns `True`
(non-empty string is truthy). The tick records `pass=True`; the
`--check-gate` score-plateau brake, the tick log, and the orchestrator
downstream all see a "passing" tick when the verifier said "failing".
This is a silent correctness bug. Real LLMs do emit JSON booleans most
of the time, but OpenAI-grade models occasionally emit `"false"` /
`"true"` strings (quote-wrapped). The runner should accept both.
2. **Unclamped `score`**: if the verifier emits `"score": 1.5` or
`"score": -0.3` (out-of-contract), the value is stored as-is. The
`score_history` cap and the score-plateau brake assume `[0, 1]`. A
`1.5` value inflates the rolling-average computation; a `-0.3`
value causes `gate_score_plateau` to compute a negative trend that
looks like degradation when none exists. The verifier prompt
(`prompts/loop-verifier.md`) declares `score` is a float in `[0, 1]`,
but the runner should not rely on prompt-discipline alone.
## Requirements
### R1 — Coerce `pass` from string or bool
`parse_verdict` accepts the following as `data["pass"]`:
- `true` / `false` (JSON bool) — already correct via `json.loads`.
- `"true"` / `"false"` (JSON string) — case-insensitive. The string `"true"`
→ `True`; the string `"false"` → `False`. Any other non-empty string
→ fall through to the existing `bool(...)` semantics (i.e. truthy).
Empty string → `False`.
Implementation shape:
```python
raw_pass = data.get("pass")
if isinstance(raw_pass, str):
verdict_pass = raw_pass.strip().lower() == "true"
else:
verdict_pass = bool(raw_pass)
```
This handles `true`/`false` strings AND retains current behavior for
actual booleans (`True`/`False`) AND numbers (`0`/`1` — `bool(0)`
returns `False`; current behavior is preserved).
### R2 — Clamp `score` to `[0, 1]`
After parsing `float(data.get("score", 0.0))`, clamp:
```python
score = float(data.get("score", 0.0))
score = max(0.0, min(1.0, score))
```
NaN handling: `float("nan")` would propagate. If the verifier emits a
literal NaN (impossible in strict JSON; some Hermes-style models
occasionally emit it via `float('nan')` in reflowed text), the
`max/min` comparison returns NaN — both branches preserve NaN, NaN is
not equal to NaN, and score-plateau gate would see a constant NaN history.
Defensive: reject NaN / non-finite scores by treating them as 0.5 (the
verifier emitted something unusable; the midpoint is a neutral default).
Use `math.isfinite`:
```python
import math
score = float(data.get("score", 0.0))
if not math.isfinite(score):
score = 0.5
score = max(0.0, min(1.0, score))
```
### R3 — Defensive non-numeric `score`
If `data.get("score")` is a string like `"0.8"`, `float(...)` already
handles it (Python's `float` accepts string numerics). If it's a
non-numeric string, `float(...)` raises `ValueError`. Current code doesn't
catch this; would propagate as an unhandled exception mid-tick (→ halt via
the runner's finally → lock releases → tick log shows a HALT but errors
aren't categorized as `verifier_failed`). Wrap the float conversion:
```python
try:
score = float(data.get("score", 0.0))
except (TypeError, ValueError):
score = 0.5
```
### R4 — Backwards compat
- Verdict dict shape is unchanged: `{"pass": bool, "score": float,
"reasons": list[str], "next_hint": str}`. Existing callers (`cmd_tick`,
`_gate_score_plateau` indirectly via `score_history`, tick-log line
format) are unaffected.
- Strict-JSON emitters (true booleans, numeric scores in `[0,1]`) get
identical results to current behavior.
- The `score = 0.5` defaults for NaN / non-numeric are new behavior;
documented in the CHANGELOG.
### R5 — No new pip deps
`math` is stdlib.
## Test plan (`tests/test_parse_verdict.py`)
New file. Pure unit tests against `parse_verdict`; no subprocess, no
fixtures, no tmp_path needed. All use the function directly with literal
input strings.
1. `test_pass_true_bool` — `{"pass": true, "score": 0.8}` →
`pass=True, score=0.8`.
2. `test_pass_false_bool` — `{"pass": false, "score": 0.2}` →
`pass=False, score=0.2`.
3. `test_pass_true_string` — `{"pass": "true", "score": 0.9}` →
`pass=True, score=0.9` (the O6 bug).
4. `test_pass_false_string` — `{"pass": "false", "score": 0.1}` →
`pass=False, score=0.1` (the O6 bug).
5. `test_pass_string_case_insensitive` — `{"pass": "FALSE", "score": 0.1}`
→ `pass=False`. `{"pass": "True", "score": 0.9}` → `pass=True`.
6. `test_score_clamped_high` — `{"pass": true, "score": 1.5}` →
`pass=True, score=1.0`.
7. `test_score_clamped_low` — `{"pass": true, "score": -0.3}` →
`pass=True, score=0.0`.
8. `test_score_nan_to_neutral` — `{"pass": true, "score": NaN}` →
`pass=True, score=0.5`. (Use `float("nan")` literal in the test JSON
`__import__('math').nan` — actually use the string `"NaN"` to
simulate.
9. `test_score_non_numeric_string` — `{"pass": true, "score": "great"}`
→ `pass=True, score=0.5`.
10. `test_score_numeric_string_ok` — `{"pass": true, "score": "0.75"}`
→ `pass=True, score=0.75` (Python's float() already handles this; no
regression).
11. `test_empty_pass_string` — `{"pass": "", "score": 0.5}` →
`pass=False` (per R1's `bool(raw_pass)` fallback for non-`true`/`false`
strings: empty string → `bool("")` → `False`).
12. `test_other_truthy_string_pass` — `{"pass": "yes", "score": 0.5}` →
`pass=True` (`"yes"` is non-empty, non-`true` → `bool("yes")` is `True`).
Backwards-compat with prior semantics.
13. `test_existing_fence_block_behavior` — ```` ```json {"pass": true,
"score": 0.8} ``` ```` → still parses; new clamp/coerce don't break
the fence-extractor path.
## Concrete code shape
```python
import math
def parse_verdict(text: str) -> Optional[dict]:
"""Parse verifier JSON verdict. Accepts raw, fenced, or commented JSON.
Required keys: pass (bool — also accepts "true"/"false" strings),
score (float — clamped to [0, 1]; NaN/non-finite defaults to 0.5).
Optional: reasons (list[str]), next_hint (str). Returns None on parse
failure.
"""
if not text or not text.strip():
return None
candidates = []
fence_match = _FENCE_RE.search(text)
if fence_match:
candidates.append(fence_match.group(1))
candidates.append(text)
for body in candidates:
body = _strip_comments(body).strip()
if not body:
continue
try:
data = json.loads(body)
except json.JSONDecodeError:
continue
if not isinstance(data, dict):
continue
if "pass" not in data:
continue
raw_pass = data.get("pass")
if isinstance(raw_pass, str):
verdict_pass = raw_pass.strip().lower() == "true"
else:
verdict_pass = bool(raw_pass)
try:
score = float(data.get("score", 0.0))
except (TypeError, ValueError):
score = 0.5
if not math.isfinite(score):
score = 0.5
score = max(0.0, min(1.0, score))
verdict = {
"pass": verdict_pass,
"score": score,
}
if "reasons" in data and isinstance(data["reasons"], list):
verdict["reasons"] = [str(r) for r in data["reasons"]]
else:
verdict["reasons"] = []
if "next_hint" in data and isinstance(data["next_hint"], str):
verdict["next_hint"] = data["next_hint"]
return verdict
return None
```
## D-items (decisions locked for this task)
- **D-V1**: `"true"` / `"false"` strings → bool via case-insensitive
equality with `"true"`. Other strings defer to current `bool(...)` for
backwards compat (a verifier emitting `pass: "yes"` keeps current
truthy behavior).
- **D-V2**: `score` NaN / non-finite → `0.5` (neutral midpoint). This is
arbitrary but defensible; documented in CHANGELOG.
- **D-V3**: `score` non-numeric string → `0.5` (same neutral default).
Documented.
- **D-V4**: No CLI flag to opt out of clamping. Strict emitters in `[0,1]`
are unaffected; loose emitters get a deterministic value rather than a
raw one.
- **D-V5**: Tests are pure-functional; no subprocess; no monkeypatch.
## Non-goals
- No `parse_verdict` rewrite in `status.py`'s `_parse_verdict_status_line`
(different function, different concern — parses `VERDICT.md`
STATUS:PASS / FAIL strings; not in scope for this task).
- No `loop-verifier.md` prompt changes (the prompt still asks for JSON
booleans; the runner-side coercion is defense-in-depth). Verifier
prompt changes are tracked separately.
- No schema change to `.state.loop` `score_history` — existing floats
already in `[0,1]` from prior ticks are unaffected; new ticks are
clamped.
- No `verifier_failed` halt prompt change.
## Verification
- `python3 -m py_compile scripts/loop-runner.py`
- `python3 -m pytest tests/test_parse_verdict.py -v`
- `python3 -m pytest tests/ -q` (full suite stays green; baseline 447 +
new)