Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
This commit is contained in:
@@ -0,0 +1,55 @@
|
||||
# Referee Verdict: harden-parse-verdict
|
||||
|
||||
## Status: PASS
|
||||
|
||||
## Artifacts reviewed
|
||||
|
||||
- `SPEC.md` — R1-R5 + D-V1 to D-V5; 13-item test plan
|
||||
- `IMPLEMENTATION.md` — files touched, decisions locked, inline bug found and fixed, tests enumerated
|
||||
- `CODE_REVIEW.md` — SPEC coverage table, defensive correctness check, cross-script impact, spot-check, PASS verdict
|
||||
- `BUG_REPORT.md` — O1-O4 observations; all non-blocking; documented behaviors per SPEC
|
||||
- `ADVERSARIAL_BUG_REPORT.md` — A1-A8 sweep; no blockers; documented behaviors per SPEC
|
||||
- `DOC_REVIEW.md` — docs touched: technical.md §7, functional.md §10, CHANGELOG.md; AGENTS/README/prompt intentionally untouched; cross-references confirmed
|
||||
|
||||
## Phase gates satisfied
|
||||
|
||||
| Phase | Artifact |
|
||||
|-------|----------|
|
||||
| research | SPEC.md ✓ |
|
||||
| research:awaiting_approval | approved ✓ |
|
||||
| implement | IMPLEMENTATION.md ✓ |
|
||||
| code_review | CODE_REVIEW.md ✓ |
|
||||
| code_review:awaiting_approval | approved ✓ |
|
||||
| bug_find | BUG_REPORT.md ✓ |
|
||||
| adversarial_bug_find | ADVERSARIAL_BUG_REPORT.md ✓ |
|
||||
| doc_review | DOC_REVIEW.md ✓ |
|
||||
| referee | VERDICT.md (this file) ✓ |
|
||||
|
||||
## Final acceptance criteria
|
||||
|
||||
1. **R1 (pass string coercion)**: ✓ three-way dispatch; `"true"`→True, `"false"`→False, others→`bool(...)`.
|
||||
2. **R2 (score clamp [0,1])**: ✓ `max(0.0, min(1.0, score))`.
|
||||
3. **R3 (non-numeric score → 0.5)**: ✓ `try/except (TypeError, ValueError)`.
|
||||
4. **R4 (backwards compat)**: ✓ strict emitters unaffected; tested.
|
||||
5. **R5 (no new deps)**: ✓ `math` stdlib only.
|
||||
6. **Tests pass**: ✓ 469 passed (447 baseline + 22 new; 0 regressions).
|
||||
7. **Docs in sync**: ✓ technical.md §7 + functional.md §10 + CHANGELOG.md updated.
|
||||
|
||||
## Inline bug found during implementation
|
||||
|
||||
The first iteration of the three-way branch set `verdict_pass = (raw_pass.strip().lower() == "true")`, mapping every non-`"true"` string to False. SPEC R1's `bool(...)` fallback clause was violated (`"yes"` would have regressed from True to False). The implementer caught this via `test_other_truthy_string_pass` before running the full suite, fixed the branch to explicit `if/elif/else: bool(...)`, and the test now guards the contract.
|
||||
|
||||
This is exactly the failure mode the phase pipeline is designed to surface: test-driven discovery of SPEC non-conformance during implement, not after deploy.
|
||||
|
||||
## Adversarial highlights
|
||||
|
||||
- `pass: null`/`[]`/`{}`/`0` → False, `pass: [false]` → True (Python truthy non-empty list). All match v1 `bool(...)` semantics; no regression.
|
||||
- `score: NaN`/`Infinity`/`-Infinity` literals (json.loads accepts) → 0.5 via `math.isfinite`. Confirmed.
|
||||
- `score: "2.0"` (out-of-range numeric string) → clamped to 1.0. The clamp fires after the try/except float() parse, so numeric strings are clamped too. Confirmed.
|
||||
- `score: true` / `score: false` (JSON bool) → 1.0 / 0.0 via `float(True)` / `float(False)`. Documented quirk; downstream plateau gate handles consistently.
|
||||
|
||||
## Verdict
|
||||
|
||||
PASS — task is complete; all artifacts present; all phase gates satisfied; no blockers; no outstanding follow-ups for this task. The score-clamping improvement to plateau detection (BUG_REPORT.md O3) is a positive side effect noted in the CHANGELOG.
|
||||
|
||||
Approve transition to complete.
|
||||
Reference in New Issue
Block a user