Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test

- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
This commit is contained in:
Lap Tran
2026-06-26 10:05:18 -04:00
parent fe43b9e1fc
commit bc7daf8590
666 changed files with 15994 additions and 69 deletions
+55
View File
@@ -0,0 +1,55 @@
# Referee Verdict: harden-parse-verdict
## Status: PASS
## Artifacts reviewed
- `SPEC.md` — R1-R5 + D-V1 to D-V5; 13-item test plan
- `IMPLEMENTATION.md` — files touched, decisions locked, inline bug found and fixed, tests enumerated
- `CODE_REVIEW.md` — SPEC coverage table, defensive correctness check, cross-script impact, spot-check, PASS verdict
- `BUG_REPORT.md` — O1-O4 observations; all non-blocking; documented behaviors per SPEC
- `ADVERSARIAL_BUG_REPORT.md` — A1-A8 sweep; no blockers; documented behaviors per SPEC
- `DOC_REVIEW.md` — docs touched: technical.md §7, functional.md §10, CHANGELOG.md; AGENTS/README/prompt intentionally untouched; cross-references confirmed
## Phase gates satisfied
| Phase | Artifact |
|-------|----------|
| research | SPEC.md ✓ |
| research:awaiting_approval | approved ✓ |
| implement | IMPLEMENTATION.md ✓ |
| code_review | CODE_REVIEW.md ✓ |
| code_review:awaiting_approval | approved ✓ |
| bug_find | BUG_REPORT.md ✓ |
| adversarial_bug_find | ADVERSARIAL_BUG_REPORT.md ✓ |
| doc_review | DOC_REVIEW.md ✓ |
| referee | VERDICT.md (this file) ✓ |
## Final acceptance criteria
1. **R1 (pass string coercion)**: ✓ three-way dispatch; `"true"`→True, `"false"`→False, others→`bool(...)`.
2. **R2 (score clamp [0,1])**: ✓ `max(0.0, min(1.0, score))`.
3. **R3 (non-numeric score → 0.5)**: ✓ `try/except (TypeError, ValueError)`.
4. **R4 (backwards compat)**: ✓ strict emitters unaffected; tested.
5. **R5 (no new deps)**: ✓ `math` stdlib only.
6. **Tests pass**: ✓ 469 passed (447 baseline + 22 new; 0 regressions).
7. **Docs in sync**: ✓ technical.md §7 + functional.md §10 + CHANGELOG.md updated.
## Inline bug found during implementation
The first iteration of the three-way branch set `verdict_pass = (raw_pass.strip().lower() == "true")`, mapping every non-`"true"` string to False. SPEC R1's `bool(...)` fallback clause was violated (`"yes"` would have regressed from True to False). The implementer caught this via `test_other_truthy_string_pass` before running the full suite, fixed the branch to explicit `if/elif/else: bool(...)`, and the test now guards the contract.
This is exactly the failure mode the phase pipeline is designed to surface: test-driven discovery of SPEC non-conformance during implement, not after deploy.
## Adversarial highlights
- `pass: null`/`[]`/`{}`/`0` → False, `pass: [false]` → True (Python truthy non-empty list). All match v1 `bool(...)` semantics; no regression.
- `score: NaN`/`Infinity`/`-Infinity` literals (json.loads accepts) → 0.5 via `math.isfinite`. Confirmed.
- `score: "2.0"` (out-of-range numeric string) → clamped to 1.0. The clamp fires after the try/except float() parse, so numeric strings are clamped too. Confirmed.
- `score: true` / `score: false` (JSON bool) → 1.0 / 0.0 via `float(True)` / `float(False)`. Documented quirk; downstream plateau gate handles consistently.
## Verdict
PASS — task is complete; all artifacts present; all phase gates satisfied; no blockers; no outstanding follow-ups for this task. The score-clamping improvement to plateau detection (BUG_REPORT.md O3) is a positive side effect noted in the CHANGELOG.
Approve transition to complete.