Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

3.2 KiB

Referee Verdict: harden-parse-verdict

Status: PASS

Artifacts reviewed

  • SPEC.md — R1-R5 + D-V1 to D-V5; 13-item test plan
  • IMPLEMENTATION.md — files touched, decisions locked, inline bug found and fixed, tests enumerated
  • CODE_REVIEW.md — SPEC coverage table, defensive correctness check, cross-script impact, spot-check, PASS verdict
  • BUG_REPORT.md — O1-O4 observations; all non-blocking; documented behaviors per SPEC
  • ADVERSARIAL_BUG_REPORT.md — A1-A8 sweep; no blockers; documented behaviors per SPEC
  • DOC_REVIEW.md — docs touched: technical.md §7, functional.md §10, CHANGELOG.md; AGENTS/README/prompt intentionally untouched; cross-references confirmed

Phase gates satisfied

Phase Artifact
research SPEC.md ✓
research:awaiting_approval approved ✓
implement IMPLEMENTATION.md ✓
code_review CODE_REVIEW.md ✓
code_review:awaiting_approval approved ✓
bug_find BUG_REPORT.md ✓
adversarial_bug_find ADVERSARIAL_BUG_REPORT.md ✓
doc_review DOC_REVIEW.md ✓
referee VERDICT.md (this file) ✓

Final acceptance criteria

  1. R1 (pass string coercion): ✓ three-way dispatch; "true"→True, "false"→False, others→bool(...).
  2. R2 (score clamp [0,1]): ✓ max(0.0, min(1.0, score)).
  3. R3 (non-numeric score → 0.5): ✓ try/except (TypeError, ValueError).
  4. R4 (backwards compat): ✓ strict emitters unaffected; tested.
  5. R5 (no new deps): ✓ math stdlib only.
  6. Tests pass: ✓ 469 passed (447 baseline + 22 new; 0 regressions).
  7. Docs in sync: ✓ technical.md §7 + functional.md §10 + CHANGELOG.md updated.

Inline bug found during implementation

The first iteration of the three-way branch set verdict_pass = (raw_pass.strip().lower() == "true"), mapping every non-"true" string to False. SPEC R1's bool(...) fallback clause was violated ("yes" would have regressed from True to False). The implementer caught this via test_other_truthy_string_pass before running the full suite, fixed the branch to explicit if/elif/else: bool(...), and the test now guards the contract.

This is exactly the failure mode the phase pipeline is designed to surface: test-driven discovery of SPEC non-conformance during implement, not after deploy.

Adversarial highlights

  • pass: null/[]/{}/0 → False, pass: [false] → True (Python truthy non-empty list). All match v1 bool(...) semantics; no regression.
  • score: NaN/Infinity/-Infinity literals (json.loads accepts) → 0.5 via math.isfinite. Confirmed.
  • score: "2.0" (out-of-range numeric string) → clamped to 1.0. The clamp fires after the try/except float() parse, so numeric strings are clamped too. Confirmed.
  • score: true / score: false (JSON bool) → 1.0 / 0.0 via float(True) / float(False). Documented quirk; downstream plateau gate handles consistently.

Verdict

PASS — task is complete; all artifacts present; all phase gates satisfied; no blockers; no outstanding follow-ups for this task. The score-clamping improvement to plateau detection (BUG_REPORT.md O3) is a positive side effect noted in the CHANGELOG.

Approve transition to complete.