Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

4.2 KiB

Fix Verdict Parsing and State Machine Alignment

Goal

Fix the critical verdict-parsing bug that causes PASS verdicts to be falsely classified as BLOCKED, and align the dashboard's determine_task_state() with the orchestrator's state machine specification.

Requirements

automaton/dashboard/core/task.py:159-169 currently uses substring search for FAIL/NEEDS_REVIEW/PASS. This means a PASS verdict that mentions a previous failure (which referee.md explicitly requires when comparing bug finder outputs) gets misclassified as BLOCKED.

Fix: Parse the actual status line (## Status: PASS, **Status**: FAIL, etc.) extracted from the verdict content, falling back to substring search only when no structured status line is found.

R2. Fix verdict check ordering

The current code checks FAIL/NEEDS_REVIEW substrings before PASS. A correctly parsed status line makes this irrelevant for structured verdicts — only fall back to substring search for unstructured verdicts, using the same check order (check FAIL/NEEDS_REVIEW first, then PASS) but document the limitation.

R3. Use same parsing in parse_sub_tasks

task.py:215-220 has the same substring-search issue for sub-task verdicts. Apply the same fix.

R4. Align determine_task_state() with orchestrate.md state machine

Four concrete divergences between orchestrate.md:266-282 and task.py:135-198:

Orchestrator says Dashboard does Fix
IMPLEMENTATION.md → Bug Find IMPLEMENTATION.md → Implement Match orchestrator: show Bug Find when IMPLEMENTATION.md exists but no BUG_REPORT.md or ADVERSARIAL_BUG_REPORT.md
BUG_REPORT.md + SPEC.md (no ADV) → Adversarial Bug Find BUG_REPORT.md alone → Bug Find Match orchestrator: BUG_REPORT.md → Bug Find, ADVERSARIAL_BUG_REPORT.md alone → Adversarial Bug Find. When both exist, advance to Doc Review or Referee.
ADVERSARIAL_BUG_REPORT.md alone → not specified ADVERSARIAL_BUG_REPORT.md alone → ADV_BUG_FIND Follow orchestrator's intent: a lone ADVERSARIAL_BUG_REPORT without BUG_REPORT technically doesn't reach Adversarial Bug Find per spec. Treat ADV alone same as BUG alone for the dashboard (Bug Find).
SPEC.md alone → Design or Implement SPEC.md alone → Research Keep dashboard behavior. The orchestrator spec says "Design or Implement" meaning those are the next steps the orchestrator would drive. The dashboard should show the task in its current state (Research). No change needed.

R5. Update parse_sub_tasks to match the same logic

Sub-task state determination uses the same function, so these fixes propagate automatically. Verify that sub-tasks with only PARENT_SPEC.md or VRAM_CONFIG.md correctly show as BACKLOG.

R6. Document the minimal verdict schema

Add a note in prompts/referee.md requiring that VERDICT.md include ## Status: PASS / ## Status: FAIL / ## Status: NEEDS_REVIEW as a structured machine-parseable field. The dashboard relies on this for correct classification.

Acceptance Criteria

  • ## Status: PASS verdict mentioning the word "FAIL" in findings → DONE (not BLOCKED)
  • ## Status: PASS verdict mentioning "NEEDS_REVIEW" in body → DONE (not BLOCKED)
  • ## Status: FAIL verdict → BLOCKED
  • ## Status: NEEDS_REVIEW verdict → BLOCKED
  • IMPLEMENTATION.md alone (no BUG_REPORT, no ADVERSARIAL_BUG_REPORT) → BUG_FIND (not IMPLEMENT)
  • BUG_REPORT.md + SPEC.md (no ADVERSARIAL_BUG_REPORT) → BUG_FIND
  • ADVERSARIAL_BUG_REPORT.md + BUG_REPORT.md + SPEC.md → ADV_BUG_FIND (or higher if DOC_REVIEW/VERDICT present)
  • SPEC.md alone → RESEARCH (unchanged, confirmed as correct)
  • Existing tests in tests/test_task.py still pass
  • New tests cover: PASS-verdict-mentions-FAIL, unstructured-verdict-fallback, implement-to-bug-find transition
  • parse_sub_tasks correctly parses structured sub-task verdicts

Non-Goals

  • Not removing substring fallback entirely (backward compat for unstructured verdicts)
  • Not changing orchestrator.md (that spec is the authority)
  • Not modifying ui/app.py verdict display logic