- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
level (all <7 days old per the cleanup policy; premature bulk archive
was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
the board via innerHTML every 2s, destroying each column-body's
scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
view.scrollTop before rebuild and restores after (matched by
PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
section cards, transition buttons, inline artifact editor (textarea for
writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
pointing at a pytest temp dir (test isolation leak). Rewired to point
at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
but a real install creates it. Now snapshots mtime before run, asserts
unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
board renders tasks, column scroll survives auto-refresh tick.
Verified the test fails without the scroll fix (scrollTop resets to 0).
Skipped via importorskip when playwright is absent (main CI stays
green).
- **Clarify SI loop scope in README** — new-project onboarding section
documents the framework-scoped self-improvement loop and options
(leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
gap (mde tasks marked complete but per-role model binding was never
implemented).
The review system (REVIEW.md, approve/changes_requested buttons,
review filter, review badges) was purely cosmetic — only the dashboard
read/wrote it. No workflow component (status.py, autopilot.py,
loop-runner.py, prompts) ever enforced it.
The 'Approve' button in the task detail panel confused users into
thinking it approved the task's phase gate. In reality it only wrote
to REVIEW.md, which had zero effect on transitions.
Removed:
- Review section (buttons, textarea, status badge) from detail panel
- Review badge from task cards
- Review filter from toolbar
- Pending-review counter from header
- All review-related CSS
Users now use the single '🔓 Approve Phase' button in the detail
panel, which calls status.py --approve and actually transitions the
task.
Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)
Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.
Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.