- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
2.1 KiB
2.1 KiB
Implementation: Autopilot Gate Integration
Changes Made
1. Gate-between-phases in autopilot
The orchestrator prompt (prompts/orchestrate.md) now defines an explicit gate-check loop:
- Read
.state→ confirm current phase - Run
status.py --validate-folder→ check for out-of-order artifacts - If violations found → STOP and report
- Load phase prompt → confirm ALLOWED/FORBIDDEN boundaries
- Execute phase → produce required artifact
- If phase requires approval →
--transition {phase}:awaiting_approval, pause for user sign-off,--approve,--transition {next-phase} - If phase does NOT require approval →
--transition {next-phase}
2. Resumption from .state
- The orchestrator reads
.statefor each task, no artifact re-derivation needed - Approval sub-states are preserved across sessions
3. Persona switching
- Orchestrator loads the prompt for the current phase based on
.state - FORBIDDEN sections in phase prompts constrain what the orchestrator can do
- Orchestrator must NOT override phase-level FORBIDDEN rules
4. Approval gates in autopilot
- Research, decomposition, design, and test_design phases ALWAYS pause for user approval in autopilot
- The pause is enforced by
status.py --transitionrefusing past:awaiting_approval - After user says "APPROVED",
status.py --approveis called, then transition proceeds
5. Session break recovery
.statefile records the last completed phase (including approval sub-states)- Next session reads
.stateand resumes exactly where it left off - No phase progress is lost on session break
6. Manual mode coexistence
- Orchestrator reads
.stateand reports current phase - User triggers phases manually, orchestrator calls
status.py --transitionandstatus.py --approve
7. Periodic audit
- Orchestrator calls
status.py --auditat session start and after task completion - Catches violations that might slip through individual phase gates
Files Modified
prompts/orchestrate.md(rewritten, 143 lines with gate-check loop)prompts/workflow.md(referenced from orchestrate.md)