- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
7.8 KiB
SPEC: Autopilot Gate Integration
Goal
Update the autopilot mode to work with the new enforcement mechanisms (.state file, phase-scoped prompts, status.py) so that the Orchestrator drives tasks through phases with hard gates between them instead of the current soft advisory approach.
Background
Autopilot currently works by having one long orchestrator prompt that tries to act as every persona. The Orchestrator executes research, then design, then implementation — all in one session with no hard boundaries. This allows the agent to blur phases or skip ahead when it encounters something familiar. The new enforcement mechanisms need to be integrated so that autopilot retains its drive-to-completion behavior while respecting phase gates.
Requirements
1. Gate-between-phases in autopilot
When the Orchestrator completes a phase in autopilot mode, it must:
- Call
status.py --validate-folder --task {task-name}to check for out-of-order artifacts - If violations are found, report them and STOP — do not proceed past a phase-skipping violation
- If the phase requires approval (research, decomposition, design, test_design):
a. Call
status.py --transition {phase}:awaiting_approvalto move to the awaiting_approval sub-state b. Present the draft artifact to the user for sign-off c. STOP and wait for user approval — do NOT proceed past the approval gate in autopilot d. After user says "APPROVED", callstatus.py --approveto record the approval e. Callstatus.py --transition {next-phase}to move to the next phase - If the phase does NOT require approval (implement, bug_find, adversarial_bug_find, doc_review, referee):
a. Call
status.py --transition {next-phase}to validate and record the transition - If the transition is rejected (forbidden artifacts, missing required artifact, illegal transition, awaiting approval without approval), report the error and STOP
- If the transition is accepted, load the next phase's prompt and continue
- This replaces the current approach where the Orchestrator just "knows" what to do next
Approval gates in autopilot: Approval-requiring phases (research, decomposition, design, test_design) always pause autopilot for user sign-off, even in Autopilot mode. This is consistent with the current research/design prompts which require interactive sign-off. The difference is that now the pause is enforced by status.py --transition refusing to proceed past :awaiting_approval.
2. Resumption from .state
When the user says "orchestrate" or "continue" and the Orchestrator needs to resume:
- Read
.statefor each task (or callstatus.py --list) - Start from the recorded phase — no need to re-derive from artifacts
- This is a hard resumption point — if
.statesays "implement", the Orchestrator starts at implement, not at research
3. Persona switching
In autopilot, the Orchestrator acts as different personas (Researcher, Designer, Implementer). With phase-scoped prompts:
- The Orchestrator loads the prompt for the current phase (based on
.state) - The loaded prompt's FORBIDDEN section constrains what the Orchestrator can do in that phase
- When the phase completes, the Orchestrator transitions
.stateand loads the next prompt - The Orchestrator's own prompt must NOT override phase-level FORBIDDEN rules
4. Orchestrator prompt updates
Update orchestrate.md autopilot section:
- Replace the
drive_all()pseudocode with an explicit gate-check loop:For each phase in autopilot: 1. Read .state → confirm current phase 2. Call status.py --validate-folder → check for out-of-order artifacts 3. If violations found → STOP and report (phase-skipping detected) 4. Load phase prompt → confirm ALLOWED/FORBIDDEN boundaries 5. Execute phase → produce required artifact 6. If phase requires approval (research, decomposition, design, test_design): a. Call status.py --transition {phase}:awaiting_approval b. STOP and wait for user to say "APPROVED" c. Call status.py --approve d. Call status.py --transition {next-phase} 7. If phase does NOT require approval: a. Call status.py --transition {next-phase} 8. If transition accepted → load next phase prompt, continue 9. If transition rejected → stop and report - Remove the current auto-execution rules that allow the Orchestrator to skip ahead
- Add: "The Orchestrator MUST NOT perform actions that are FORBIDDEN in the current phase prompt, even in autopilot mode"
5. Session break recovery
If an autopilot session breaks (context limit, error, user interrupt):
- The
.statefile records the last completed phase - The next session reads
.stateand resumes from there - No phase progress is lost
- This is a major improvement over the current system where session breaks require re-deriving state from artifacts
6. Manual mode coexistence
Manual mode (Autopilot: Disabled) should also use .state:
- The Orchestrator reads
.stateand reports current phase - The user must manually trigger each phase
- The Orchestrator uses
status.py --transitionto record each transition - For approval-requiring phases, the user explicitly says "APPROVED" and the Orchestrator calls
status.py --approve - The manual mode flow is: read
.state→ report to user → user says "implement" → Orchestrator callsstatus.py --transition implement→ user executes phase
7. Parallel sub-task execution
In autopilot, when sub-tasks are in the same wave:
- Each sub-task has its own
.statefile - The Orchestrator can drive them in parallel
- The
status.py --listcommand shows all sub-task states - When all Wave 1 sub-tasks reach
completeorhuman_intervention, Wave 2 starts
8. Periodic audit during autopilot
During long autopilot runs, the Orchestrator should call status.py --audit:
- At the start of each session (before driving any tasks)
- After completing a full task lifecycle
- If the orchestrator detects unexpected behavior (e.g., an artifact appeared that it didn't create)
- The audit catches violations that might slip through individual phase gates (e.g., code edits during research that don't create an artifact file)
Acceptance Criteria
- Orchestrator autopilot uses
status.py --transitionbetween phases - Orchestrator calls
status.py --validate-folderbefore each transition - Orchestrator STOPS on validation violations (no proceeding past phase-skipping)
- Orchestrator pauses at approval gates (research, decomposition, design, test_design) even in autopilot
- Orchestrator calls
status.py --transition {phase}:awaiting_approvalbefore user sign-off - Orchestrator calls
status.py --approveonly after user says "APPROVED" - Orchestrator calls
status.py --transition {next-phase}after approval - Non-approval phases (implement, bug_find, etc.) transition automatically in autopilot
- Orchestrator reads
.statefor resumption (no artifact re-derivation needed) - Orchestrator loads phase-specific prompt for each phase (persona switching)
- Orchestrator respects FORBIDDEN actions even in autopilot
- Session break recovery works via
.statefile (including approval sub-states) - Manual mode uses
.state,status.py --transition, andstatus.py --approve - Parallel sub-task execution uses per-sub-task
.statefiles orchestrate.mdautopilot section updated with gate-check loop (including validate-folder and approval steps)- No duplicate state determination logic between orchestrate.md and workflow.md
- Periodic audit during autopilot runs
Non-Goals
- This spec does not cover the
.statefile format (covered by state-file-enforcement) - This spec does not cover prompt restructuring (covered by phase-scoped-prompts)
- This spec does not cover
status.pyimplementation (covered by status-script) - This spec does not cover dashboard updates