Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

7.8 KiB

SPEC: Autopilot Gate Integration

Goal

Update the autopilot mode to work with the new enforcement mechanisms (.state file, phase-scoped prompts, status.py) so that the Orchestrator drives tasks through phases with hard gates between them instead of the current soft advisory approach.

Background

Autopilot currently works by having one long orchestrator prompt that tries to act as every persona. The Orchestrator executes research, then design, then implementation — all in one session with no hard boundaries. This allows the agent to blur phases or skip ahead when it encounters something familiar. The new enforcement mechanisms need to be integrated so that autopilot retains its drive-to-completion behavior while respecting phase gates.

Requirements

1. Gate-between-phases in autopilot

When the Orchestrator completes a phase in autopilot mode, it must:

  1. Call status.py --validate-folder --task {task-name} to check for out-of-order artifacts
  2. If violations are found, report them and STOP — do not proceed past a phase-skipping violation
  3. If the phase requires approval (research, decomposition, design, test_design): a. Call status.py --transition {phase}:awaiting_approval to move to the awaiting_approval sub-state b. Present the draft artifact to the user for sign-off c. STOP and wait for user approval — do NOT proceed past the approval gate in autopilot d. After user says "APPROVED", call status.py --approve to record the approval e. Call status.py --transition {next-phase} to move to the next phase
  4. If the phase does NOT require approval (implement, bug_find, adversarial_bug_find, doc_review, referee): a. Call status.py --transition {next-phase} to validate and record the transition
  5. If the transition is rejected (forbidden artifacts, missing required artifact, illegal transition, awaiting approval without approval), report the error and STOP
  6. If the transition is accepted, load the next phase's prompt and continue
  7. This replaces the current approach where the Orchestrator just "knows" what to do next

Approval gates in autopilot: Approval-requiring phases (research, decomposition, design, test_design) always pause autopilot for user sign-off, even in Autopilot mode. This is consistent with the current research/design prompts which require interactive sign-off. The difference is that now the pause is enforced by status.py --transition refusing to proceed past :awaiting_approval.

2. Resumption from .state

When the user says "orchestrate" or "continue" and the Orchestrator needs to resume:

  1. Read .state for each task (or call status.py --list)
  2. Start from the recorded phase — no need to re-derive from artifacts
  3. This is a hard resumption point — if .state says "implement", the Orchestrator starts at implement, not at research

3. Persona switching

In autopilot, the Orchestrator acts as different personas (Researcher, Designer, Implementer). With phase-scoped prompts:

  • The Orchestrator loads the prompt for the current phase (based on .state)
  • The loaded prompt's FORBIDDEN section constrains what the Orchestrator can do in that phase
  • When the phase completes, the Orchestrator transitions .state and loads the next prompt
  • The Orchestrator's own prompt must NOT override phase-level FORBIDDEN rules

4. Orchestrator prompt updates

Update orchestrate.md autopilot section:

  • Replace the drive_all() pseudocode with an explicit gate-check loop:
    For each phase in autopilot:
      1. Read .state → confirm current phase
      2. Call status.py --validate-folder → check for out-of-order artifacts
      3. If violations found → STOP and report (phase-skipping detected)
      4. Load phase prompt → confirm ALLOWED/FORBIDDEN boundaries
      5. Execute phase → produce required artifact
      6. If phase requires approval (research, decomposition, design, test_design):
         a. Call status.py --transition {phase}:awaiting_approval
         b. STOP and wait for user to say "APPROVED"
         c. Call status.py --approve
         d. Call status.py --transition {next-phase}
      7. If phase does NOT require approval:
         a. Call status.py --transition {next-phase}
      8. If transition accepted → load next phase prompt, continue
      9. If transition rejected → stop and report
    
  • Remove the current auto-execution rules that allow the Orchestrator to skip ahead
  • Add: "The Orchestrator MUST NOT perform actions that are FORBIDDEN in the current phase prompt, even in autopilot mode"

5. Session break recovery

If an autopilot session breaks (context limit, error, user interrupt):

  • The .state file records the last completed phase
  • The next session reads .state and resumes from there
  • No phase progress is lost
  • This is a major improvement over the current system where session breaks require re-deriving state from artifacts

6. Manual mode coexistence

Manual mode (Autopilot: Disabled) should also use .state:

  • The Orchestrator reads .state and reports current phase
  • The user must manually trigger each phase
  • The Orchestrator uses status.py --transition to record each transition
  • For approval-requiring phases, the user explicitly says "APPROVED" and the Orchestrator calls status.py --approve
  • The manual mode flow is: read .state → report to user → user says "implement" → Orchestrator calls status.py --transition implement → user executes phase

7. Parallel sub-task execution

In autopilot, when sub-tasks are in the same wave:

  • Each sub-task has its own .state file
  • The Orchestrator can drive them in parallel
  • The status.py --list command shows all sub-task states
  • When all Wave 1 sub-tasks reach complete or human_intervention, Wave 2 starts

8. Periodic audit during autopilot

During long autopilot runs, the Orchestrator should call status.py --audit:

  • At the start of each session (before driving any tasks)
  • After completing a full task lifecycle
  • If the orchestrator detects unexpected behavior (e.g., an artifact appeared that it didn't create)
  • The audit catches violations that might slip through individual phase gates (e.g., code edits during research that don't create an artifact file)

Acceptance Criteria

  • Orchestrator autopilot uses status.py --transition between phases
  • Orchestrator calls status.py --validate-folder before each transition
  • Orchestrator STOPS on validation violations (no proceeding past phase-skipping)
  • Orchestrator pauses at approval gates (research, decomposition, design, test_design) even in autopilot
  • Orchestrator calls status.py --transition {phase}:awaiting_approval before user sign-off
  • Orchestrator calls status.py --approve only after user says "APPROVED"
  • Orchestrator calls status.py --transition {next-phase} after approval
  • Non-approval phases (implement, bug_find, etc.) transition automatically in autopilot
  • Orchestrator reads .state for resumption (no artifact re-derivation needed)
  • Orchestrator loads phase-specific prompt for each phase (persona switching)
  • Orchestrator respects FORBIDDEN actions even in autopilot
  • Session break recovery works via .state file (including approval sub-states)
  • Manual mode uses .state, status.py --transition, and status.py --approve
  • Parallel sub-task execution uses per-sub-task .state files
  • orchestrate.md autopilot section updated with gate-check loop (including validate-folder and approval steps)
  • No duplicate state determination logic between orchestrate.md and workflow.md
  • Periodic audit during autopilot runs

Non-Goals

  • This spec does not cover the .state file format (covered by state-file-enforcement)
  • This spec does not cover prompt restructuring (covered by phase-scoped-prompts)
  • This spec does not cover status.py implementation (covered by status-script)
  • This spec does not cover dashboard updates