Files
automaton/tasks/autopilot-gate-integration/SPEC.md
T
gitea 05c76852a2
CI / build (push) Has been cancelled
v2.0: state enforcement, project scoping, harness integration
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks

Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)

Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)

Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2026-06-15 14:16:46 -04:00

7.8 KiB

SPEC: Autopilot Gate Integration

Goal

Update the autopilot mode to work with the new enforcement mechanisms (.state file, phase-scoped prompts, status.py) so that the Orchestrator drives tasks through phases with hard gates between them instead of the current soft advisory approach.

Background

Autopilot currently works by having one long orchestrator prompt that tries to act as every persona. The Orchestrator executes research, then design, then implementation — all in one session with no hard boundaries. This allows the agent to blur phases or skip ahead when it encounters something familiar. The new enforcement mechanisms need to be integrated so that autopilot retains its drive-to-completion behavior while respecting phase gates.

Requirements

1. Gate-between-phases in autopilot

When the Orchestrator completes a phase in autopilot mode, it must:

  1. Call status.py --validate-folder --task {task-name} to check for out-of-order artifacts
  2. If violations are found, report them and STOP — do not proceed past a phase-skipping violation
  3. If the phase requires approval (research, decomposition, design, test_design): a. Call status.py --transition {phase}:awaiting_approval to move to the awaiting_approval sub-state b. Present the draft artifact to the user for sign-off c. STOP and wait for user approval — do NOT proceed past the approval gate in autopilot d. After user says "APPROVED", call status.py --approve to record the approval e. Call status.py --transition {next-phase} to move to the next phase
  4. If the phase does NOT require approval (implement, bug_find, adversarial_bug_find, doc_review, referee): a. Call status.py --transition {next-phase} to validate and record the transition
  5. If the transition is rejected (forbidden artifacts, missing required artifact, illegal transition, awaiting approval without approval), report the error and STOP
  6. If the transition is accepted, load the next phase's prompt and continue
  7. This replaces the current approach where the Orchestrator just "knows" what to do next

Approval gates in autopilot: Approval-requiring phases (research, decomposition, design, test_design) always pause autopilot for user sign-off, even in Autopilot mode. This is consistent with the current research/design prompts which require interactive sign-off. The difference is that now the pause is enforced by status.py --transition refusing to proceed past :awaiting_approval.

2. Resumption from .state

When the user says "orchestrate" or "continue" and the Orchestrator needs to resume:

  1. Read .state for each task (or call status.py --list)
  2. Start from the recorded phase — no need to re-derive from artifacts
  3. This is a hard resumption point — if .state says "implement", the Orchestrator starts at implement, not at research

3. Persona switching

In autopilot, the Orchestrator acts as different personas (Researcher, Designer, Implementer). With phase-scoped prompts:

  • The Orchestrator loads the prompt for the current phase (based on .state)
  • The loaded prompt's FORBIDDEN section constrains what the Orchestrator can do in that phase
  • When the phase completes, the Orchestrator transitions .state and loads the next prompt
  • The Orchestrator's own prompt must NOT override phase-level FORBIDDEN rules

4. Orchestrator prompt updates

Update orchestrate.md autopilot section:

  • Replace the drive_all() pseudocode with an explicit gate-check loop:
    For each phase in autopilot:
      1. Read .state → confirm current phase
      2. Call status.py --validate-folder → check for out-of-order artifacts
      3. If violations found → STOP and report (phase-skipping detected)
      4. Load phase prompt → confirm ALLOWED/FORBIDDEN boundaries
      5. Execute phase → produce required artifact
      6. If phase requires approval (research, decomposition, design, test_design):
         a. Call status.py --transition {phase}:awaiting_approval
         b. STOP and wait for user to say "APPROVED"
         c. Call status.py --approve
         d. Call status.py --transition {next-phase}
      7. If phase does NOT require approval:
         a. Call status.py --transition {next-phase}
      8. If transition accepted → load next phase prompt, continue
      9. If transition rejected → stop and report
    
  • Remove the current auto-execution rules that allow the Orchestrator to skip ahead
  • Add: "The Orchestrator MUST NOT perform actions that are FORBIDDEN in the current phase prompt, even in autopilot mode"

5. Session break recovery

If an autopilot session breaks (context limit, error, user interrupt):

  • The .state file records the last completed phase
  • The next session reads .state and resumes from there
  • No phase progress is lost
  • This is a major improvement over the current system where session breaks require re-deriving state from artifacts

6. Manual mode coexistence

Manual mode (Autopilot: Disabled) should also use .state:

  • The Orchestrator reads .state and reports current phase
  • The user must manually trigger each phase
  • The Orchestrator uses status.py --transition to record each transition
  • For approval-requiring phases, the user explicitly says "APPROVED" and the Orchestrator calls status.py --approve
  • The manual mode flow is: read .state → report to user → user says "implement" → Orchestrator calls status.py --transition implement → user executes phase

7. Parallel sub-task execution

In autopilot, when sub-tasks are in the same wave:

  • Each sub-task has its own .state file
  • The Orchestrator can drive them in parallel
  • The status.py --list command shows all sub-task states
  • When all Wave 1 sub-tasks reach complete or human_intervention, Wave 2 starts

8. Periodic audit during autopilot

During long autopilot runs, the Orchestrator should call status.py --audit:

  • At the start of each session (before driving any tasks)
  • After completing a full task lifecycle
  • If the orchestrator detects unexpected behavior (e.g., an artifact appeared that it didn't create)
  • The audit catches violations that might slip through individual phase gates (e.g., code edits during research that don't create an artifact file)

Acceptance Criteria

  • Orchestrator autopilot uses status.py --transition between phases
  • Orchestrator calls status.py --validate-folder before each transition
  • Orchestrator STOPS on validation violations (no proceeding past phase-skipping)
  • Orchestrator pauses at approval gates (research, decomposition, design, test_design) even in autopilot
  • Orchestrator calls status.py --transition {phase}:awaiting_approval before user sign-off
  • Orchestrator calls status.py --approve only after user says "APPROVED"
  • Orchestrator calls status.py --transition {next-phase} after approval
  • Non-approval phases (implement, bug_find, etc.) transition automatically in autopilot
  • Orchestrator reads .state for resumption (no artifact re-derivation needed)
  • Orchestrator loads phase-specific prompt for each phase (persona switching)
  • Orchestrator respects FORBIDDEN actions even in autopilot
  • Session break recovery works via .state file (including approval sub-states)
  • Manual mode uses .state, status.py --transition, and status.py --approve
  • Parallel sub-task execution uses per-sub-task .state files
  • orchestrate.md autopilot section updated with gate-check loop (including validate-folder and approval steps)
  • No duplicate state determination logic between orchestrate.md and workflow.md
  • Periodic audit during autopilot runs

Non-Goals

  • This spec does not cover the .state file format (covered by state-file-enforcement)
  • This spec does not cover prompt restructuring (covered by phase-scoped-prompts)
  • This spec does not cover status.py implementation (covered by status-script)
  • This spec does not cover dashboard updates