Fix autopilot drive-all loop — agent was stopping after one task

The orchestrator only drove ONE task per invocation, then stopped with
ORCHESTRATION_COMPLETE. User had to manually re-trigger for each task.

Changes:
- Added drive_all() outer loop that scans EVERY task and drives them all
- Tasks needing user review are flagged; orchestrator continues to next
- Only stops when ALL tasks are terminal or ALL remaining need user input
- Output format now shows session summary (completed, awaiting review, blocked)
- Session-starter references 'orchestrate' as primary command
- Updated continue-from-existing to use drive_all()
This commit is contained in:
2026-06-13 22:11:44 -04:00
parent da1799cd5c
commit a1dcf09d67
2 changed files with 68 additions and 33 deletions
+67 -33
View File
@@ -160,43 +160,70 @@ Each task is a state machine. The Orchestrator determines the current state and
## Autopilot Mode (Autopilot: Enabled in .agent.md) ## Autopilot Mode (Autopilot: Enabled in .agent.md)
In Autopilot mode, the Orchestrator MUST **drive the task all the way** to completion or until human intervention is needed. It does this by: In Autopilot mode, the Orchestrator MUST **drive ALL tasks to completion** or until every task either reaches a terminal state or awaits user input. It does this by:
1. **Scanning**: Determine the current state of each task by checking artifacts 1. **Scan all tasks**: Determine the current state of EVERY task by checking artifacts
2. **Executing**: Run the next phase directly (the agent should execute the phase) 2. **Prioritize**: Work on the most advanced task first (closest to done)
3. **Looping**: After each phase completes (check for `CONTRACT_MET` or the phase's stop condition), re-scan and continue to the next phase 3. **Execute**: Run the next phase directly
4. **Stopping**: Stop when the task reaches a terminal state (Complete or Human Intervention) 4. **Loop**: After each phase completes, RE-SCAN all tasks — if any remain non-terminal, drive the next one
5. **Parallelize**: When tasks are independent (different parent, same stage), work them in parallel
6. **Defer user blocks**: If a task requires user approval, flag it and move to the next task that doesn't
7. **Stop only when**: ALL tasks are terminal (Complete or Human Intervention) or ALL remaining tasks are blocked by user input
### Auto-Execution Loop ### Drive-All Loop
``` ```
while task is not in terminal state: function drive_all():
if iteration_count >= MAX_ITERATIONS (default: 10): tasks = scan_all_tasks()
break (human intervention needed — too many iterations) non_terminal = [t for t in tasks if not is_terminal(t)]
if total_time_elapsed >= MAX_TOTAL_TIME (default: 24 hours):
break (human intervention needed — too much time elapsed) while non_terminal:
if phase_time_elapsed >= MAX_PHASE_TIME (default: 1 hour): unblocked = [t for t in non_terminal if not needs_user_input(t)]
break (human intervention needed — phase took too long)
determine current state if not unblocked:
execute the phase that moves the task forward # All remaining tasks need user input — report and stop
wait for phase to complete (CONTRACT_MET or stop condition) report_pending_reviews(non_terminal)
if phase failed (FAIL/NEEDS_REVIEW verdict): output "ORCHESTRATION_COMPLETE — awaiting user review"
break (human intervention needed) return
if phase artifact is empty or malformed:
break (human intervention needed — artifact validation failed) # Sort by advancement (most advanced first)
if phase succeeded: sort_by_advancement(unblocked)
task = unblocked[0]
drive_task(task)
# After completing a task phase, re-scan
non_terminal = [t for t in scan_all_tasks() if not is_terminal(t)]
output "ORCHESTRATION_COMPLETE — all tasks done"
function drive_task(task):
iteration_count = 0
while not is_terminal(task):
if iteration_count >= MAX_ITERATIONS (default: 10):
flag_human_intervention(task, "too many iterations")
return
phase = determine_next_phase(task)
if phase == REVIEW_REQUIRED:
flag_review_needed(task)
return
execute_phase(phase)
wait_for_completion()
if phase_failed():
flag_human_intervention(task, "phase failed")
return
iteration_count++ iteration_count++
continue loop
``` ```
### Task Creation in Autopilot ### Task Creation in Autopilot
#### Continue from existing tasks #### Continue from existing tasks
If the user says "orchestrate" or "continue" with no new task description, the Orchestrator should: If the user says "orchestrate" or "continue" with no new task description, the Orchestrator should run `drive_all()`:
1. Scan all tasks in `{project}/.automaton/tasks/`, including sub-task folders under `{project}/.automaton/tasks/{parent-task}/subtasks/` 1. Scan ALL tasks in `{project}/.automaton/tasks/`, including sub-task folders
2. Find the most advanced task (the one closest to completion) — **prioritize sub-tasks over parent tasks** (because the parent depends on the sub-tasks) 2. Work through every non-terminal task in order of advancement (most advanced first)
3. When choosing among sub-tasks in the same wave, prioritize those in later phases (e.g., Bug Find over Research) because they are closer to completion 3. For tasks needing user review — flag them, report to user, and continue with tasks that don't
4. Drive that task through the remaining phases 4. Stop only when ALL tasks are terminal or ALL remaining tasks need user input
5. Output a final summary showing which tasks completed and which await review
#### New tasks from user input #### New tasks from user input
If {task-description} contains a description for a NEW task, the Orchestrator MUST: If {task-description} contains a description for a NEW task, the Orchestrator MUST:
@@ -260,19 +287,26 @@ Examine the `{project}/.automaton/tasks/` directory and determine the state of e
### Default Mode — Autopilot (Autopilot: Enabled) ### Default Mode — Autopilot (Autopilot: Enabled)
In Autopilot mode, the Orchestrator auto-executes all phases until completion or human intervention: In Autopilot mode, the Orchestrator runs `drive_all()` — driving every task forward until all are complete or blocked by user input.
**Task: {task-folder-name}** **Drive-All Summary:**
- **Tasks completed this session**: {count}
- **Tasks awaiting review**: {count} — see flagged tasks below
- **Tasks remaining**: {count} — blocked by dependencies
- **Phase**: {Current Phase for active work}
**Active task: {task-folder-name}**
- **Status**: {Current Phase} - **Status**: {Current Phase}
- **Next Step**: {Next Phase} - **Next Step**: {Next Phase}
- **Auto-Execute**: YES
- **Command**: - **Command**:
> "{Command to trigger the next phase}" > "{Command to trigger the next phase}"
If a task has `FAIL` or `NEEDS_REVIEW` verdict or Tie-Breaks, after the task status output, explicitly state: **Flagged for review:**
"⚠️ **HUMAN INTERVENTION REQUIRED**: {Reason}" {For each task needing user review}
- **{task-name}**: {Phase completed} — awaiting approval to proceed
When finished, output "ORCHESTRATION_COMPLETE". When ALL tasks are terminal, output "ORCHESTRATION_COMPLETE — all tasks done".
When all remaining tasks need user input, output "ORCHESTRATION_COMPLETE — awaiting user review".
### Manual Mode (Autopilot: Disabled) ### Manual Mode (Autopilot: Disabled)
+1
View File
@@ -20,6 +20,7 @@ After reading these files, acknowledge with: ".agent.md and .rules.md loaded. Re
1. Start a fresh session with your agent. 1. Start a fresh session with your agent.
2. Paste the block above (replace `{project}` with the actual path). 2. Paste the block above (replace `{project}` with the actual path).
3. Then give your real request, for example: 3. Then give your real request, for example:
- "orchestrate" — drive ALL remaining tasks to completion
- "What should we do next on this project?" - "What should we do next on this project?"
- "Implement the fix-alert-test task" - "Implement the fix-alert-test task"
- "Onboard this project to the agent framework" - "Onboard this project to the agent framework"