Fix autopilot drive-all loop — agent was stopping after one task

The orchestrator only drove ONE task per invocation, then stopped with
ORCHESTRATION_COMPLETE. User had to manually re-trigger for each task.

Changes:
- Added drive_all() outer loop that scans EVERY task and drives them all
- Tasks needing user review are flagged; orchestrator continues to next
- Only stops when ALL tasks are terminal or ALL remaining need user input
- Output format now shows session summary (completed, awaiting review, blocked)
- Session-starter references 'orchestrate' as primary command
- Updated continue-from-existing to use drive_all()
This commit is contained in:
2026-06-13 22:11:44 -04:00
parent da1799cd5c
commit a1dcf09d67
2 changed files with 68 additions and 33 deletions
+67 -33
View File
@@ -160,43 +160,70 @@ Each task is a state machine. The Orchestrator determines the current state and
## Autopilot Mode (Autopilot: Enabled in .agent.md)
In Autopilot mode, the Orchestrator MUST **drive the task all the way** to completion or until human intervention is needed. It does this by:
In Autopilot mode, the Orchestrator MUST **drive ALL tasks to completion** or until every task either reaches a terminal state or awaits user input. It does this by:
1. **Scanning**: Determine the current state of each task by checking artifacts
2. **Executing**: Run the next phase directly (the agent should execute the phase)
3. **Looping**: After each phase completes (check for `CONTRACT_MET` or the phase's stop condition), re-scan and continue to the next phase
4. **Stopping**: Stop when the task reaches a terminal state (Complete or Human Intervention)
1. **Scan all tasks**: Determine the current state of EVERY task by checking artifacts
2. **Prioritize**: Work on the most advanced task first (closest to done)
3. **Execute**: Run the next phase directly
4. **Loop**: After each phase completes, RE-SCAN all tasks — if any remain non-terminal, drive the next one
5. **Parallelize**: When tasks are independent (different parent, same stage), work them in parallel
6. **Defer user blocks**: If a task requires user approval, flag it and move to the next task that doesn't
7. **Stop only when**: ALL tasks are terminal (Complete or Human Intervention) or ALL remaining tasks are blocked by user input
### Auto-Execution Loop
### Drive-All Loop
```
while task is not in terminal state:
if iteration_count >= MAX_ITERATIONS (default: 10):
break (human intervention needed — too many iterations)
if total_time_elapsed >= MAX_TOTAL_TIME (default: 24 hours):
break (human intervention needed — too much time elapsed)
if phase_time_elapsed >= MAX_PHASE_TIME (default: 1 hour):
break (human intervention needed — phase took too long)
determine current state
execute the phase that moves the task forward
wait for phase to complete (CONTRACT_MET or stop condition)
if phase failed (FAIL/NEEDS_REVIEW verdict):
break (human intervention needed)
if phase artifact is empty or malformed:
break (human intervention needed — artifact validation failed)
if phase succeeded:
function drive_all():
tasks = scan_all_tasks()
non_terminal = [t for t in tasks if not is_terminal(t)]
while non_terminal:
unblocked = [t for t in non_terminal if not needs_user_input(t)]
if not unblocked:
# All remaining tasks need user input — report and stop
report_pending_reviews(non_terminal)
output "ORCHESTRATION_COMPLETE — awaiting user review"
return
# Sort by advancement (most advanced first)
sort_by_advancement(unblocked)
task = unblocked[0]
drive_task(task)
# After completing a task phase, re-scan
non_terminal = [t for t in scan_all_tasks() if not is_terminal(t)]
output "ORCHESTRATION_COMPLETE — all tasks done"
function drive_task(task):
iteration_count = 0
while not is_terminal(task):
if iteration_count >= MAX_ITERATIONS (default: 10):
flag_human_intervention(task, "too many iterations")
return
phase = determine_next_phase(task)
if phase == REVIEW_REQUIRED:
flag_review_needed(task)
return
execute_phase(phase)
wait_for_completion()
if phase_failed():
flag_human_intervention(task, "phase failed")
return
iteration_count++
continue loop
```
### Task Creation in Autopilot
#### Continue from existing tasks
If the user says "orchestrate" or "continue" with no new task description, the Orchestrator should:
1. Scan all tasks in `{project}/.automaton/tasks/`, including sub-task folders under `{project}/.automaton/tasks/{parent-task}/subtasks/`
2. Find the most advanced task (the one closest to completion) — **prioritize sub-tasks over parent tasks** (because the parent depends on the sub-tasks)
3. When choosing among sub-tasks in the same wave, prioritize those in later phases (e.g., Bug Find over Research) because they are closer to completion
4. Drive that task through the remaining phases
If the user says "orchestrate" or "continue" with no new task description, the Orchestrator should run `drive_all()`:
1. Scan ALL tasks in `{project}/.automaton/tasks/`, including sub-task folders
2. Work through every non-terminal task in order of advancement (most advanced first)
3. For tasks needing user review — flag them, report to user, and continue with tasks that don't
4. Stop only when ALL tasks are terminal or ALL remaining tasks need user input
5. Output a final summary showing which tasks completed and which await review
#### New tasks from user input
If {task-description} contains a description for a NEW task, the Orchestrator MUST:
@@ -260,19 +287,26 @@ Examine the `{project}/.automaton/tasks/` directory and determine the state of e
### Default Mode — Autopilot (Autopilot: Enabled)
In Autopilot mode, the Orchestrator auto-executes all phases until completion or human intervention:
In Autopilot mode, the Orchestrator runs `drive_all()` — driving every task forward until all are complete or blocked by user input.
**Task: {task-folder-name}**
**Drive-All Summary:**
- **Tasks completed this session**: {count}
- **Tasks awaiting review**: {count} — see flagged tasks below
- **Tasks remaining**: {count} — blocked by dependencies
- **Phase**: {Current Phase for active work}
**Active task: {task-folder-name}**
- **Status**: {Current Phase}
- **Next Step**: {Next Phase}
- **Auto-Execute**: YES
- **Command**:
> "{Command to trigger the next phase}"
If a task has `FAIL` or `NEEDS_REVIEW` verdict or Tie-Breaks, after the task status output, explicitly state:
"⚠️ **HUMAN INTERVENTION REQUIRED**: {Reason}"
**Flagged for review:**
{For each task needing user review}
- **{task-name}**: {Phase completed} — awaiting approval to proceed
When finished, output "ORCHESTRATION_COMPLETE".
When ALL tasks are terminal, output "ORCHESTRATION_COMPLETE — all tasks done".
When all remaining tasks need user input, output "ORCHESTRATION_COMPLETE — awaiting user review".
### Manual Mode (Autopilot: Disabled)