Fix all 30 bugs: Orchestrator state determination, auto-execution loop, sub-task management, VRAM detection, and more

This commit is contained in:
2026-06-11 21:54:15 -04:00
parent a74eadfb86
commit 8897852cc0
5 changed files with 522 additions and 174 deletions
+105 -48
View File
@@ -16,8 +16,8 @@ When VRAM configuration is needed (during task decomposition, sub-task creation,
### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`.
2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
1. **Auto-detect via script**: Check if `{project}/.agent-framework/scripts/vram_detect.sh` exists. If it does, run it to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`.
2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window. **Important**: Only read the specific lines needed (e.g., the model name line), not the entire file. Limit file reads to 10KB to prevent memory exhaustion.
3. **Manual override**: Check if `~/.agent-framework/config.md` has `Auto-detect: No` under VRAM Configuration. If so, use the manually specified values.
4. **Fallback**: Use 8k tokens as default, with 25% headroom.
@@ -41,9 +41,9 @@ When model context window is needed, the Orchestrator MUST attempt to detect it
#### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to detect the model name and its context window. Parse the JSON output for `model_context_kb`.
1. **Auto-detect via script**: Check if `{project}/.agent-framework/scripts/vram_detect.sh` exists. If it does, run it to detect the model name and its context window. Parse the JSON output for `model_context_kb`.
2. **Auto-detect via config**: Check `~/.agent-framework/config.md` for the model name and override context window.
3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window. **Important**: Only read the specific lines needed (e.g., the model name line), not the entire file. Limit file reads to 10KB to prevent memory exhaustion.
4. **Fallback**: Use 128k tokens as default (common for modern models).
#### How to Read Model Config from config.md
@@ -91,6 +91,8 @@ The Orchestrator should run VRAM detection in the following scenarios:
3. **When sub-tasks are created** — to propagate VRAM config to sub-task folders.
4. **When a sub-task's VRAM_CONFIG.md is missing** — to create one with auto-detected values.
**VRAM Detection Caching**: When the Orchestrator is invoked multiple times (e.g., the user says "orchestrate" twice), it MUST cache the VRAM detection results and reuse them instead of running the detection script again. This prevents performance degradation from repeated GPU/RAM probing. The cache should be stored in a temporary file (e.g., `{project}/.agent-framework/.vram_cache.json`) and invalidated when a new task is created or Decomposition is triggered.
### Reporting Detection Results
When auto-detecting VRAM, the Orchestrator should report:
@@ -112,6 +114,13 @@ VRAM Detection Results:
- **Using: 16k tokens** (auto-detected)
```
### Error Handling
- If the detection script does not exist, skip to the next detection method.
- If the detection script fails (e.g., `nvidia-smi` is not installed, the GPU is busy, etc.), check the exit code and fall back to the next detection method.
- If API config files are large (>10KB), read only the specific lines needed (e.g., the model name line) and skip the rest.
- If API config files contain API keys, warn the user that the VRAM detection script may be reading them.
## Task
{task-description}
@@ -151,12 +160,21 @@ In Autopilot mode, the Orchestrator MUST **drive the task all the way** to compl
```
while task is not in terminal state:
if iteration_count >= MAX_ITERATIONS (default: 10):
break (human intervention needed — too many iterations)
if total_time_elapsed >= MAX_TOTAL_TIME (default: 24 hours):
break (human intervention needed — too much time elapsed)
if phase_time_elapsed >= MAX_PHASE_TIME (default: 1 hour):
break (human intervention needed — phase took too long)
determine current state
execute the phase that moves the task forward
wait for phase to complete (CONTRACT_MET or stop condition)
if phase failed (FAIL/NEEDS_REVIEW verdict):
break (human intervention needed)
if phase artifact is empty or malformed:
break (human intervention needed — artifact validation failed)
if phase succeeded:
iteration_count++
continue loop
```
@@ -164,9 +182,10 @@ while task is not in terminal state:
#### Continue from existing tasks
If the user says "orchestrate" or "continue" with no new task description, the Orchestrator should:
1. Scan all tasks in the tasks/ directory
2. Find the most advanced task (the one closest to completion)
3. Drive that task through the remaining phases
1. Scan all tasks in the tasks/ directory, including sub-task folders under `tasks/{parent-task}/subtasks/`
2. Find the most advanced task (the one closest to completion) — **prioritize sub-tasks over parent tasks** (because the parent depends on the sub-tasks)
3. When choosing among sub-tasks in the same wave, prioritize those in later phases (e.g., Bug Find over Research) because they are closer to completion
4. Drive that task through the remaining phases
#### New tasks from user input
If {task-description} contains a description for a NEW task, the Orchestrator MUST:
@@ -174,7 +193,7 @@ If {task-description} contains a description for a NEW task, the Orchestrator MU
2. Create the task folder: `{project}/tasks/{task-name}/` (empty — no artifact files)
3. **Immediately drive it to completion** using the auto-execution loop
Note: `IMPLEMENTATION.md` is the artifact produced by the implementation phase, not the Orchestrator. Do not pre-create it.
Note: `IMPLEMENTATION.md` is the artifact produced by the implementation phase, not the Orchestrator. Do not pre-create it. **Do not create an empty `IMPLEMENTATION.md` for new tasks** — this is inconsistent with the Orchestrator's own rule and can confuse the Research phase.
#### Tasks from bug verdicts (FAIL / NEEDS_REVIEW)
If the Orchestrator detects a `VERDICT.md` with `FAIL` or `NEEDS_REVIEW` for an existing task, the behavior depends on the mode:
@@ -207,30 +226,31 @@ In manual mode, the Orchestrator only **reports** the current state and the next
## State Determination
Examine the tasks/ directory and determine the state of each task folder. Check from the most advanced state backward:
1. Has `VERDICT.md` with `PASS` → **Complete**
2. Has `VERDICT.md` with `NEEDS_REVIEW` or `FAIL` → **Human Intervention**
3. Has `DOC_REVIEW.md` → **Referee**
4. Has `ADVERSARIAL_BUG_REPORT.md` and `BUG_REPORT.md` and `SPEC.md` → **Doc Review**
5. Has `BUG_REPORT.md` and `SPEC.md` but no `ADVERSARIAL_BUG_REPORT.md` → **Adversarial Bug Find**
6. Has `SPEC.md` but no `BUG_REPORT.md` and no `ADVERSARIAL_BUG_REPORT.md` → **Bug Find**
7. Has `IMPLEMENTATION.md` → **Bug Find**
8. Has `TEST_PLAN.md` → **Implement**
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
12. Has `SPEC.md` and `DECOMPOSITION.md` → **Decomposition** (sub-tasks will be created)
13. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
Examine the tasks/ directory and determine the state of each task folder. Check from the most advanced state backward. **Important**: Always check that artifacts are non-empty before considering them as indicators of task state.
1. Has `VERDICT.md` with `PASS` (non-empty) → **Complete**
2. Has `VERDICT.md` with `NEEDS_REVIEW` or `FAIL` (non-empty) → **Human Intervention**
3. Has `DOC_REVIEW.md` (non-empty) → **Referee**
4. Has `ADVERSARIAL_BUG_REPORT.md` (non-empty) and `BUG_REPORT.md` (non-empty) and `SPEC.md` (non-empty) → **Doc Review**
5. Has `BUG_REPORT.md` (non-empty) and `SPEC.md` (non-empty) but no `ADVERSARIAL_BUG_REPORT.md` → **Adversarial Bug Find**
6. Has `SPEC.md` (non-empty) but no `BUG_REPORT.md` and no `ADVERSARIAL_BUG_REPORT.md` → **Bug Find**
7. Has `IMPLEMENTATION.md` (non-empty) → **Bug Find**
8. Has `TEST_PLAN.md` (non-empty) → **Implement**
9. Has `SPEC.md` (non-empty) and `TEST_PLAN.md` (non-empty) → **Implement**
10. Has `DESIGN.md` (non-empty) → **Test Design** (optional) or **Implement** (if user skips test design)
11. Has `SPEC.md` (non-empty) and `DESIGN.md` (non-empty) → **Test Design** (optional) or **Implement** (if user skips test design)
12. Has `SPEC.md` (non-empty) and `DECOMPOSITION.md` (non-empty) → **Decomposition** (sub-tasks will be created)
13. Has `SPEC.md` (non-empty) → **Design** (optional) or **Implement** (if user skips design)
14. No artifacts → **New**
**Note on overlapping conditions**: If a task has both `TEST_PLAN.md` and `DESIGN.md`, the Orchestrator should prioritize the more advanced state (TEST_PLAN.md → Implement) over the optional state (DESIGN.md → Test Design). Similarly, if a task has both `IMPLEMENTATION.md` and `BUG_REPORT.md`, the Orchestrator should prioritize the more advanced state (BUG_REPORT.md → Adversarial Bug Find) over the earlier state (IMPLEMENTATION.md → Bug Find).
## Output Format
### Default Mode — Autopilot (Autopilot: Enabled)
In Autopilot mode, the Orchestrator auto-executes all phases until completion or human intervention:
In Autopilot mode, the Orchestrator auto-executes all phases until completion or human intervention:
**Task: {task-folder-name}**
- **Status**: {Current Phase}
- **Next Step**: {Next Phase}
@@ -275,10 +295,13 @@ In Autopilot mode, after outputting the task statuses, the Orchestrator MUST aut
4. Wait for the phase to complete (check for `CONTRACT_MET` or the phase's stop condition)
5. If the phase completes successfully, continue to the next phase
6. If the phase fails (FAIL verdict, HUMAN INTERVENTION REQUIRED), stop and report
7. If the phase artifact is empty or malformed, stop and report human intervention
8. If the phase takes too long (exceeds MAX_PHASE_TIME), stop and report human intervention
9. If the total iterations exceed MAX_ITERATIONS, stop and report human intervention
When finished, output "ORCHESTRATION_COMPLETE".
Only recommend one phase at a time. Do not suggest running multiple phases in parallel.
**Sub-task Parallel Execution**: When sub-tasks are in the same wave and can run in parallel, the Orchestrator should drive them simultaneously instead of sequentially. The Orchestrator should output the commands for all sub-tasks in the wave and wait for all of them to complete before moving to the next wave. This is the only time multiple phases should be suggested in parallel.
## Sub-Task Management
@@ -317,17 +340,21 @@ When a parent task reaches the **Decomposition** phase (has `SPEC.md` and `DECOM
5. **Verify VRAM constraints**:
- For each sub-task, check if its estimated peak context (from DECOMPOSITION.md) is within the max peak context.
- If a sub-task exceeds the VRAM limit, create a warning and suggest splitting it further.
6. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`:
6. **Check for existing sub-task folders**: For each sub-task, check if the folder `tasks/{parent-task-name}/subtasks/{sub-task-name}/` already exists. If it does, skip the creation and report that the sub-task has already been created. This prevents duplicate sub-tasks when the Orchestrator is invoked multiple times.
7. **Check for DECOMPOSITION.md updates**: Compare the DECOMPOSITION.md with the existing sub-task folders. If the DECOMPOSITION.md has been updated (new sub-tasks added or existing sub-tasks removed), update the sub-task folders accordingly:
- For new sub-tasks, create the folders.
- For removed sub-tasks, report the orphaned sub-tasks and delete the folders.
8. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`:
- For each sub-task, create an empty folder (no files yet — the sub-task starts at the Research phase).
- The sub-task name should be a kebab-case version of the sub-task description from DECOMPOSITION.md.
7. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration.
8. **Create a PARENT_SPEC.md** file for each sub-task with:
- The parent task's SPEC.md content
- A reference to the parent task name
- The VRAM configuration (auto-detected or manual)
- Any context the sub-task needs from the parent
9. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently.
10. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention).
9. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration (see VRAM Config Propagation section).
10. **Create a PARENT_SPEC.md** file for each sub-task with:
- **The sub-task's own scope/acceptance criteria from the DECOMPOSITION.md** — this is the most important part, as it tells the sub-task's Research phase what the sub-task is responsible for.
- A reference to the parent task name.
- The VRAM configuration (auto-detected or manual).
- **Do NOT include the parent task's full SPEC.md** — this can cause circular references if the parent's SPEC.md references the sub-task's SPEC.md files. Instead, include only the sub-task's own scope from the DECOMPOSITION.md.
11. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently.
12. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention).
### Sub-Task Lifecycle
@@ -352,41 +379,55 @@ When creating a sub-task folder, the Orchestrator should create a `VRAM_CONFIG.m
# VRAM Configuration for this sub-task
- **Auto-detect**: Yes/No
- **Target VRAM context**: {from detection script or config.md override}
- **Headroom**: {from detection script or config.md override}
- **Max peak context per sub-task**: {from detection script or config.md override}
- **Target VRAM context**: {from detection script or config.md override}k tokens (e.g., "16k")
- **Headroom**: {from detection script or config.md override}% (e.g., "25%")
- **Max peak context per sub-task**: {from detection script or config.md override}k tokens (e.g., "12k")
- **GPU VRAM detected**: {value}GB (or "None")
- **RAM detected**: {value}GB
- **Model context window**: {value}k tokens (or "Unknown")
- **Framework overhead**: ~{value} tokens
- **This sub-task's estimated peak context**: {from DECOMPOSITION.md}
- **This sub-task's estimated peak context**: {from DECOMPOSITION.md}k tokens (e.g., "10k")
- **Fits within VRAM**: Yes/No
```
This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly.
**Important**: The units must be clarified. `Target VRAM context` uses "k" units (e.g., "16k" for 16k tokens), not raw values like "16". `Max peak context per sub-task` uses the raw value from the detection script (e.g., "12000" for 12000 tokens). `This sub-task's estimated peak context` uses "k" units (e.g., "10k" for 10k tokens). This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly.
### Sub-Task Parent Specification
When creating a sub-task folder, the Orchestrator should also create a `PARENT_SPEC.md` file that contains:
- The parent task's SPEC.md content
- **The sub-task's own scope/acceptance criteria from the DECOMPOSITION.md** — this is the most important part, as it tells the sub-task's Research phase what the sub-task is responsible for.
- A reference to the parent task name
- Any context the sub-task needs from the parent
- The VRAM configuration (auto-detected or manual)
- **Do NOT include the parent task's full SPEC.md** — this can cause circular references if the parent's SPEC.md references the sub-task's SPEC.md files. Instead, include only the sub-task's own scope from the DECOMPOSITION.md.
This ensures sub-tasks have all the information they need to implement their portion of the parent spec.
This ensures sub-tasks have all the information they need to implement their portion of the parent spec without causing circular references.
### Sub-Task Dependencies and Wave Management
Sub-tasks that are in the same wave (can run in parallel) should be driven through the lifecycle in parallel. Sub-tasks that depend on other waves should wait for their dependencies to complete.
**Wave Enforcement**: The Orchestrator MUST check sub-task dependencies and only start driving Wave 2 sub-tasks when all Wave 1 sub-tasks are in terminal state. If a sub-task in Wave 2 depends on a sub-task in Wave 1, the Orchestrator should not drive the Wave 2 sub-task until the Wave 1 sub-task is in terminal state. This ensures that sub-tasks are driven in the correct order and that the parent task does not complete prematurely.
### Sub-Task Completion and Parent Task
When a sub-task reaches a terminal state:
- If the sub-task **PASSes**: The parent task can move to the next wave of sub-tasks.
- If a sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator pauses and reports human intervention is required. The parent task is stuck until the failing sub-task is resolved.
When a sub-task reaches a terminal state, the Orchestrator MUST verify that the sub-task has a `VERDICT.md` before evaluating the verdict. If the sub-task does not have a `VERDICT.md`, it is NOT in a terminal state.
When ALL sub-tasks are in terminal state:
When a sub-task reaches a terminal state:
- If the sub-task **PASSes** (VERDICT.md with PASS): The parent task can move to the next wave of sub-tasks.
- If a sub-task **FAILs or NEEDS_REVIEW** (VERDICT.md with FAIL/NEEDS_REVIEW):
- The Orchestrator pauses and reports human intervention is required.
- **In Autopilot Mode**: The Orchestrator does NOT auto-create fix/review tasks — it should pause and report that human intervention is required. The user must decide whether to create fix tasks and how to proceed.
- **In Manual Mode**: The Orchestrator MUST create a new task for the failing sub-task:
- Task name: `{parent-task-name}-fix-{sub-task-name}` (e.g., `add-user-auth-fix-auth-gateway`)
- The Orchestrator creates the folder with an empty `IMPLEMENTATION.md`
- The task starts at the **Bug Find** phase
- The Orchestrator copies the sub-task's `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` (if they exist) into the new task folder
- If the sub-task failed during the Referee phase (before producing BUG_REPORT.md and ADVERSARIAL_BUG_REPORT.md), the Orchestrator still creates the fix task but only copies the artifacts that exist (BUG_REPORT.md if it exists, ADVERSARIAL_BUG_REPORT.md if it exists).
When ALL sub-tasks are in terminal state (each sub-task has a VERDICT.md):
- **Check for orphaned sub-tasks**: Before checking completion, verify that each sub-task is still referenced in the DECOMPOSITION.md. If a sub-task is orphaned (no longer in the DECOMPOSITION.md), remove it from the parent's completion check.
- If ALL sub-tasks **PASS**: The parent task is **Complete**.
- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks and reports human intervention is required.
- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks (as described above) and reports human intervention is required.
### Sub-Task Verdict Reporting
@@ -395,4 +436,20 @@ When a sub-task reaches the Referee phase, the VERDICT.md should include:
- A reference to the parent task name
- Any findings that affect the parent task
The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status.
### Sub-Task Verdict Aggregation
The Orchestrator MUST aggregate sub-task verdicts when reporting the parent task's status:
- If ANY sub-task **FAILs or NEEDS_REVIEW**, the parent task should be marked as **Human Intervention** regardless of whether the parent's own VERDICT.md says PASS.
- The parent task's status should include a summary of all sub-task verdicts:
- PASS: {count}
- FAIL: {count}
- NEEDS_REVIEW: {count}
- If the parent task has a VERDICT.md with PASS, but a sub-task has a VERDICT.md with FAIL, the Orchestrator should report: "⚠️ **HUMAN INTERVENTION REQUIRED**: Parent task VERDICT.md says PASS, but sub-task {sub-task-name} has VERDICT.md with FAIL. Create a fix task for the failing sub-task."
### Sub-Task Tie-Breaks
If a sub-task has "Tasks for Review / Tie-Breaks" in its VERDICT.md, the Orchestrator should create tie-break tasks for the sub-task:
- Task name: `{parent-task-name}-tiebreak-{sub-task-name}` (e.g., `add-user-auth-tiebreak-auth-gateway`)
- The Orchestrator creates the folder with an empty `IMPLEMENTATION.md`
- The task starts at the **Research** phase
- The Orchestrator copies the sub-task's `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` (if they exist) into the new task folder
+21 -9
View File
@@ -7,15 +7,15 @@ This file defines the linear progression of a task in the agent-framework. The O
| Current State | Signal (Artifact) | Next Phase | Action |
| :--- | :--- | :--- | :--- |
| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` |
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement | Generate `DECOMPOSITION.md` or `DESIGN.md` or code |
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research | Orchestrator creates sub-task folders |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
| **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests |
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review | Generate `DOC_REVIEW.md` |
| **Doc Review** | Has `DOC_REVIEW.md` | Referee | Generate `VERDICT.md` |
| **Referee** | Has `VERDICT.md` | Complete / Review | Finalize or request user intervention |
| **Research** | Has `SPEC.md` (non-empty) | Decomposition (optional) or Design (optional) or Implement | Generate `DECOMPOSITION.md` or `DESIGN.md` or code |
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` (both non-empty) | Sub-task Research | Orchestrator creates sub-task folders |
| **Design** | Has `DESIGN.md` (non-empty) | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
| **Test Design** | Has `TEST_PLAN.md` (non-empty) | Implement | Generate code and tests |
| **Implementation** | Has `IMPLEMENTATION.md` (non-empty) | Bug Find | Generate `BUG_REPORT.md` |
| **Bug Find** | Has `BUG_REPORT.md` (non-empty) | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` (non-empty) | Doc Review | Generate `DOC_REVIEW.md` |
| **Doc Review** | Has `DOC_REVIEW.md` (non-empty) | Referee | Generate `VERDICT.md` |
| **Referee** | Has `VERDICT.md` (non-empty) | Complete / Review | Finalize or request user intervention |
## Task Creation (Orchestrator Responsibility)
@@ -54,6 +54,18 @@ When the Orchestrator detects a `VERDICT.md` with `FAIL` or `NEEDS_REVIEW`, the
- The task starts at the **Research** phase (the tie-break may require spec changes)
- The Orchestrator copies the original `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` into the new task folder
4. **From sub-task `FAIL` verdict**: For each sub-task that FAILs or NEEDS_REVIEW, create a new task:
- Task name: `{parent-task-name}-fix-{sub-task-name}` (e.g., `add-user-auth-fix-auth-gateway`)
- The Orchestrator creates the folder with an empty `IMPLEMENTATION.md`
- The task starts at the **Bug Find** phase
- The Orchestrator copies the sub-task's `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` (if they exist) into the new task folder
5. **From sub-task "Tasks for Review / Tie-Breaks"**: For each sub-task that has tie-breaks, create a new task:
- Task name: `{parent-task-name}-tiebreak-{sub-task-name}` (e.g., `add-user-auth-tiebreak-auth-gateway`)
- The Orchestrator creates the folder with an empty `IMPLEMENTATION.md`
- The task starts at the **Research** phase
- The Orchestrator copies the sub-task's `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` (if they exist) into the new task folder
**In Autopilot Mode**: The Orchestrator should NOT auto-create fix/review/tiebreak tasks — it should pause and report that human intervention is required. The user must decide whether to create fix tasks and how to proceed.
**Note**: When a task starts at the **Bug Find** phase (fix tasks), the Orchestrator skips the Research phase. The implementation agent should first review the existing bugs and spec before fixing them. The Orchestrator signals this by checking for `BUG_REPORT.md` and `ADVERSARIAL_BUG_REPORT.md` in the new task folder.