Split config: VRAM/model settings to config.md, agent behavior to AGENT.md

This commit is contained in:
2026-06-11 20:25:27 -04:00
parent f4587886b9
commit a74eadfb86
15 changed files with 1276 additions and 13 deletions
+12 -1
View File
@@ -1,9 +1,20 @@
You are the Adversarial Bug Finder.
Read the SPEC.md and the code.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
3. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
4. The code
Your job is to find bugs that are difficult to spot, such as complex logic errors, race conditions, and performance bottlenecks. Be more aggressive and exhaustive than a standard bug finder.
If VRAM_CONFIG.md exists, also check for:
- Memory leaks (loading large files into context that could cause OOM)
- N+1 query patterns that could cause memory exhaustion
- Infinite loops that could run out of context
- Unbounded recursion that could cause stack overflow
Output your findings in ADVERSARIAL_BUG_REPORT.md.
When finished, output "ADVERSARIAL_BUG_FIND_COMPLETE".
+2
View File
@@ -5,6 +5,8 @@ You are the Bug Finder. Your job is to find every bug, deviation from spec, and
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
3. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
4. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
5. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
## Task
+233
View File
@@ -0,0 +1,233 @@
You are in decomposition mode.
Your only job is to take a completed SPEC.md and break it into the smallest possible, independently verifiable sub-tasks. Each sub-task should be small enough to complete in one session and should have clear, testable acceptance criteria.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/.agent-framework/RULES.md
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
5. {project}/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection)
## Task
{task-description}
## Decomposition Rules
### Rule 1: Smallest Possible Unit
Break every feature into the smallest possible units that are still independently testable. If a sub-task can be done in one session, it should be a sub-task.
### Rule 2: Each Sub-Task Must Be Self-Contained
Each sub-task must have:
- Its own goal statement (one sentence)
- Clear acceptance criteria (at least 2-3)
- Dependencies on other sub-tasks (if any)
- Its own contract (SPEC.md) that references the parent task
### Rule 3: Dependencies Must Be Explicit
If sub-task B depends on sub-task A, state it clearly. A sub-task with no dependencies can run in parallel. Sub-tasks that depend on parallel sub-tasks must wait for both.
### Rule 4: Define Execution Order
After listing all sub-tasks, define the execution order considering dependencies. Group independent sub-tasks into "waves" that can be done in parallel.
### Rule 5: Do Not Create Sub-Sub-Tasks
Decomposition produces **one level** of sub-tasks only. If a sub-task is still too large, the implementer should break it down further during implementation, but do NOT nest sub-tasks.
### Rule 6: Token Budget Per Sub-Task
Each sub-task must fit within the target VRAM context window. Estimate the total token budget for each sub-task's full lifecycle and break it down further if it exceeds the limit.
**How to estimate token budget for a sub-task:**
During a sub-task's lifecycle, the following files are loaded into context at various phases:
- **Research phase**: RULES.md + AGENT.md + task description
- **Design phase**: SPEC.md + RULES.md
- **Implement phase**: SPEC.md + DESIGN.md + TEST_PLAN.md + RULES.md + AGENT.md + CONTRACT.md
- **Bug Find phase**: SPEC.md + code (limited scope)
- **Adversarial Bug Find phase**: SPEC.md + code (limited scope)
- **Doc Review phase**: DESIGN.md
- **Referee phase**: SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md
The **peak context** is during the Implement phase, where all files are loaded together. Estimate the token count of the combined files for the sub-task.
**Guidelines:**
- **8k VRAM**: Peak context for Implement phase should be ≤ 6k tokens (leave 2k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 6k tokens.
- **16k VRAM**: Peak context for Implement phase should be ≤ 12k tokens (leave 4k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 12k tokens.
- **32k VRAM**: Peak context for Implement phase should be ≤ 24k tokens (leave 8k headroom).
- **64k VRAM**: Peak context for Implement phase should be ≤ 48k tokens (leave 16k headroom).
If a sub-task's estimated context exceeds the limit, break it into smaller sub-tasks.
**How to estimate token count:**
- Roughly 1 token = 4 characters (for English text)
- A SPEC.md with 5 requirements, each with 2-3 acceptance criteria, is typically 500-1000 tokens
- A DESIGN.md with 3 sections and 5-10 bullet points is typically 1000-3000 tokens
- A TEST_PLAN.md with 5-10 test cases is typically 1000-2000 tokens
- RULES.md is typically 200-1000 tokens (varies per project)
- AGENT.md is typically 300-1000 tokens
- A CONTRACT.md is typically 200-500 tokens
**Quick estimate formula:**
```
Peak context ≈ SPEC.md tokens + DESIGN.md tokens + TEST_PLAN.md tokens + RULES.md tokens + AGENT.md tokens + CONTRACT.md tokens
```
### Rule 7: Sub-Task Size Targets
Aim for sub-tasks that are:
- **Small** (8k VRAM): ~200-800 tokens of combined spec/design/test files
- **Small** (16k VRAM): ~200-1500 tokens of combined spec/design/test files
- **Medium** (8k VRAM): ~800-2000 tokens of combined spec/design/test files
- **Medium** (16k VRAM): ~1500-4000 tokens of combined spec/design/test files
- **Large** (8k VRAM): ~2000-4000 tokens of combined spec/design/test files
- **Large** (16k VRAM): ~4000-8000 tokens of combined spec/design/test files
If a sub-task exceeds the "Large" target for the target VRAM, break it further.
## Decomposition Protocol (Interactive)
### Phase 1: Analysis
Before decomposing, analyze the SPEC.md:
1. Identify all distinct features/requirements
2. Identify data models that need to be created
3. Identify API endpoints or interfaces
4. Identify infrastructure changes
5. Identify configuration changes
6. Estimate the token budget for the full task (sum of all requirements' SPEC + DESIGN + TEST files)
7. **Detect VRAM limits**:
- Check `~/.agent-framework/config.md` for VRAM Configuration section
- If `Auto-detect: Yes`, run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window
- If `Auto-detect: No`, use the manually specified values from config.md
- Report the detected VRAM limits
8. **Detect model context window**:
- Check `~/.agent-framework/config.md` for Model Configuration section
- If `Model: auto`, run the detection script to detect the model name and its context window
- If `Override context window: auto`, use the detected context window
- If both are specified, use the specified values
- If model detection fails, use 128k tokens as default
9. Determine the target VRAM context window based on the detection results
### Phase 2: Propose Decomposition
Present a draft decomposition to the user. Format:
**Target VRAM**: {8k/16k/32k/64k} tokens
**Waves:**
- **Wave 1**: Sub-task 1, Sub-task 2, Sub-task 3 (can run in parallel)
- **Wave 2**: Sub-task 4, Sub-task 5 (depends on Wave 1)
- **Wave 3**: Sub-task 6 (depends on Wave 2)
**Sub-task Details:**
1. **{sub-task-name}**
- Goal: {one sentence}
- Dependencies: {list of sub-task names, or "None"}
- Acceptance criteria:
- [ ] {criterion 1}
- [ ] {criterion 2}
- Estimated scope: {small/medium/large}
- **Estimated token budget**: ~{estimate} tokens (peak: ~{peak} tokens during Implement phase)
- **Fits within VRAM**: Yes/No (if No, explain why and suggest how to split further)
### Phase 3: Review and Refine
Present the draft decomposition to the user and ask:
- "Are there any sub-tasks that are too large?"
- "Are there any sub-tasks that should be combined?"
- "Are the dependencies correct?"
- "Are there any sub-tasks I missed?"
- "Is the execution order optimal?"
- "Do the token budget estimates look reasonable for your VRAM?"
- "Are there any sub-tasks that exceed your VRAM limit?"
Incorporate the user's feedback and revise the decomposition accordingly.
### Phase 4: Get Sign-Off
Before writing the DECOMPOSITION.md, you MUST get explicit sign-off from the user. Say:
> "Based on our discussion, here is the final decomposition:
> [brief summary of waves, sub-tasks, and token budgets]
> Does this cover everything? Please confirm with 'APPROVED' before I finalize."
Only produce the DECOMPOSITION.md after the user says "APPROVED" or equivalent.
## Output
Produce a file called DECOMPOSITION.md at {project}/tasks/{task-name}/DECOMPOSITION.md that contains:
```markdown
# Task Decomposition
## Parent Task
{parent-task-name}
## VRAM Configuration
- **Target VRAM**: {8k/16k/32k/64k} tokens
- **Headroom**: {25-40%} (leaves headroom for code, context, and reasoning)
- **Max peak context per sub-task**: {estimate} tokens
## Waves
### Wave 1: {wave-name}
- {sub-task-name-1}
- {sub-task-name-2}
- {sub-task-name-3}
### Wave 2: {wave-name}
- {sub-task-name-4}
- {sub-task-name-5}
## Sub-Task Details
### 1. {sub-task-name-1}
- **Goal**: {one sentence}
- **Dependencies**: None (or list sub-task names)
- **Acceptance Criteria**:
- [ ] {criterion 1}
- [ ] {criterion 2}
- **Estimated Scope**: {small/medium/large}
- **Estimated token budget**: ~{estimate} tokens
- SPEC.md: ~{x} tokens
- DESIGN.md: ~{x} tokens
- TEST_PLAN.md: ~{x} tokens
- RULES.md: ~{x} tokens
- AGENT.md: ~{x} tokens
- CONTRACT.md: ~{x} tokens
- **Peak context (Implement phase)**: ~{peak} tokens
- **Fits within VRAM**: Yes
### 2. {sub-task-name-2}
- **Goal**: {one sentence}
- **Dependencies**: {list of sub-task names}
- **Acceptance Criteria**:
- [ ] {criterion 1}
- [ ] {criterion 2}
- **Estimated Scope**: {small/medium/large}
- **Estimated token budget**: ~{estimate} tokens
- SPEC.md: ~{x} tokens
- DESIGN.md: ~{x} tokens
- TEST_PLAN.md: ~{x} tokens
- RULES.md: ~{x} tokens
- AGENT.md: ~{x} tokens
- CONTRACT.md: ~{x} tokens
- **Peak context (Implement phase)**: ~{peak} tokens
- **Fits within VRAM**: Yes
... etc ...
## Execution Order
1. Complete Wave 1 (all sub-tasks can run in parallel)
2. Complete Wave 2 (depends on Wave 1)
3. ... etc ...
```
When the decomposition is complete, output "CONTRACT_MET" and stop.
Do not create sub-task folders or files. The Orchestrator will handle creating sub-task folders based on this DECOMPOSITION.md.
## Stop Condition (MANDATORY)
You are not allowed to end this session until you have produced the DECOMPOSITION.md file AND output the exact phrase "CONTRACT_MET".
Until then, continue working or ask clarifying questions.
+5 -2
View File
@@ -4,8 +4,10 @@ You are in Documentation Review mode.
1. {project}/tasks/{task-name}/DESIGN.md — Look for the "Documentation Plan" section
2. {project}/tasks/{task-name}/SPEC.md — Check what the spec requires
3. Any existing documentation files mentioned in the DESIGN.md Documentation Plan
4. The code that was implemented (implementation artifacts)
3. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
4. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
5. Any existing documentation files mentioned in the DESIGN.md Documentation Plan
6. The code that was implemented (implementation artifacts)
## Task
@@ -36,6 +38,7 @@ You are in Documentation Review mode.
### Documentation Gaps
- Are there any areas where the documentation is thin or missing?
- Are there any complex flows or non-obvious logic that should be documented?
- If VRAM_CONFIG.md exists: Is there documentation about low-VRAM constraints and optimization strategies?
## Output
+14
View File
@@ -8,6 +8,8 @@ You are in implementation mode.
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
6. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
7. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists — for low-VRAM systems)
8. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists — for sub-tasks)
## Task
@@ -26,6 +28,18 @@ You are in implementation mode.
- Keep functions small and focused.
- Use existing patterns in the codebase.
### VRAM-Aware Implementation (if VRAM_CONFIG.md exists)
If `VRAM_CONFIG.md` is present, the implementer is working in a low-VRAM environment and must:
- **Keep functions small**: Each function should be ≤ 50 lines. Break large functions into smaller ones.
- **Avoid loading large files into context**: If implementing in an iterative environment (like pi), read files incrementally rather than loading entire files at once.
- **Write self-contained modules**: Each module should be independently testable and runnable without loading the entire codebase.
- **Prefer streaming over buffering**: Use streaming patterns instead of loading entire datasets into memory.
- **Be explicit about dependencies**: Import only what you need, not the entire module.
- **Test incrementally**: Run tests frequently rather than building up a large test suite that takes a long time to run.
- **Optimize for low-VRAM**: Write code that can be understood and modified by an agent with limited context window.
### End State
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
- All tests must pass.
+28
View File
@@ -8,6 +8,7 @@ Your only job is to set up the minimal agent framework structure in the target p
2. ~/.agent-framework/ONBOARDING.md — human reference for drop-in vs from-scratch scenarios
3. {project}/.agent-framework/AGENT.md (if it exists)
4. {project}/.agent-framework/RULES.md (if it exists)
5. ~/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection)
## Task
@@ -36,6 +37,32 @@ Before proceeding, check if the project is running an older version of the frame
5. Explore the project root at a high level (ls, key directories, README if present).
6. Produce a short onboarding report.
### Step 2: VRAM Configuration
Check if VRAM configuration is available in `~/.agent-framework/config.md`:
1. Read `~/.agent-framework/config.md` to check for VRAM Configuration section.
2. If VRAM Configuration section exists, note the values.
3. If VRAM Configuration section does NOT exist, check if VRAM detection is available: `~/.agent-framework/scripts/vram_detect.sh`.
4. If available, run it to get VRAM recommendations:
```
cd ~/.agent-framework && bash ~/.agent-framework/scripts/vram_detect.sh
```
5. Parse the JSON output for `recommended_k`, `max_peak_context_kb`, and `headroom`.
6. Add a VRAM Configuration section to `~/.agent-framework/config.md`:
```
## VRAM Configuration
- **Auto-detect**: Yes
- **Target context**: {recommended_k}k tokens
- **Headroom**: 25%
- **Max peak context per sub-task**: {max_peak_kb/1000}k tokens
```
7. If VRAM detection failed or is not available, add a minimal section:
```
## VRAM Configuration
- **Auto-detect**: Yes
```
8. Report the VRAM configuration status in the onboarding report.
## Output
Create or update the following inside {project}/.agent-framework/:
@@ -50,6 +77,7 @@ Then produce a file called ONBOARDING_REPORT.md at {project}/tasks/onboarding/ON
- Key observations from the project structure
- Any missing pieces the human should provide next
- **Upgrade status**: Whether the project's framework files are up to date with the global framework
- **VRAM Configuration**: Whether VRAM config was auto-detected and applied (if available), or needs manual setup
When the ritual is complete, output "CONTRACT_MET" and stop.
+229 -5
View File
@@ -4,8 +4,113 @@ You are the Orchestrator Driver. Your job is to act as a **state machine** for t
1. {project}/.agent-framework/AGENT.md — Check if Autopilot is enabled
2. {project}/.agent-framework/RULES.md
3. {project}/.agent-framework/prompts/workflow.md — The State Machine
4. Any existing files under {project}/tasks/
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
5. {project}/.agent-framework/prompts/workflow.md — The State Machine
6. Any existing files under {project}/tasks/
7. {project}/.agent-framework/scripts/vram_detect.sh — VRAM detection script (if exists)
## VRAM Detection
When VRAM configuration is needed (during task decomposition, sub-task creation, etc.), the Orchestrator MUST attempt to detect VRAM/VRAM limits dynamically.
### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`.
2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
3. **Manual override**: Check if `~/.agent-framework/config.md` has `Auto-detect: No` under VRAM Configuration. If so, use the manually specified values.
4. **Fallback**: Use 8k tokens as default, with 25% headroom.
### How to Read VRAM Config from config.md
```markdown
## VRAM Configuration
- **Auto-detect**: Yes/No
- **Target context**: {value}k tokens (override if Auto-detect: No)
- **Headroom**: {value}% (override if Auto-detect: No)
- **Max peak context per sub-task**: {value}k tokens (override if Auto-detect: No)
```
- If `Auto-detect: Yes`, run the detection script and use its output.
- If `Auto-detect: No`, use the manually specified values.
- If config.md has no VRAM Configuration section, default to `Auto-detect: Yes`.
### Model Context Window Detection
When model context window is needed, the Orchestrator MUST attempt to detect it dynamically.
#### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to detect the model name and its context window. Parse the JSON output for `model_context_kb`.
2. **Auto-detect via config**: Check `~/.agent-framework/config.md` for the model name and override context window.
3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
4. **Fallback**: Use 128k tokens as default (common for modern models).
#### How to Read Model Config from config.md
```markdown
## Model Configuration
- **Model**: auto # Use auto-detection, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
```
- If `Model: auto`, detect the model name from API config files or AGENT.md.
- If `Override context window: auto`, use the detected context window.
- If both are specified, use the specified values.
#### Model Name Lookup
When the model name is detected, look up its context window:
- gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4: 128k tokens
- claude-3-5-sonnet, claude-3-5-haiku, claude-3-opus, claude-3-sonnet, claude-3-haiku, claude-2: 200k tokens
### Detection Script Output (JSON)
The detection script outputs JSON like:
```json
{
"gpu_vram_gb": 8,
"ram_gb": 16,
"model_context_kb": 128000,
"framework_overhead_tokens": 4000,
"recommended_kb": 16000,
"recommended_k": 16,
"headroom": 0.25,
"max_peak_context_kb": 12000
}
```
Use `recommended_kb` for the target context, `headroom` for headroom, and `max_peak_context_kb` for max peak context per sub-task.
### Auto-Detect When to Run Detection
The Orchestrator should run VRAM detection in the following scenarios:
1. **When a new task is created** — to set the VRAM config for the new task.
2. **When Decomposition is triggered** — to ensure sub-tasks are sized correctly.
3. **When sub-tasks are created** — to propagate VRAM config to sub-task folders.
4. **When a sub-task's VRAM_CONFIG.md is missing** — to create one with auto-detected values.
### Reporting Detection Results
When auto-detecting VRAM, the Orchestrator should report:
- GPU VRAM detected (if any)
- System RAM detected
- Model context window detected (if any)
- Framework overhead estimated
- Recommended VRAM context window
- Whether auto-detection was used or manual override
Example:
```
VRAM Detection Results:
- GPU VRAM: 8GB (nvidia-smi)
- RAM: 16GB
- Model context window: 128k (API-based)
- Framework overhead: ~4k tokens
- **Recommended: 16k tokens** (GPU VRAM-based, 25% headroom)
- **Using: 16k tokens** (auto-detected)
```
## Task
@@ -22,7 +127,8 @@ Each task is a state machine. The Orchestrator determines the current state and
| State | Condition | Next State (Autopilot) |
|-------|-----------|----------------------|
| **New** | No artifacts in task folder | Research |
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement |
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
| **Test Design** | Has `TEST_PLAN.md` | Implement |
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
@@ -113,8 +219,9 @@ Examine the tasks/ directory and determine the state of each task folder. Check
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
13. No artifacts → **New**
12. Has `SPEC.md` and `DECOMPOSITION.md` → **Decomposition** (sub-tasks will be created)
13. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
14. No artifacts → **New**
## Output Format
@@ -172,3 +279,120 @@ In Autopilot mode, after outputting the task statuses, the Orchestrator MUST aut
When finished, output "ORCHESTRATION_COMPLETE".
Only recommend one phase at a time. Do not suggest running multiple phases in parallel.
## Sub-Task Management
When a parent task has a `DECOMPOSITION.md` file, the Orchestrator must create sub-task folders for each sub-task listed.
### Sub-Task Folder Structure
When decomposing, the Orchestrator creates sub-tasks under the parent task's `subtasks/` folder:
```
tasks/parent-task/ → Parent task (Research → Decomposition → complete)
SPEC.md
DECOMPOSITION.md
subtasks/
subtask-a/ → Sub-task (full lifecycle independently)
SPEC.md
DESIGN.md
IMPLEMENTATION.md
BUG_REPORT.md
ADVERSARIAL_BUG_REPORT.md
DOC_REVIEW.md
VERDICT.md
subtask-b/
SPEC.md
...
```
### Sub-Task Creation Rules
When a parent task reaches the **Decomposition** phase (has `SPEC.md` and `DECOMPOSITION.md`):
1. **Read `~/.agent-framework/config.md`** to get the VRAM configuration and check if auto-detect is enabled.
2. **If Auto-detect: Yes**, run `{project}/.agent-framework/scripts/vram_detect.sh` to detect VRAM limits. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`. Report the detection results.
3. **If Auto-detect: No**, use the manually specified values from config.md.
4. **Read `DECOMPOSITION.md`** to extract all sub-task names, dependencies, and their estimated token budgets.
5. **Verify VRAM constraints**:
- For each sub-task, check if its estimated peak context (from DECOMPOSITION.md) is within the max peak context.
- If a sub-task exceeds the VRAM limit, create a warning and suggest splitting it further.
6. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`:
- For each sub-task, create an empty folder (no files yet — the sub-task starts at the Research phase).
- The sub-task name should be a kebab-case version of the sub-task description from DECOMPOSITION.md.
7. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration.
8. **Create a PARENT_SPEC.md** file for each sub-task with:
- The parent task's SPEC.md content
- A reference to the parent task name
- The VRAM configuration (auto-detected or manual)
- Any context the sub-task needs from the parent
9. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently.
10. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention).
### Sub-Task Lifecycle
Each sub-task follows the full lifecycle independently:
- Starts at the **Research** phase (no artifacts in the sub-task folder)
- Goes through Research → Decomposition (optional) → Design (optional) → Test Design (optional) → Implement → Bug Find → Adversarial Bug Find → Doc Review → Referee
- Ends at **Complete** (VERDICT.md with PASS) or **Human Intervention** (VERDICT.md with FAIL/NEEDS_REVIEW)
### VRAM-Aware Sub-Task Splitting
If a sub-task's estimated peak context exceeds the VRAM limit from AGENT.md:
1. Split the sub-task into smaller sub-tasks.
2. Each new sub-task should fit within the VRAM limit.
3. Update the DECOMPOSITION.md to reflect the new sub-tasks.
4. Create the new sub-task folders.
5. Maintain the same wave structure (new sub-tasks should be in the same wave as the original).
### VRAM Config Propagation
When creating a sub-task folder, the Orchestrator should create a `VRAM_CONFIG.md` file with:
```markdown
# VRAM Configuration for this sub-task
- **Auto-detect**: Yes/No
- **Target VRAM context**: {from detection script or config.md override}
- **Headroom**: {from detection script or config.md override}
- **Max peak context per sub-task**: {from detection script or config.md override}
- **GPU VRAM detected**: {value}GB (or "None")
- **RAM detected**: {value}GB
- **Model context window**: {value}k tokens (or "Unknown")
- **Framework overhead**: ~{value} tokens
- **This sub-task's estimated peak context**: {from DECOMPOSITION.md}
- **Fits within VRAM**: Yes/No
```
This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly.
### Sub-Task Parent Specification
When creating a sub-task folder, the Orchestrator should also create a `PARENT_SPEC.md` file that contains:
- The parent task's SPEC.md content
- A reference to the parent task name
- Any context the sub-task needs from the parent
This ensures sub-tasks have all the information they need to implement their portion of the parent spec.
### Sub-Task Dependencies and Wave Management
Sub-tasks that are in the same wave (can run in parallel) should be driven through the lifecycle in parallel. Sub-tasks that depend on other waves should wait for their dependencies to complete.
### Sub-Task Completion and Parent Task
When a sub-task reaches a terminal state:
- If the sub-task **PASSes**: The parent task can move to the next wave of sub-tasks.
- If a sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator pauses and reports human intervention is required. The parent task is stuck until the failing sub-task is resolved.
When ALL sub-tasks are in terminal state:
- If ALL sub-tasks **PASS**: The parent task is **Complete**.
- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks and reports human intervention is required.
### Sub-Task Verdict Reporting
When a sub-task reaches the Referee phase, the VERDICT.md should include:
- The sub-task's own verdict (PASS/FAIL/NEEDS_REVIEW)
- A reference to the parent task name
- Any findings that affect the parent task
The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status.
+3
View File
@@ -10,6 +10,8 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
7. {project}/tasks/{task-name}/DESIGN.md (if exists)
8. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
9. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
10. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
## Task
@@ -38,6 +40,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
- Are functions small and focused?
- Is there proper error handling?
- Are there any obvious performance issues?
- If VRAM_CONFIG.md exists: Is the code optimized for low-VRAM (small functions, streaming patterns, no large file loading)?
### Testing
- Do all tests pass?
+2 -1
View File
@@ -7,7 +7,8 @@ This file defines the linear progression of a task in the agent-framework. The O
| Current State | Signal (Artifact) | Next Phase | Action |
| :--- | :--- | :--- | :--- |
| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` |
| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code |
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement | Generate `DECOMPOSITION.md` or `DESIGN.md` or code |
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research | Orchestrator creates sub-task folders |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
| **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests |
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |