Split config: VRAM/model settings to config.md, agent behavior to AGENT.md
This commit is contained in:
+229
-5
@@ -4,8 +4,113 @@ You are the Orchestrator Driver. Your job is to act as a **state machine** for t
|
||||
|
||||
1. {project}/.agent-framework/AGENT.md — Check if Autopilot is enabled
|
||||
2. {project}/.agent-framework/RULES.md
|
||||
3. {project}/.agent-framework/prompts/workflow.md — The State Machine
|
||||
4. Any existing files under {project}/tasks/
|
||||
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
|
||||
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
|
||||
5. {project}/.agent-framework/prompts/workflow.md — The State Machine
|
||||
6. Any existing files under {project}/tasks/
|
||||
7. {project}/.agent-framework/scripts/vram_detect.sh — VRAM detection script (if exists)
|
||||
|
||||
## VRAM Detection
|
||||
|
||||
When VRAM configuration is needed (during task decomposition, sub-task creation, etc.), the Orchestrator MUST attempt to detect VRAM/VRAM limits dynamically.
|
||||
|
||||
### Detection Priority
|
||||
|
||||
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`.
|
||||
2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
|
||||
3. **Manual override**: Check if `~/.agent-framework/config.md` has `Auto-detect: No` under VRAM Configuration. If so, use the manually specified values.
|
||||
4. **Fallback**: Use 8k tokens as default, with 25% headroom.
|
||||
|
||||
### How to Read VRAM Config from config.md
|
||||
|
||||
```markdown
|
||||
## VRAM Configuration
|
||||
- **Auto-detect**: Yes/No
|
||||
- **Target context**: {value}k tokens (override if Auto-detect: No)
|
||||
- **Headroom**: {value}% (override if Auto-detect: No)
|
||||
- **Max peak context per sub-task**: {value}k tokens (override if Auto-detect: No)
|
||||
```
|
||||
|
||||
- If `Auto-detect: Yes`, run the detection script and use its output.
|
||||
- If `Auto-detect: No`, use the manually specified values.
|
||||
- If config.md has no VRAM Configuration section, default to `Auto-detect: Yes`.
|
||||
|
||||
### Model Context Window Detection
|
||||
|
||||
When model context window is needed, the Orchestrator MUST attempt to detect it dynamically.
|
||||
|
||||
#### Detection Priority
|
||||
|
||||
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to detect the model name and its context window. Parse the JSON output for `model_context_kb`.
|
||||
2. **Auto-detect via config**: Check `~/.agent-framework/config.md` for the model name and override context window.
|
||||
3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
|
||||
4. **Fallback**: Use 128k tokens as default (common for modern models).
|
||||
|
||||
#### How to Read Model Config from config.md
|
||||
|
||||
```markdown
|
||||
## Model Configuration
|
||||
- **Model**: auto # Use auto-detection, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
|
||||
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
|
||||
```
|
||||
|
||||
- If `Model: auto`, detect the model name from API config files or AGENT.md.
|
||||
- If `Override context window: auto`, use the detected context window.
|
||||
- If both are specified, use the specified values.
|
||||
|
||||
#### Model Name Lookup
|
||||
|
||||
When the model name is detected, look up its context window:
|
||||
- gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4: 128k tokens
|
||||
- claude-3-5-sonnet, claude-3-5-haiku, claude-3-opus, claude-3-sonnet, claude-3-haiku, claude-2: 200k tokens
|
||||
|
||||
### Detection Script Output (JSON)
|
||||
|
||||
The detection script outputs JSON like:
|
||||
```json
|
||||
{
|
||||
"gpu_vram_gb": 8,
|
||||
"ram_gb": 16,
|
||||
"model_context_kb": 128000,
|
||||
"framework_overhead_tokens": 4000,
|
||||
"recommended_kb": 16000,
|
||||
"recommended_k": 16,
|
||||
"headroom": 0.25,
|
||||
"max_peak_context_kb": 12000
|
||||
}
|
||||
```
|
||||
|
||||
Use `recommended_kb` for the target context, `headroom` for headroom, and `max_peak_context_kb` for max peak context per sub-task.
|
||||
|
||||
### Auto-Detect When to Run Detection
|
||||
|
||||
The Orchestrator should run VRAM detection in the following scenarios:
|
||||
|
||||
1. **When a new task is created** — to set the VRAM config for the new task.
|
||||
2. **When Decomposition is triggered** — to ensure sub-tasks are sized correctly.
|
||||
3. **When sub-tasks are created** — to propagate VRAM config to sub-task folders.
|
||||
4. **When a sub-task's VRAM_CONFIG.md is missing** — to create one with auto-detected values.
|
||||
|
||||
### Reporting Detection Results
|
||||
|
||||
When auto-detecting VRAM, the Orchestrator should report:
|
||||
- GPU VRAM detected (if any)
|
||||
- System RAM detected
|
||||
- Model context window detected (if any)
|
||||
- Framework overhead estimated
|
||||
- Recommended VRAM context window
|
||||
- Whether auto-detection was used or manual override
|
||||
|
||||
Example:
|
||||
```
|
||||
VRAM Detection Results:
|
||||
- GPU VRAM: 8GB (nvidia-smi)
|
||||
- RAM: 16GB
|
||||
- Model context window: 128k (API-based)
|
||||
- Framework overhead: ~4k tokens
|
||||
- **Recommended: 16k tokens** (GPU VRAM-based, 25% headroom)
|
||||
- **Using: 16k tokens** (auto-detected)
|
||||
```
|
||||
|
||||
## Task
|
||||
|
||||
@@ -22,7 +127,8 @@ Each task is a state machine. The Orchestrator determines the current state and
|
||||
| State | Condition | Next State (Autopilot) |
|
||||
|-------|-----------|----------------------|
|
||||
| **New** | No artifacts in task folder | Research |
|
||||
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
|
||||
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement |
|
||||
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research |
|
||||
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
|
||||
| **Test Design** | Has `TEST_PLAN.md` | Implement |
|
||||
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
|
||||
@@ -113,8 +219,9 @@ Examine the tasks/ directory and determine the state of each task folder. Check
|
||||
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
|
||||
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
|
||||
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
|
||||
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
|
||||
13. No artifacts → **New**
|
||||
12. Has `SPEC.md` and `DECOMPOSITION.md` → **Decomposition** (sub-tasks will be created)
|
||||
13. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
|
||||
14. No artifacts → **New**
|
||||
|
||||
## Output Format
|
||||
|
||||
@@ -172,3 +279,120 @@ In Autopilot mode, after outputting the task statuses, the Orchestrator MUST aut
|
||||
When finished, output "ORCHESTRATION_COMPLETE".
|
||||
|
||||
Only recommend one phase at a time. Do not suggest running multiple phases in parallel.
|
||||
|
||||
## Sub-Task Management
|
||||
|
||||
When a parent task has a `DECOMPOSITION.md` file, the Orchestrator must create sub-task folders for each sub-task listed.
|
||||
|
||||
### Sub-Task Folder Structure
|
||||
|
||||
When decomposing, the Orchestrator creates sub-tasks under the parent task's `subtasks/` folder:
|
||||
|
||||
```
|
||||
tasks/parent-task/ → Parent task (Research → Decomposition → complete)
|
||||
SPEC.md
|
||||
DECOMPOSITION.md
|
||||
subtasks/
|
||||
subtask-a/ → Sub-task (full lifecycle independently)
|
||||
SPEC.md
|
||||
DESIGN.md
|
||||
IMPLEMENTATION.md
|
||||
BUG_REPORT.md
|
||||
ADVERSARIAL_BUG_REPORT.md
|
||||
DOC_REVIEW.md
|
||||
VERDICT.md
|
||||
subtask-b/
|
||||
SPEC.md
|
||||
...
|
||||
```
|
||||
|
||||
### Sub-Task Creation Rules
|
||||
|
||||
When a parent task reaches the **Decomposition** phase (has `SPEC.md` and `DECOMPOSITION.md`):
|
||||
|
||||
1. **Read `~/.agent-framework/config.md`** to get the VRAM configuration and check if auto-detect is enabled.
|
||||
2. **If Auto-detect: Yes**, run `{project}/.agent-framework/scripts/vram_detect.sh` to detect VRAM limits. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`. Report the detection results.
|
||||
3. **If Auto-detect: No**, use the manually specified values from config.md.
|
||||
4. **Read `DECOMPOSITION.md`** to extract all sub-task names, dependencies, and their estimated token budgets.
|
||||
5. **Verify VRAM constraints**:
|
||||
- For each sub-task, check if its estimated peak context (from DECOMPOSITION.md) is within the max peak context.
|
||||
- If a sub-task exceeds the VRAM limit, create a warning and suggest splitting it further.
|
||||
6. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`:
|
||||
- For each sub-task, create an empty folder (no files yet — the sub-task starts at the Research phase).
|
||||
- The sub-task name should be a kebab-case version of the sub-task description from DECOMPOSITION.md.
|
||||
7. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration.
|
||||
8. **Create a PARENT_SPEC.md** file for each sub-task with:
|
||||
- The parent task's SPEC.md content
|
||||
- A reference to the parent task name
|
||||
- The VRAM configuration (auto-detected or manual)
|
||||
- Any context the sub-task needs from the parent
|
||||
9. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently.
|
||||
10. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention).
|
||||
|
||||
### Sub-Task Lifecycle
|
||||
|
||||
Each sub-task follows the full lifecycle independently:
|
||||
- Starts at the **Research** phase (no artifacts in the sub-task folder)
|
||||
- Goes through Research → Decomposition (optional) → Design (optional) → Test Design (optional) → Implement → Bug Find → Adversarial Bug Find → Doc Review → Referee
|
||||
- Ends at **Complete** (VERDICT.md with PASS) or **Human Intervention** (VERDICT.md with FAIL/NEEDS_REVIEW)
|
||||
|
||||
### VRAM-Aware Sub-Task Splitting
|
||||
|
||||
If a sub-task's estimated peak context exceeds the VRAM limit from AGENT.md:
|
||||
1. Split the sub-task into smaller sub-tasks.
|
||||
2. Each new sub-task should fit within the VRAM limit.
|
||||
3. Update the DECOMPOSITION.md to reflect the new sub-tasks.
|
||||
4. Create the new sub-task folders.
|
||||
5. Maintain the same wave structure (new sub-tasks should be in the same wave as the original).
|
||||
|
||||
### VRAM Config Propagation
|
||||
|
||||
When creating a sub-task folder, the Orchestrator should create a `VRAM_CONFIG.md` file with:
|
||||
```markdown
|
||||
# VRAM Configuration for this sub-task
|
||||
|
||||
- **Auto-detect**: Yes/No
|
||||
- **Target VRAM context**: {from detection script or config.md override}
|
||||
- **Headroom**: {from detection script or config.md override}
|
||||
- **Max peak context per sub-task**: {from detection script or config.md override}
|
||||
- **GPU VRAM detected**: {value}GB (or "None")
|
||||
- **RAM detected**: {value}GB
|
||||
- **Model context window**: {value}k tokens (or "Unknown")
|
||||
- **Framework overhead**: ~{value} tokens
|
||||
- **This sub-task's estimated peak context**: {from DECOMPOSITION.md}
|
||||
- **Fits within VRAM**: Yes/No
|
||||
```
|
||||
|
||||
This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly.
|
||||
|
||||
### Sub-Task Parent Specification
|
||||
|
||||
When creating a sub-task folder, the Orchestrator should also create a `PARENT_SPEC.md` file that contains:
|
||||
- The parent task's SPEC.md content
|
||||
- A reference to the parent task name
|
||||
- Any context the sub-task needs from the parent
|
||||
|
||||
This ensures sub-tasks have all the information they need to implement their portion of the parent spec.
|
||||
|
||||
### Sub-Task Dependencies and Wave Management
|
||||
|
||||
Sub-tasks that are in the same wave (can run in parallel) should be driven through the lifecycle in parallel. Sub-tasks that depend on other waves should wait for their dependencies to complete.
|
||||
|
||||
### Sub-Task Completion and Parent Task
|
||||
|
||||
When a sub-task reaches a terminal state:
|
||||
- If the sub-task **PASSes**: The parent task can move to the next wave of sub-tasks.
|
||||
- If a sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator pauses and reports human intervention is required. The parent task is stuck until the failing sub-task is resolved.
|
||||
|
||||
When ALL sub-tasks are in terminal state:
|
||||
- If ALL sub-tasks **PASS**: The parent task is **Complete**.
|
||||
- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks and reports human intervention is required.
|
||||
|
||||
### Sub-Task Verdict Reporting
|
||||
|
||||
When a sub-task reaches the Referee phase, the VERDICT.md should include:
|
||||
- The sub-task's own verdict (PASS/FAIL/NEEDS_REVIEW)
|
||||
- A reference to the parent task name
|
||||
- Any findings that affect the parent task
|
||||
|
||||
The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status.
|
||||
|
||||
Reference in New Issue
Block a user