Split config: VRAM/model settings to config.md, agent behavior to AGENT.md

This commit is contained in:
2026-06-11 20:25:27 -04:00
parent f4587886b9
commit a74eadfb86
15 changed files with 1276 additions and 13 deletions
+229 -5
View File
@@ -4,8 +4,113 @@ You are the Orchestrator Driver. Your job is to act as a **state machine** for t
1. {project}/.agent-framework/AGENT.md — Check if Autopilot is enabled
2. {project}/.agent-framework/RULES.md
3. {project}/.agent-framework/prompts/workflow.md — The State Machine
4. Any existing files under {project}/tasks/
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
5. {project}/.agent-framework/prompts/workflow.md — The State Machine
6. Any existing files under {project}/tasks/
7. {project}/.agent-framework/scripts/vram_detect.sh — VRAM detection script (if exists)
## VRAM Detection
When VRAM configuration is needed (during task decomposition, sub-task creation, etc.), the Orchestrator MUST attempt to detect VRAM/VRAM limits dynamically.
### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`.
2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
3. **Manual override**: Check if `~/.agent-framework/config.md` has `Auto-detect: No` under VRAM Configuration. If so, use the manually specified values.
4. **Fallback**: Use 8k tokens as default, with 25% headroom.
### How to Read VRAM Config from config.md
```markdown
## VRAM Configuration
- **Auto-detect**: Yes/No
- **Target context**: {value}k tokens (override if Auto-detect: No)
- **Headroom**: {value}% (override if Auto-detect: No)
- **Max peak context per sub-task**: {value}k tokens (override if Auto-detect: No)
```
- If `Auto-detect: Yes`, run the detection script and use its output.
- If `Auto-detect: No`, use the manually specified values.
- If config.md has no VRAM Configuration section, default to `Auto-detect: Yes`.
### Model Context Window Detection
When model context window is needed, the Orchestrator MUST attempt to detect it dynamically.
#### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to detect the model name and its context window. Parse the JSON output for `model_context_kb`.
2. **Auto-detect via config**: Check `~/.agent-framework/config.md` for the model name and override context window.
3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
4. **Fallback**: Use 128k tokens as default (common for modern models).
#### How to Read Model Config from config.md
```markdown
## Model Configuration
- **Model**: auto # Use auto-detection, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
```
- If `Model: auto`, detect the model name from API config files or AGENT.md.
- If `Override context window: auto`, use the detected context window.
- If both are specified, use the specified values.
#### Model Name Lookup
When the model name is detected, look up its context window:
- gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4: 128k tokens
- claude-3-5-sonnet, claude-3-5-haiku, claude-3-opus, claude-3-sonnet, claude-3-haiku, claude-2: 200k tokens
### Detection Script Output (JSON)
The detection script outputs JSON like:
```json
{
"gpu_vram_gb": 8,
"ram_gb": 16,
"model_context_kb": 128000,
"framework_overhead_tokens": 4000,
"recommended_kb": 16000,
"recommended_k": 16,
"headroom": 0.25,
"max_peak_context_kb": 12000
}
```
Use `recommended_kb` for the target context, `headroom` for headroom, and `max_peak_context_kb` for max peak context per sub-task.
### Auto-Detect When to Run Detection
The Orchestrator should run VRAM detection in the following scenarios:
1. **When a new task is created** — to set the VRAM config for the new task.
2. **When Decomposition is triggered** — to ensure sub-tasks are sized correctly.
3. **When sub-tasks are created** — to propagate VRAM config to sub-task folders.
4. **When a sub-task's VRAM_CONFIG.md is missing** — to create one with auto-detected values.
### Reporting Detection Results
When auto-detecting VRAM, the Orchestrator should report:
- GPU VRAM detected (if any)
- System RAM detected
- Model context window detected (if any)
- Framework overhead estimated
- Recommended VRAM context window
- Whether auto-detection was used or manual override
Example:
```
VRAM Detection Results:
- GPU VRAM: 8GB (nvidia-smi)
- RAM: 16GB
- Model context window: 128k (API-based)
- Framework overhead: ~4k tokens
- **Recommended: 16k tokens** (GPU VRAM-based, 25% headroom)
- **Using: 16k tokens** (auto-detected)
```
## Task
@@ -22,7 +127,8 @@ Each task is a state machine. The Orchestrator determines the current state and
| State | Condition | Next State (Autopilot) |
|-------|-----------|----------------------|
| **New** | No artifacts in task folder | Research |
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement |
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
| **Test Design** | Has `TEST_PLAN.md` | Implement |
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
@@ -113,8 +219,9 @@ Examine the tasks/ directory and determine the state of each task folder. Check
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
13. No artifacts → **New**
12. Has `SPEC.md` and `DECOMPOSITION.md` → **Decomposition** (sub-tasks will be created)
13. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
14. No artifacts → **New**
## Output Format
@@ -172,3 +279,120 @@ In Autopilot mode, after outputting the task statuses, the Orchestrator MUST aut
When finished, output "ORCHESTRATION_COMPLETE".
Only recommend one phase at a time. Do not suggest running multiple phases in parallel.
## Sub-Task Management
When a parent task has a `DECOMPOSITION.md` file, the Orchestrator must create sub-task folders for each sub-task listed.
### Sub-Task Folder Structure
When decomposing, the Orchestrator creates sub-tasks under the parent task's `subtasks/` folder:
```
tasks/parent-task/ → Parent task (Research → Decomposition → complete)
SPEC.md
DECOMPOSITION.md
subtasks/
subtask-a/ → Sub-task (full lifecycle independently)
SPEC.md
DESIGN.md
IMPLEMENTATION.md
BUG_REPORT.md
ADVERSARIAL_BUG_REPORT.md
DOC_REVIEW.md
VERDICT.md
subtask-b/
SPEC.md
...
```
### Sub-Task Creation Rules
When a parent task reaches the **Decomposition** phase (has `SPEC.md` and `DECOMPOSITION.md`):
1. **Read `~/.agent-framework/config.md`** to get the VRAM configuration and check if auto-detect is enabled.
2. **If Auto-detect: Yes**, run `{project}/.agent-framework/scripts/vram_detect.sh` to detect VRAM limits. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`. Report the detection results.
3. **If Auto-detect: No**, use the manually specified values from config.md.
4. **Read `DECOMPOSITION.md`** to extract all sub-task names, dependencies, and their estimated token budgets.
5. **Verify VRAM constraints**:
- For each sub-task, check if its estimated peak context (from DECOMPOSITION.md) is within the max peak context.
- If a sub-task exceeds the VRAM limit, create a warning and suggest splitting it further.
6. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`:
- For each sub-task, create an empty folder (no files yet — the sub-task starts at the Research phase).
- The sub-task name should be a kebab-case version of the sub-task description from DECOMPOSITION.md.
7. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration.
8. **Create a PARENT_SPEC.md** file for each sub-task with:
- The parent task's SPEC.md content
- A reference to the parent task name
- The VRAM configuration (auto-detected or manual)
- Any context the sub-task needs from the parent
9. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently.
10. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention).
### Sub-Task Lifecycle
Each sub-task follows the full lifecycle independently:
- Starts at the **Research** phase (no artifacts in the sub-task folder)
- Goes through Research → Decomposition (optional) → Design (optional) → Test Design (optional) → Implement → Bug Find → Adversarial Bug Find → Doc Review → Referee
- Ends at **Complete** (VERDICT.md with PASS) or **Human Intervention** (VERDICT.md with FAIL/NEEDS_REVIEW)
### VRAM-Aware Sub-Task Splitting
If a sub-task's estimated peak context exceeds the VRAM limit from AGENT.md:
1. Split the sub-task into smaller sub-tasks.
2. Each new sub-task should fit within the VRAM limit.
3. Update the DECOMPOSITION.md to reflect the new sub-tasks.
4. Create the new sub-task folders.
5. Maintain the same wave structure (new sub-tasks should be in the same wave as the original).
### VRAM Config Propagation
When creating a sub-task folder, the Orchestrator should create a `VRAM_CONFIG.md` file with:
```markdown
# VRAM Configuration for this sub-task
- **Auto-detect**: Yes/No
- **Target VRAM context**: {from detection script or config.md override}
- **Headroom**: {from detection script or config.md override}
- **Max peak context per sub-task**: {from detection script or config.md override}
- **GPU VRAM detected**: {value}GB (or "None")
- **RAM detected**: {value}GB
- **Model context window**: {value}k tokens (or "Unknown")
- **Framework overhead**: ~{value} tokens
- **This sub-task's estimated peak context**: {from DECOMPOSITION.md}
- **Fits within VRAM**: Yes/No
```
This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly.
### Sub-Task Parent Specification
When creating a sub-task folder, the Orchestrator should also create a `PARENT_SPEC.md` file that contains:
- The parent task's SPEC.md content
- A reference to the parent task name
- Any context the sub-task needs from the parent
This ensures sub-tasks have all the information they need to implement their portion of the parent spec.
### Sub-Task Dependencies and Wave Management
Sub-tasks that are in the same wave (can run in parallel) should be driven through the lifecycle in parallel. Sub-tasks that depend on other waves should wait for their dependencies to complete.
### Sub-Task Completion and Parent Task
When a sub-task reaches a terminal state:
- If the sub-task **PASSes**: The parent task can move to the next wave of sub-tasks.
- If a sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator pauses and reports human intervention is required. The parent task is stuck until the failing sub-task is resolved.
When ALL sub-tasks are in terminal state:
- If ALL sub-tasks **PASS**: The parent task is **Complete**.
- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks and reports human intervention is required.
### Sub-Task Verdict Reporting
When a sub-task reaches the Referee phase, the VERDICT.md should include:
- The sub-task's own verdict (PASS/FAIL/NEEDS_REVIEW)
- A reference to the parent task name
- Any findings that affect the parent task
The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status.