Split config: VRAM/model settings to config.md, agent behavior to AGENT.md

This commit is contained in:
2026-06-11 20:25:27 -04:00
parent f4587886b9
commit a74eadfb86
15 changed files with 1276 additions and 13 deletions
+4 -1
View File
@@ -3,6 +3,8 @@
## Autopilot
Autopilot: Enabled
> Note: For global framework settings (VRAM, model, system requirements), see `~/.agent-framework/config.md`.
## Routing
IF task type = research → load prompts/research.md + RULES.md
@@ -13,7 +15,8 @@ IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code
IF task type = adversarial_bug_find → load prompts/adversarial_bug_find.md + SPEC.md + code
IF task type = referee → load prompts/referee.md + SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md
IF task type = doc_review → load prompts/doc_review.md + DESIGN.md
IF task type = decompose → load prompts/decompose.md + SPEC.md
IF task type = orchestrate → load prompts/orchestrate.md + project structure
IF task type = compaction → load prompts/compaction.md
Always start by reading this file to determine mode.
Always start by reading this file to determine mode.
+77 -3
View File
@@ -66,7 +66,8 @@ The framework features an **Autopilot** mode that allows the agent to drive a pr
### Lifecycle of a Task
1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. **Interactive** — agent grills you for requirements and gets sign-off.
2. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off.
2. **Decomposition** (optional): Break the task into sub-tasks sized for your VRAM. **Interactive** — agent analyzes the spec and proposes sub-tasks with token budget estimates, gets sign-off. See [VRAM Configuration](#vram-configuration) below.
3. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off.
3. **Test Design** (optional): Produce a `TEST_PLAN.md` (Test specification). No code allowed. **Interactive** — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off.
4. **Implement**: Write code and tests based *only* on the `SPEC.md`, `DESIGN.md` (if present), and `TEST_PLAN.md` (if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows.
5. **Bug Find**: Aggressive search for bugs and spec deviations.
@@ -79,11 +80,69 @@ The framework features an **Autopilot** mode that allows the agent to drive a pr
#### Autopilot mode (default)
The default mode is **Autopilot: Enabled**. The Orchestrator automatically drives tasks through all phases. Just say:
- **Start a new task**: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion)
- **Decompose a task**: "Decompose the add-user-auth task" (the Orchestrator breaks it into sub-tasks sized for your VRAM)
- **Continue**: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it)
#### VRAM Configuration
For low-VRAM systems (8GB, 16GB), set your VRAM limits in `~/.agent-framework/config.md`:
```markdown
## VRAM Configuration
- **Auto-detect**: Yes # Let the agent detect your VRAM automatically
- **Target VRAM context**: 16k tokens # Override auto-detect if needed
- **Headroom**: 25%
- **Max peak context per sub-task**: 12k tokens
```
**Auto-detection**: When `Auto-detect: Yes`, the Orchestrator runs `scripts/vram_detect.sh` to probe:
- GPU VRAM (via `nvidia-smi`)
- System RAM (via `free`)
- Model context window (from config.md or API config files)
- Framework overhead (by counting token load in loaded prompts)
**Manual override**: When `Auto-detect: No`, use the manually specified values:
```markdown
## VRAM Configuration
- **Auto-detect**: No
- **Target VRAM context**: 8k
- **Headroom**: 30%
- **Max peak context per sub-task**: 5.6k
```
#### Model Configuration
When using a local LLM or a specific API model, set the model in `~/.agent-framework/config.md`:
```markdown
## Model Configuration
- **Model**: auto # Use auto-detection from API config files
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
```
**Auto-detection**: When `Model: auto`, the framework detects the model name from API config files (`.env`, `config.yaml`, etc.) and looks up its context window.
**Manual override**: When you know your model name, specify it:
```markdown
## Model Configuration
- **Model**: gpt-4o
- **Override context window**: 128k
```
When you run "Decompose the X task", the Orchestrator will:
1. Analyze the task's SPEC.md
2. Detect VRAM limits (auto or manual)
3. Break it into sub-tasks, each sized to fit within your VRAM limit
4. Estimate the token budget for each sub-task
5. Create sub-task folders under `tasks/{parent-task}/subtasks/{sub-task}/`
6. Propagate VRAM config to each sub-task
Sub-tasks run independently through the full lifecycle. The parent task is complete only when ALL sub-tasks pass.
#### Manual mode (opt-in)
Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like:
- "Research add user authentication" — starts a new task
- "Decompose the add-user-auth task" — breaks into VRAM-sized sub-tasks (optional)
- "Design the add-user-auth task" — designs the architecture (optional)
- "Design tests for the add-user-auth task" — designs test cases (optional)
- "Implement the add-user-auth task" — implements the task with tests
@@ -93,12 +152,27 @@ Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you p
- "Review the add-user-auth task" — referee evaluates
- "orchestrate" — asks the Orchestrator what to do next
## Sub-Task Management
When a task is decomposed, the Orchestrator creates sub-tasks under `tasks/{parent-task}/subtasks/`. Each sub-task:
- Has its own full lifecycle (Research → Decomposition → Design → ... → Complete)
- Starts at the Research phase with an empty folder
- Receives a `PARENT_SPEC.md` with the parent task's context
- Receives a `VRAM_CONFIG.md` with the VRAM constraints
- Runs independently — sub-tasks in the same wave can run in parallel
- The parent task is NOT complete until ALL sub-tasks pass
## Key Components
- `AGENT.md`: Project-specific configuration and mode selection.
- `AGENT.md`: Project-specific agent behavior (Autopilot mode, routing rules).
- `config.md`: Global framework settings (VRAM, model, system requirements).
- `RULES.md`: Living document of project constraints and past failure modes.
- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, etc.).
- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, Decompose, etc.).
- `workflow.md`: The state machine governing the Autopilot lifecycle.
- `test_design.md`: Produces a TEST_PLAN.md — an explicit test specification before implementation.
- `decompose.md`: Breaks a task into VRAM-sized sub-tasks.
- `scripts/vram_detect.sh`: Auto-detects GPU VRAM, RAM, model context window, and framework overhead.
- `contracts/vram_config.md`: Contract for VRAM-aware task decomposition.
## Contact & Support
[Insert Contact Info]
+61
View File
@@ -0,0 +1,61 @@
# Framework Configuration
This file contains global framework settings that apply across all projects.
## VRAM Configuration
Settings for task decomposition based on available VRAM.
- **Auto-detect**: Yes # Detect GPU VRAM, RAM, and model context window automatically
- **Target context**: 16k tokens # Override auto-detect if needed
- **Headroom**: 25% # Leave headroom for code, context, and reasoning
- **Max peak context per sub-task**: 12k tokens # Max context for any single sub-task
### Auto-detection
When `Auto-detect: Yes`, the framework probes your system to detect:
- GPU VRAM (via `nvidia-smi` or `lspci`)
- System RAM (via `free`)
- Model context window (via API config or model name lookup)
- Framework overhead (by reading all loaded prompt files)
To disable auto-detection and use manual values:
```
## VRAM Configuration
- **Auto-detect**: No
- **Target context**: 8k
- **Headroom**: 30%
- **Max peak context per sub-task**: 5.6k
```
## Model Configuration
Settings for the LLM model being used.
- **Model**: auto # Use auto-detection from API config, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
### Auto-detection
When `Model: auto`, the framework detects the model name from:
1. `AGENT.md` in the project (if specified there)
2. API config files (`.env`, `config.yaml`, `config.json`, etc.)
3. Model name lookup by name (e.g., gpt-4o → 128k, claude-3-5-sonnet → 200k)
To disable auto-detection and use manual values:
```
## Model Configuration
- **Model**: gpt-4o
- **Override context window**: 128k
```
## System Requirements
Requirements for the environment the framework runs in.
- **nvidia-smi**: Required if NVIDIA GPU (for VRAM detection)
- **lspci**: Fallback if NVIDIA GPU not available (for AMD GPU VRAM detection)
- **/proc/meminfo**: Required for RAM detection (Linux)
- **sysctl**: Fallback for RAM detection (macOS)
+11
View File
@@ -0,0 +1,11 @@
## VRAM Configuration Contract
### Acceptance Criteria
- [ ] Target VRAM context is specified
- [ ] Headroom is specified (recommended: 25-40%)
- [ ] Max peak context per sub-task is calculated
- [ ] Sub-tasks are sized to fit within the max peak context
- [ ] VRAM_CONFIG.md is propagated to each sub-task
### Stop Condition
When all checkboxes are checked, output "CONTRACT_MET" and stop.
Regular → Executable
+46
View File
@@ -12,5 +12,51 @@ fi
echo "Cloning agent-framework to $FRAMEWORK_DIR..."
git clone https://gitea.yourdomain.com/you/agent-framework.git "$FRAMEWORK_DIR"
echo ""
echo "=== VRAM / Context Detection ==="
echo "Detecting your system's VRAM to recommend task decomposition settings..."
echo ""
# Run VRAM detection script if it exists
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.sh" ]; then
# Run in project-dir context so it can read framework overhead
detection_output=$(cd "$FRAMEWORK_DIR" && bash "$FRAMEWORK_DIR/scripts/vram_detect.sh" 2>&1)
# Extract JSON output (last section after "=== JSON Output ===")
json_output=$(echo "$detection_output" | sed -n '/=== JSON Output ===/,/EOF/p' | grep -v '=== JSON Output ===' | grep -v '^EOF$')
if [ -n "$json_output" ]; then
echo "$detection_output"
# Extract key values from JSON for display
recommended_k=$(echo "$json_output" | grep '"recommended_k"' | grep -oP '\d+')
max_peak_kb=$(echo "$json_output" | grep '"max_peak_context_kb"' | grep -oP '\d+')
headroom=$(echo "$json_output" | grep '"headroom"' | grep -oP '\d+\.\d+')
gpu_vram=$(echo "$json_output" | grep '"gpu_vram_gb"' | grep -oP '\d+')
ram_gb=$(echo "$json_output" | grep '"ram_gb"' | grep -oP '\d+')
model_context=$(echo "$json_output" | grep '"model_context_kb"' | grep -oP '\d+')
echo ""
echo "=== Recommended VRAM Configuration ==="
echo "For low-VRAM systems (8GB, 16GB VRAM), add this to ~/.agent-framework/config.md:"
echo ""
echo "## VRAM Configuration"
echo "- **Auto-detect**: Yes # Let the agent detect automatically"
echo "- **Target context**: ${recommended_k}k tokens # Override auto-detect if needed"
echo "- **Headroom**: ${headroom}%"
echo "- **Max peak context per sub-task**: $((max_peak_kb / 1000))k tokens"
echo ""
echo "This ensures tasks are decomposed into sub-tasks that fit within your"
echo "available VRAM. For more information, see the README."
else
echo "Could not detect VRAM. You can manually set your VRAM configuration in ~/.agent-framework/config.md."
echo "See the README for details."
fi
else
echo "VRAM detection script not found. You can manually set your VRAM configuration in ~/.agent-framework/config.md."
echo "See the README for details."
fi
echo ""
echo "Installation complete."
echo "Next step: cd into a project and run the onboarding prompt."
+12 -1
View File
@@ -1,9 +1,20 @@
You are the Adversarial Bug Finder.
Read the SPEC.md and the code.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
3. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
4. The code
Your job is to find bugs that are difficult to spot, such as complex logic errors, race conditions, and performance bottlenecks. Be more aggressive and exhaustive than a standard bug finder.
If VRAM_CONFIG.md exists, also check for:
- Memory leaks (loading large files into context that could cause OOM)
- N+1 query patterns that could cause memory exhaustion
- Infinite loops that could run out of context
- Unbounded recursion that could cause stack overflow
Output your findings in ADVERSARIAL_BUG_REPORT.md.
When finished, output "ADVERSARIAL_BUG_FIND_COMPLETE".
+2
View File
@@ -5,6 +5,8 @@ You are the Bug Finder. Your job is to find every bug, deviation from spec, and
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
3. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
4. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
5. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
## Task
+233
View File
@@ -0,0 +1,233 @@
You are in decomposition mode.
Your only job is to take a completed SPEC.md and break it into the smallest possible, independently verifiable sub-tasks. Each sub-task should be small enough to complete in one session and should have clear, testable acceptance criteria.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/.agent-framework/RULES.md
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
5. {project}/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection)
## Task
{task-description}
## Decomposition Rules
### Rule 1: Smallest Possible Unit
Break every feature into the smallest possible units that are still independently testable. If a sub-task can be done in one session, it should be a sub-task.
### Rule 2: Each Sub-Task Must Be Self-Contained
Each sub-task must have:
- Its own goal statement (one sentence)
- Clear acceptance criteria (at least 2-3)
- Dependencies on other sub-tasks (if any)
- Its own contract (SPEC.md) that references the parent task
### Rule 3: Dependencies Must Be Explicit
If sub-task B depends on sub-task A, state it clearly. A sub-task with no dependencies can run in parallel. Sub-tasks that depend on parallel sub-tasks must wait for both.
### Rule 4: Define Execution Order
After listing all sub-tasks, define the execution order considering dependencies. Group independent sub-tasks into "waves" that can be done in parallel.
### Rule 5: Do Not Create Sub-Sub-Tasks
Decomposition produces **one level** of sub-tasks only. If a sub-task is still too large, the implementer should break it down further during implementation, but do NOT nest sub-tasks.
### Rule 6: Token Budget Per Sub-Task
Each sub-task must fit within the target VRAM context window. Estimate the total token budget for each sub-task's full lifecycle and break it down further if it exceeds the limit.
**How to estimate token budget for a sub-task:**
During a sub-task's lifecycle, the following files are loaded into context at various phases:
- **Research phase**: RULES.md + AGENT.md + task description
- **Design phase**: SPEC.md + RULES.md
- **Implement phase**: SPEC.md + DESIGN.md + TEST_PLAN.md + RULES.md + AGENT.md + CONTRACT.md
- **Bug Find phase**: SPEC.md + code (limited scope)
- **Adversarial Bug Find phase**: SPEC.md + code (limited scope)
- **Doc Review phase**: DESIGN.md
- **Referee phase**: SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md
The **peak context** is during the Implement phase, where all files are loaded together. Estimate the token count of the combined files for the sub-task.
**Guidelines:**
- **8k VRAM**: Peak context for Implement phase should be ≤ 6k tokens (leave 2k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 6k tokens.
- **16k VRAM**: Peak context for Implement phase should be ≤ 12k tokens (leave 4k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 12k tokens.
- **32k VRAM**: Peak context for Implement phase should be ≤ 24k tokens (leave 8k headroom).
- **64k VRAM**: Peak context for Implement phase should be ≤ 48k tokens (leave 16k headroom).
If a sub-task's estimated context exceeds the limit, break it into smaller sub-tasks.
**How to estimate token count:**
- Roughly 1 token = 4 characters (for English text)
- A SPEC.md with 5 requirements, each with 2-3 acceptance criteria, is typically 500-1000 tokens
- A DESIGN.md with 3 sections and 5-10 bullet points is typically 1000-3000 tokens
- A TEST_PLAN.md with 5-10 test cases is typically 1000-2000 tokens
- RULES.md is typically 200-1000 tokens (varies per project)
- AGENT.md is typically 300-1000 tokens
- A CONTRACT.md is typically 200-500 tokens
**Quick estimate formula:**
```
Peak context ≈ SPEC.md tokens + DESIGN.md tokens + TEST_PLAN.md tokens + RULES.md tokens + AGENT.md tokens + CONTRACT.md tokens
```
### Rule 7: Sub-Task Size Targets
Aim for sub-tasks that are:
- **Small** (8k VRAM): ~200-800 tokens of combined spec/design/test files
- **Small** (16k VRAM): ~200-1500 tokens of combined spec/design/test files
- **Medium** (8k VRAM): ~800-2000 tokens of combined spec/design/test files
- **Medium** (16k VRAM): ~1500-4000 tokens of combined spec/design/test files
- **Large** (8k VRAM): ~2000-4000 tokens of combined spec/design/test files
- **Large** (16k VRAM): ~4000-8000 tokens of combined spec/design/test files
If a sub-task exceeds the "Large" target for the target VRAM, break it further.
## Decomposition Protocol (Interactive)
### Phase 1: Analysis
Before decomposing, analyze the SPEC.md:
1. Identify all distinct features/requirements
2. Identify data models that need to be created
3. Identify API endpoints or interfaces
4. Identify infrastructure changes
5. Identify configuration changes
6. Estimate the token budget for the full task (sum of all requirements' SPEC + DESIGN + TEST files)
7. **Detect VRAM limits**:
- Check `~/.agent-framework/config.md` for VRAM Configuration section
- If `Auto-detect: Yes`, run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window
- If `Auto-detect: No`, use the manually specified values from config.md
- Report the detected VRAM limits
8. **Detect model context window**:
- Check `~/.agent-framework/config.md` for Model Configuration section
- If `Model: auto`, run the detection script to detect the model name and its context window
- If `Override context window: auto`, use the detected context window
- If both are specified, use the specified values
- If model detection fails, use 128k tokens as default
9. Determine the target VRAM context window based on the detection results
### Phase 2: Propose Decomposition
Present a draft decomposition to the user. Format:
**Target VRAM**: {8k/16k/32k/64k} tokens
**Waves:**
- **Wave 1**: Sub-task 1, Sub-task 2, Sub-task 3 (can run in parallel)
- **Wave 2**: Sub-task 4, Sub-task 5 (depends on Wave 1)
- **Wave 3**: Sub-task 6 (depends on Wave 2)
**Sub-task Details:**
1. **{sub-task-name}**
- Goal: {one sentence}
- Dependencies: {list of sub-task names, or "None"}
- Acceptance criteria:
- [ ] {criterion 1}
- [ ] {criterion 2}
- Estimated scope: {small/medium/large}
- **Estimated token budget**: ~{estimate} tokens (peak: ~{peak} tokens during Implement phase)
- **Fits within VRAM**: Yes/No (if No, explain why and suggest how to split further)
### Phase 3: Review and Refine
Present the draft decomposition to the user and ask:
- "Are there any sub-tasks that are too large?"
- "Are there any sub-tasks that should be combined?"
- "Are the dependencies correct?"
- "Are there any sub-tasks I missed?"
- "Is the execution order optimal?"
- "Do the token budget estimates look reasonable for your VRAM?"
- "Are there any sub-tasks that exceed your VRAM limit?"
Incorporate the user's feedback and revise the decomposition accordingly.
### Phase 4: Get Sign-Off
Before writing the DECOMPOSITION.md, you MUST get explicit sign-off from the user. Say:
> "Based on our discussion, here is the final decomposition:
> [brief summary of waves, sub-tasks, and token budgets]
> Does this cover everything? Please confirm with 'APPROVED' before I finalize."
Only produce the DECOMPOSITION.md after the user says "APPROVED" or equivalent.
## Output
Produce a file called DECOMPOSITION.md at {project}/tasks/{task-name}/DECOMPOSITION.md that contains:
```markdown
# Task Decomposition
## Parent Task
{parent-task-name}
## VRAM Configuration
- **Target VRAM**: {8k/16k/32k/64k} tokens
- **Headroom**: {25-40%} (leaves headroom for code, context, and reasoning)
- **Max peak context per sub-task**: {estimate} tokens
## Waves
### Wave 1: {wave-name}
- {sub-task-name-1}
- {sub-task-name-2}
- {sub-task-name-3}
### Wave 2: {wave-name}
- {sub-task-name-4}
- {sub-task-name-5}
## Sub-Task Details
### 1. {sub-task-name-1}
- **Goal**: {one sentence}
- **Dependencies**: None (or list sub-task names)
- **Acceptance Criteria**:
- [ ] {criterion 1}
- [ ] {criterion 2}
- **Estimated Scope**: {small/medium/large}
- **Estimated token budget**: ~{estimate} tokens
- SPEC.md: ~{x} tokens
- DESIGN.md: ~{x} tokens
- TEST_PLAN.md: ~{x} tokens
- RULES.md: ~{x} tokens
- AGENT.md: ~{x} tokens
- CONTRACT.md: ~{x} tokens
- **Peak context (Implement phase)**: ~{peak} tokens
- **Fits within VRAM**: Yes
### 2. {sub-task-name-2}
- **Goal**: {one sentence}
- **Dependencies**: {list of sub-task names}
- **Acceptance Criteria**:
- [ ] {criterion 1}
- [ ] {criterion 2}
- **Estimated Scope**: {small/medium/large}
- **Estimated token budget**: ~{estimate} tokens
- SPEC.md: ~{x} tokens
- DESIGN.md: ~{x} tokens
- TEST_PLAN.md: ~{x} tokens
- RULES.md: ~{x} tokens
- AGENT.md: ~{x} tokens
- CONTRACT.md: ~{x} tokens
- **Peak context (Implement phase)**: ~{peak} tokens
- **Fits within VRAM**: Yes
... etc ...
## Execution Order
1. Complete Wave 1 (all sub-tasks can run in parallel)
2. Complete Wave 2 (depends on Wave 1)
3. ... etc ...
```
When the decomposition is complete, output "CONTRACT_MET" and stop.
Do not create sub-task folders or files. The Orchestrator will handle creating sub-task folders based on this DECOMPOSITION.md.
## Stop Condition (MANDATORY)
You are not allowed to end this session until you have produced the DECOMPOSITION.md file AND output the exact phrase "CONTRACT_MET".
Until then, continue working or ask clarifying questions.
+5 -2
View File
@@ -4,8 +4,10 @@ You are in Documentation Review mode.
1. {project}/tasks/{task-name}/DESIGN.md — Look for the "Documentation Plan" section
2. {project}/tasks/{task-name}/SPEC.md — Check what the spec requires
3. Any existing documentation files mentioned in the DESIGN.md Documentation Plan
4. The code that was implemented (implementation artifacts)
3. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
4. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
5. Any existing documentation files mentioned in the DESIGN.md Documentation Plan
6. The code that was implemented (implementation artifacts)
## Task
@@ -36,6 +38,7 @@ You are in Documentation Review mode.
### Documentation Gaps
- Are there any areas where the documentation is thin or missing?
- Are there any complex flows or non-obvious logic that should be documented?
- If VRAM_CONFIG.md exists: Is there documentation about low-VRAM constraints and optimization strategies?
## Output
+14
View File
@@ -8,6 +8,8 @@ You are in implementation mode.
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
6. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
7. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists — for low-VRAM systems)
8. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists — for sub-tasks)
## Task
@@ -26,6 +28,18 @@ You are in implementation mode.
- Keep functions small and focused.
- Use existing patterns in the codebase.
### VRAM-Aware Implementation (if VRAM_CONFIG.md exists)
If `VRAM_CONFIG.md` is present, the implementer is working in a low-VRAM environment and must:
- **Keep functions small**: Each function should be ≤ 50 lines. Break large functions into smaller ones.
- **Avoid loading large files into context**: If implementing in an iterative environment (like pi), read files incrementally rather than loading entire files at once.
- **Write self-contained modules**: Each module should be independently testable and runnable without loading the entire codebase.
- **Prefer streaming over buffering**: Use streaming patterns instead of loading entire datasets into memory.
- **Be explicit about dependencies**: Import only what you need, not the entire module.
- **Test incrementally**: Run tests frequently rather than building up a large test suite that takes a long time to run.
- **Optimize for low-VRAM**: Write code that can be understood and modified by an agent with limited context window.
### End State
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
- All tests must pass.
+28
View File
@@ -8,6 +8,7 @@ Your only job is to set up the minimal agent framework structure in the target p
2. ~/.agent-framework/ONBOARDING.md — human reference for drop-in vs from-scratch scenarios
3. {project}/.agent-framework/AGENT.md (if it exists)
4. {project}/.agent-framework/RULES.md (if it exists)
5. ~/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection)
## Task
@@ -36,6 +37,32 @@ Before proceeding, check if the project is running an older version of the frame
5. Explore the project root at a high level (ls, key directories, README if present).
6. Produce a short onboarding report.
### Step 2: VRAM Configuration
Check if VRAM configuration is available in `~/.agent-framework/config.md`:
1. Read `~/.agent-framework/config.md` to check for VRAM Configuration section.
2. If VRAM Configuration section exists, note the values.
3. If VRAM Configuration section does NOT exist, check if VRAM detection is available: `~/.agent-framework/scripts/vram_detect.sh`.
4. If available, run it to get VRAM recommendations:
```
cd ~/.agent-framework && bash ~/.agent-framework/scripts/vram_detect.sh
```
5. Parse the JSON output for `recommended_k`, `max_peak_context_kb`, and `headroom`.
6. Add a VRAM Configuration section to `~/.agent-framework/config.md`:
```
## VRAM Configuration
- **Auto-detect**: Yes
- **Target context**: {recommended_k}k tokens
- **Headroom**: 25%
- **Max peak context per sub-task**: {max_peak_kb/1000}k tokens
```
7. If VRAM detection failed or is not available, add a minimal section:
```
## VRAM Configuration
- **Auto-detect**: Yes
```
8. Report the VRAM configuration status in the onboarding report.
## Output
Create or update the following inside {project}/.agent-framework/:
@@ -50,6 +77,7 @@ Then produce a file called ONBOARDING_REPORT.md at {project}/tasks/onboarding/ON
- Key observations from the project structure
- Any missing pieces the human should provide next
- **Upgrade status**: Whether the project's framework files are up to date with the global framework
- **VRAM Configuration**: Whether VRAM config was auto-detected and applied (if available), or needs manual setup
When the ritual is complete, output "CONTRACT_MET" and stop.
+229 -5
View File
@@ -4,8 +4,113 @@ You are the Orchestrator Driver. Your job is to act as a **state machine** for t
1. {project}/.agent-framework/AGENT.md — Check if Autopilot is enabled
2. {project}/.agent-framework/RULES.md
3. {project}/.agent-framework/prompts/workflow.md — The State Machine
4. Any existing files under {project}/tasks/
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
5. {project}/.agent-framework/prompts/workflow.md — The State Machine
6. Any existing files under {project}/tasks/
7. {project}/.agent-framework/scripts/vram_detect.sh — VRAM detection script (if exists)
## VRAM Detection
When VRAM configuration is needed (during task decomposition, sub-task creation, etc.), the Orchestrator MUST attempt to detect VRAM/VRAM limits dynamically.
### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`.
2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
3. **Manual override**: Check if `~/.agent-framework/config.md` has `Auto-detect: No` under VRAM Configuration. If so, use the manually specified values.
4. **Fallback**: Use 8k tokens as default, with 25% headroom.
### How to Read VRAM Config from config.md
```markdown
## VRAM Configuration
- **Auto-detect**: Yes/No
- **Target context**: {value}k tokens (override if Auto-detect: No)
- **Headroom**: {value}% (override if Auto-detect: No)
- **Max peak context per sub-task**: {value}k tokens (override if Auto-detect: No)
```
- If `Auto-detect: Yes`, run the detection script and use its output.
- If `Auto-detect: No`, use the manually specified values.
- If config.md has no VRAM Configuration section, default to `Auto-detect: Yes`.
### Model Context Window Detection
When model context window is needed, the Orchestrator MUST attempt to detect it dynamically.
#### Detection Priority
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to detect the model name and its context window. Parse the JSON output for `model_context_kb`.
2. **Auto-detect via config**: Check `~/.agent-framework/config.md` for the model name and override context window.
3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
4. **Fallback**: Use 128k tokens as default (common for modern models).
#### How to Read Model Config from config.md
```markdown
## Model Configuration
- **Model**: auto # Use auto-detection, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
```
- If `Model: auto`, detect the model name from API config files or AGENT.md.
- If `Override context window: auto`, use the detected context window.
- If both are specified, use the specified values.
#### Model Name Lookup
When the model name is detected, look up its context window:
- gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4: 128k tokens
- claude-3-5-sonnet, claude-3-5-haiku, claude-3-opus, claude-3-sonnet, claude-3-haiku, claude-2: 200k tokens
### Detection Script Output (JSON)
The detection script outputs JSON like:
```json
{
"gpu_vram_gb": 8,
"ram_gb": 16,
"model_context_kb": 128000,
"framework_overhead_tokens": 4000,
"recommended_kb": 16000,
"recommended_k": 16,
"headroom": 0.25,
"max_peak_context_kb": 12000
}
```
Use `recommended_kb` for the target context, `headroom` for headroom, and `max_peak_context_kb` for max peak context per sub-task.
### Auto-Detect When to Run Detection
The Orchestrator should run VRAM detection in the following scenarios:
1. **When a new task is created** — to set the VRAM config for the new task.
2. **When Decomposition is triggered** — to ensure sub-tasks are sized correctly.
3. **When sub-tasks are created** — to propagate VRAM config to sub-task folders.
4. **When a sub-task's VRAM_CONFIG.md is missing** — to create one with auto-detected values.
### Reporting Detection Results
When auto-detecting VRAM, the Orchestrator should report:
- GPU VRAM detected (if any)
- System RAM detected
- Model context window detected (if any)
- Framework overhead estimated
- Recommended VRAM context window
- Whether auto-detection was used or manual override
Example:
```
VRAM Detection Results:
- GPU VRAM: 8GB (nvidia-smi)
- RAM: 16GB
- Model context window: 128k (API-based)
- Framework overhead: ~4k tokens
- **Recommended: 16k tokens** (GPU VRAM-based, 25% headroom)
- **Using: 16k tokens** (auto-detected)
```
## Task
@@ -22,7 +127,8 @@ Each task is a state machine. The Orchestrator determines the current state and
| State | Condition | Next State (Autopilot) |
|-------|-----------|----------------------|
| **New** | No artifacts in task folder | Research |
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement |
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
| **Test Design** | Has `TEST_PLAN.md` | Implement |
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
@@ -113,8 +219,9 @@ Examine the tasks/ directory and determine the state of each task folder. Check
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
13. No artifacts → **New**
12. Has `SPEC.md` and `DECOMPOSITION.md` → **Decomposition** (sub-tasks will be created)
13. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
14. No artifacts → **New**
## Output Format
@@ -172,3 +279,120 @@ In Autopilot mode, after outputting the task statuses, the Orchestrator MUST aut
When finished, output "ORCHESTRATION_COMPLETE".
Only recommend one phase at a time. Do not suggest running multiple phases in parallel.
## Sub-Task Management
When a parent task has a `DECOMPOSITION.md` file, the Orchestrator must create sub-task folders for each sub-task listed.
### Sub-Task Folder Structure
When decomposing, the Orchestrator creates sub-tasks under the parent task's `subtasks/` folder:
```
tasks/parent-task/ → Parent task (Research → Decomposition → complete)
SPEC.md
DECOMPOSITION.md
subtasks/
subtask-a/ → Sub-task (full lifecycle independently)
SPEC.md
DESIGN.md
IMPLEMENTATION.md
BUG_REPORT.md
ADVERSARIAL_BUG_REPORT.md
DOC_REVIEW.md
VERDICT.md
subtask-b/
SPEC.md
...
```
### Sub-Task Creation Rules
When a parent task reaches the **Decomposition** phase (has `SPEC.md` and `DECOMPOSITION.md`):
1. **Read `~/.agent-framework/config.md`** to get the VRAM configuration and check if auto-detect is enabled.
2. **If Auto-detect: Yes**, run `{project}/.agent-framework/scripts/vram_detect.sh` to detect VRAM limits. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`. Report the detection results.
3. **If Auto-detect: No**, use the manually specified values from config.md.
4. **Read `DECOMPOSITION.md`** to extract all sub-task names, dependencies, and their estimated token budgets.
5. **Verify VRAM constraints**:
- For each sub-task, check if its estimated peak context (from DECOMPOSITION.md) is within the max peak context.
- If a sub-task exceeds the VRAM limit, create a warning and suggest splitting it further.
6. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`:
- For each sub-task, create an empty folder (no files yet — the sub-task starts at the Research phase).
- The sub-task name should be a kebab-case version of the sub-task description from DECOMPOSITION.md.
7. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration.
8. **Create a PARENT_SPEC.md** file for each sub-task with:
- The parent task's SPEC.md content
- A reference to the parent task name
- The VRAM configuration (auto-detected or manual)
- Any context the sub-task needs from the parent
9. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently.
10. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention).
### Sub-Task Lifecycle
Each sub-task follows the full lifecycle independently:
- Starts at the **Research** phase (no artifacts in the sub-task folder)
- Goes through Research → Decomposition (optional) → Design (optional) → Test Design (optional) → Implement → Bug Find → Adversarial Bug Find → Doc Review → Referee
- Ends at **Complete** (VERDICT.md with PASS) or **Human Intervention** (VERDICT.md with FAIL/NEEDS_REVIEW)
### VRAM-Aware Sub-Task Splitting
If a sub-task's estimated peak context exceeds the VRAM limit from AGENT.md:
1. Split the sub-task into smaller sub-tasks.
2. Each new sub-task should fit within the VRAM limit.
3. Update the DECOMPOSITION.md to reflect the new sub-tasks.
4. Create the new sub-task folders.
5. Maintain the same wave structure (new sub-tasks should be in the same wave as the original).
### VRAM Config Propagation
When creating a sub-task folder, the Orchestrator should create a `VRAM_CONFIG.md` file with:
```markdown
# VRAM Configuration for this sub-task
- **Auto-detect**: Yes/No
- **Target VRAM context**: {from detection script or config.md override}
- **Headroom**: {from detection script or config.md override}
- **Max peak context per sub-task**: {from detection script or config.md override}
- **GPU VRAM detected**: {value}GB (or "None")
- **RAM detected**: {value}GB
- **Model context window**: {value}k tokens (or "Unknown")
- **Framework overhead**: ~{value} tokens
- **This sub-task's estimated peak context**: {from DECOMPOSITION.md}
- **Fits within VRAM**: Yes/No
```
This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly.
### Sub-Task Parent Specification
When creating a sub-task folder, the Orchestrator should also create a `PARENT_SPEC.md` file that contains:
- The parent task's SPEC.md content
- A reference to the parent task name
- Any context the sub-task needs from the parent
This ensures sub-tasks have all the information they need to implement their portion of the parent spec.
### Sub-Task Dependencies and Wave Management
Sub-tasks that are in the same wave (can run in parallel) should be driven through the lifecycle in parallel. Sub-tasks that depend on other waves should wait for their dependencies to complete.
### Sub-Task Completion and Parent Task
When a sub-task reaches a terminal state:
- If the sub-task **PASSes**: The parent task can move to the next wave of sub-tasks.
- If a sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator pauses and reports human intervention is required. The parent task is stuck until the failing sub-task is resolved.
When ALL sub-tasks are in terminal state:
- If ALL sub-tasks **PASS**: The parent task is **Complete**.
- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks and reports human intervention is required.
### Sub-Task Verdict Reporting
When a sub-task reaches the Referee phase, the VERDICT.md should include:
- The sub-task's own verdict (PASS/FAIL/NEEDS_REVIEW)
- A reference to the parent task name
- Any findings that affect the parent task
The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status.
+3
View File
@@ -10,6 +10,8 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
7. {project}/tasks/{task-name}/DESIGN.md (if exists)
8. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
9. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
10. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
## Task
@@ -38,6 +40,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
- Are functions small and focused?
- Is there proper error handling?
- Are there any obvious performance issues?
- If VRAM_CONFIG.md exists: Is the code optimized for low-VRAM (small functions, streaming patterns, no large file loading)?
### Testing
- Do all tests pass?
+2 -1
View File
@@ -7,7 +7,8 @@ This file defines the linear progression of a task in the agent-framework. The O
| Current State | Signal (Artifact) | Next Phase | Action |
| :--- | :--- | :--- | :--- |
| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` |
| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code |
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement | Generate `DECOMPOSITION.md` or `DESIGN.md` or code |
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research | Orchestrator creates sub-task folders |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
| **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests |
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |
+549
View File
@@ -0,0 +1,549 @@
#!/usr/bin/env bash
# VRAM/Context Detection Script
# Detects GPU VRAM, system RAM, and model context window to recommend
# a safe VRAM context window for task decomposition.
#
# Usage: ./vram_detect.sh [model_name]
# - If model_name is provided, looks up its context window
# - Otherwise, tries to detect from API config or config.md
set -uo pipefail # Don't exit on error - we want to continue even if detection fails
# ─── GPU VRAM Detection ───
detect_gpu_vram() {
local total_vram_kb=0
local vram_per_gpu_kb=0
local num_gpus=0
# Try nvidia-smi first (NVIDIA GPUs)
if command -v nvidia-smi &>/dev/null; then
local vram_kb
# Use timeout to avoid hanging on nvidia-smi (e.g., driver not loaded)
vram_kb=$(timeout 5 nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ' || true)
# Validate that vram_kb is a positive number
if [[ -n "$vram_kb" && "$vram_kb" =~ ^[0-9]+$ && "$vram_kb" -gt 0 ]]; then
total_vram_kb=$((vram_kb * 1024)) # MB → KB
vram_per_gpu_kb=$((total_vram_kb / (num_gpus+1)))
num_gpus=1
echo "GPU: NVIDIA (nvidia-smi available)"
echo "VRAM per GPU: $((vram_kb / 1024))GB ($vram_kb MB)"
else
echo "GPU: NVIDIA (nvidia-smi available but driver not responding)"
fi
fi
# Fallback: lspci
if [[ $total_vram_kb -eq 0 && $num_gpus -eq 0 ]]; then
local gpu_info
gpu_info=$(lspci 2>/dev/null | grep -i -E 'VGA|3D|Display' | head -5)
if [[ -n "$gpu_info" ]]; then
echo "GPU detected: $gpu_info"
# Try to get VRAM from lspci -vnn memory regions
# GPUs show VRAM as Memory regions in lspci
# Parse patterns like: Memory at f800000000 (64-bit, prefetchable) [size=256M]
local total_vram_mb=0
while IFS= read -r line; do
# Extract the size value from [size=256M] pattern
local size_num
size_num=$(echo "$line" | grep -oE '[0-9]+(M|G|K)' | head -1 || true)
if [[ -n "$size_num" ]]; then
local size_val
size_val=$(echo "$size_num" | grep -oE '[0-9]+')
local size_unit
size_unit=$(echo "$size_num" | grep -oE '(M|G|K)')
if [[ -n "$size_val" && -n "$size_unit" ]]; then
case "$size_unit" in
M) total_vram_mb=$((total_vram_mb + size_val)) ;;
G) total_vram_mb=$((total_vram_mb + size_val * 1024)) ;;
K) total_vram_mb=$((total_vram_mb + size_val / 1024)) ;;
esac
fi
fi
done < <(lspci -vnn 2>/dev/null | grep -i -A 15 -E 'VGA|3D|Display' | grep -i 'Memory at')
if [[ $total_vram_mb -gt 0 ]]; then
local total_vram_gb=$((total_vram_mb / 1024))
local total_vram_mb_remain=$((total_vram_mb % 1024))
echo "VRAM: $total_vram_mb MB ($total_vram_gb GB $total_vram_mb_remain MB)"
else
echo "VRAM: Could not determine from lspci"
fi
# Check for AMD GPU via amdgpu sysfs
if lspci -vnn 2>/dev/null | grep -qi 'amd\|ati'; then
local amdgpu_info
amdgpu_info=$(ls /sys/kernel/debug/amdgpu/ 2>/dev/null | head -1)
if [[ -n "$amdgpu_info" ]]; then
local vram_total
vram_total=$(cat /sys/kernel/debug/amdgpu/${amdgpu_info}/vram_total 2>/dev/null || echo 0)
if [[ "$vram_total" -gt 0 ]]; then
local vram_gb=$((vram_total / 1024 / 1024 / 1024))
local vram_mb=$((vram_total / 1024 / 1024))
echo "AMD GPU VRAM: ${vram_gb}GB (${vram_mb}MB)"
fi
fi
fi
fi
fi
echo "Total VRAM: $((total_vram_kb / 1024 / 1024))GB"
echo "VRAM per GPU: $((vram_per_gpu_kb / 1024 / 1024))GB"
echo "Num GPUs: $num_gpus"
}
# ─── System RAM Detection ───
detect_ram() {
local total_kb=0
local available_kb=0
if [[ -f /proc/meminfo ]]; then
total_kb=$(grep MemTotal /proc/meminfo | awk '{print $2}')
available_kb=$(grep MemAvailable /proc/meminfo | awk '{print $2}')
if [[ $total_kb -gt 0 ]]; then
echo "RAM: $((total_kb / 1024 / 1024))GB total, $((available_kb / 1024 / 1024))GB available"
echo "$available_kb $total_kb"
fi
elif command -v sysctl &>/dev/null; then
total_kb=$(sysctl -n hw.memsize 2>/dev/null | awk '{print $1 / 1024}')
if [[ -n "$total_kb" && "$total_kb" -gt 0 ]]; then
echo "RAM: $((total_kb / 1024))GB total"
echo "$total_kb $total_kb" # Assume all available
fi
else
echo "RAM: Could not detect"
echo "0 0"
fi
}
# ─── Model Context Window Detection ───
detect_model_context() {
local model_name="$1"
local context_kb=0
# If model name provided, look it up
if [[ -n "$model_name" ]]; then
case "$model_name" in
gpt-4o|gpt-4o-2024-05-13|gpt-4o-2024-08-06)
context_kb=128000; echo "Model: $model_name"
echo "Context window: 128k tokens"
;;
gpt-4o-mini|gpt-4o-mini-2024-07-18)
context_kb=128000; echo "Model: $model_name"
echo "Context window: 128k tokens"
;;
gpt-4-turbo|gpt-4-turbo-2024-04-09)
context_kb=128000; echo "Model: $model_name"
echo "Context window: 128k tokens"
;;
gpt-4|gpt-4-0125-preview|gpt-4-1106-preview)
context_kb=128000; echo "Model: $model_name"
echo "Context window: 128k tokens"
;;
claude-3-5-sonnet|claude-3-5-sonnet-20241022)
context_kb=200000; echo "Model: $model_name"
echo "Context window: 200k tokens"
;;
claude-3-5-haiku|claude-3-5-haiku-20241022)
context_kb=200000; echo "Model: $model_name"
echo "Context window: 200k tokens"
;;
claude-3-opus|claude-3-opus-20240229)
context_kb=200000; echo "Model: $model_name"
echo "Context window: 200k tokens"
;;
claude-3-sonnet|claude-3-sonnet-20240229)
context_kb=200000; echo "Model: $model_name"
echo "Context window: 200k tokens"
;;
claude-3-haiku|claude-3-haiku-20240307)
context_kb=200000; echo "Model: $model_name"
echo "Context window: 200k tokens"
;;
claude-2|claude-2.1)
context_kb=200000; echo "Model: $model_name"
echo "Context window: 200k tokens"
;;
*)
echo "Model: $model_name (unknown context window)"
echo "0"
;;
esac
echo "$context_kb"
return
fi
# Try to detect from config.md (global framework model settings)
local project_dir="${1:-.}"
local config_md="${HOME}/.agent-framework/config.md"
local model_from_config=""
local override_context=""
if [[ -f "$config_md" ]]; then
# Check for model name
model_from_config=$(grep -i "model:" "$config_md" 2>/dev/null | grep -v "#" | grep -v "model_context" | grep -v "override" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
# Check for override context window
override_context=$(grep -i "override context" "$config_md" 2>/dev/null | grep -v "#" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
fi
# If config.md specifies a model, use it
if [[ -n "$model_from_config" ]]; then
echo "Found model in config.md: $model_from_config"
# Use override context window if specified in config.md
if [[ -n "$override_context" && "$override_context" != "auto" ]]; then
echo "Using override context window from config.md: $override_context"
local override_kb
override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true)
if [[ -n "$override_kb" ]]; then
echo "Context window: ${override_context} tokens (override)"
echo "$((override_kb * 1000))"
return
fi
fi
detect_model_context "$model_from_config"
return
fi
# Try to detect from AGENT.md (project-level model override)
local agent_md="${project_dir}/.agent-framework/AGENT.md"
if [[ -f "$agent_md" ]]; then
local model_line
model_line=$(grep -i "model" "$agent_md" 2>/dev/null | grep -v "#" | grep -v "target" | grep -v "headroom" | grep -v "peak" | grep -v "Auto-detect" | head -1)
if [[ -n "$model_line" ]]; then
echo "Found model in AGENT.md: $model_line"
# Extract model name from the line
local model
model=$(echo "$model_line" | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
if [[ -n "$model" ]]; then
# Use override context window if specified in config.md
if [[ -n "$override_context" && "$override_context" != "auto" ]]; then
echo "Using override context window from config.md: $override_context"
# Convert override_context to kb (e.g., 128k -> 128000, 200k -> 200000)
local override_kb
override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true)
if [[ -n "$override_kb" ]]; then
echo "Context window: ${override_context} tokens (override)"
echo "$((override_kb * 1000))"
return
fi
fi
detect_model_context "$model"
return
fi
fi
fi
# Try to detect from common API config files
local config_files=(
".env"
".env.local"
"config.yaml"
"config.yml"
"config.json"
"settings.yaml"
".agent-framework/config.yaml"
".agent-framework/config.json"
)
for config_file in "${config_files[@]}"; do
local abs_file=""
for candidate in "${project_dir}/${config_file}" "${project_dir}/.agent-framework/${config_file}"; do
if [[ -f "$candidate" ]]; then
abs_file="$candidate"
break
fi
done
if [[ -n "$abs_file" ]]; then
local model
model=$(grep -i "model" "$abs_file" 2>/dev/null | grep -v "#" | grep -v "context" | grep -v "max_tokens" | grep -v "temperature" | grep -v "stream" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
if [[ -n "$model" ]]; then
echo "Found model in $abs_file: $model"
# Use override context window if specified in config.md
if [[ -n "$override_context" && "$override_context" != "auto" ]]; then
echo "Using override context window from config.md: $override_context"
local override_kb
override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true)
if [[ -n "$override_kb" ]]; then
echo "Context window: ${override_context} tokens (override)"
echo "$((override_kb * 1000))"
return
fi
fi
detect_model_context "$model"
return
fi
fi
done
echo "Model: Unknown (could not detect from AGENT.md or config files)"
echo "0"
}
# ─── Agent Framework Overhead Calculation ───
calculate_overhead() {
local project_dir="${1:-.}"
local overhead_tokens=0
# Count tokens for the framework files that are loaded during orchestration
# These are the files loaded during the most common phase (orchestration):
# AGENT.md + RULES.md + workflow.md + orchestrate.md
# Note: other phase files (decompose.md, implement.md, etc.) are only loaded during
# their specific phases, so they don't contribute to the peak context during orchestration.
local framework_files=(
"${project_dir}/.agent-framework/AGENT.md"
"${project_dir}/.agent-framework/RULES.md"
"${project_dir}/.agent-framework/prompts/workflow.md"
"${project_dir}/.agent-framework/prompts/orchestrate.md"
)
# Fallback: check home directory if project dir doesn't have framework
if [[ ! -f "${project_dir}/.agent-framework/AGENT.md" ]]; then
framework_files=(
"${HOME}/.agent-framework/AGENT.md"
"${HOME}/.agent-framework/RULES.md"
"${HOME}/.agent-framework/prompts/workflow.md"
"${HOME}/.agent-framework/prompts/orchestrate.md"
)
fi
for file in "${framework_files[@]}"; do
if [[ -f "$file" ]]; then
# Rough estimate: 1 token ≈ 4 characters (English text)
local chars
chars=$(wc -c < "$file" 2>/dev/null || echo 0)
local tokens=$((chars / 4))
overhead_tokens=$((overhead_tokens + tokens))
echo " ${file##*/}: ~${tokens} tokens"
fi
done
echo "Framework overhead: ~${overhead_tokens} tokens"
echo "$overhead_tokens"
}
# ─── Recommendation Engine ───
recommend_context() {
local gpu_vram_gb="$1"
local ram_gb="$2"
local model_context_kb="$3"
local overhead_tokens="$4"
# Read VRAM config from config.md if it exists
local config_md="${HOME}/.agent-framework/config.md"
local auto_detect="Yes"
local target_context_kb=0
local override_headroom=25
local override_max_peak_kb=0
if [[ -f "$config_md" ]]; then
auto_detect=$(grep -i "auto-detect:" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]' || true)
target_context_kb=$(grep -i "target.*context" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true)
override_headroom=$(grep -i "headroom" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true)
override_max_peak_kb=$(grep -i "max peak" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true)
fi
# If auto-detect is disabled, use the manually specified values
if [[ -n "$auto_detect" && "$auto_detect" == "No" ]]; then
if [[ -n "$target_context_kb" ]]; then
local max_peak_kb=${override_max_peak_kb:-0}
if [[ $max_peak_kb -eq 0 && $headroom_pct -gt 0 ]]; then
max_peak_kb=$((target_context_kb * (100 - headroom_pct) / 100))
fi
echo "$headroom_pct"
echo "$target_context_kb"
echo "$max_peak_kb"
return
fi
fi
local recommended_kb=0
local headroom_pct=${override_headroom:-25} # Use override from config.md, or default to 25%
# Output headroom_pct first (for parent to read)
# Then output recommended_kb
# Then output max_peak_kb (only for manual mode)
echo "$headroom_pct"
# Recommendation logic:
# 1. If GPU VRAM >= 4GB: use VRAM (practical for local inference)
# 2. If model context window is available: use it (for API inference)
# 3. If GPU VRAM < 4GB but > 0: use RAM (VRAM too small for local inference)
# 4. If no GPU VRAM and no model: use RAM as fallback
# If GPU VRAM >= 4GB, base it on VRAM
if [[ $gpu_vram_gb -ge 4 ]]; then
# Rule of thumb: 1GB VRAM ≈ 4k tokens for local LLMs
# But we need to leave room for the model itself
# For a model, each ~8k context tokens takes about ~3-5MB of GPU VRAM
# So VRAM available for context = VRAM - model size - agent overhead
# Conservative: 1GB VRAM ≈ 2k context tokens
local vram_context_kb=$((gpu_vram_gb * 2000))
# Leave headroom for the model itself and agent overhead
recommended_kb=$((vram_context_kb * (100 - headroom_pct) / 100))
# If model context window is available, use it (for API inference)
elif [[ $model_context_kb -gt 0 ]]; then
# For API-based, we're limited by the model's context window
# But we don't want to use the full window due to overhead
recommended_kb=$((model_context_kb * (100 - headroom_pct) / 100))
# Fallback: use RAM to estimate
else
# Moderate estimate for low-VRAM systems where VRAM is too small for local inference
# but RAM is available. Use 0.75k tokens per GB of RAM as a moderate estimate.
# This balances between being too conservative (0.5k/GB) and too generous (1k/GB).
local ram_context_kb=$((ram_gb * 750))
recommended_kb=$((ram_context_kb * (100 - headroom_pct) / 100))
fi
# Subtract framework overhead
local net_kb=$((recommended_kb - overhead_tokens))
if [[ $net_kb -lt 0 ]]; then
net_kb=0
fi
# Output: headroom_pct, recommended_kb
echo "$net_kb"
}
# ─── Main ───
main() {
local model_name=""
local project_dir="."
# Parse arguments
while [[ $# -gt 0 ]]; do
case "$1" in
--model|-m)
model_name="$2"
shift 2
;;
--project|-p)
project_dir="$2"
shift 2
;;
*)
# Could be model name as first argument
if [[ -z "$model_name" ]]; then
model_name="$1"
fi
shift
;;
esac
done
echo "=== VRAM / Context Detection ==="
echo ""
# Detect GPU VRAM
echo "--- GPU VRAM ---"
detect_gpu_vram
local gpu_vram_kb=0
local gpu_vram_gb=0
# Use timeout to avoid hanging
gpu_vram_kb=$(timeout 5 nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ' || true)
# Validate that vram_kb is a positive number
if [[ -z "$gpu_vram_kb" || ! "$gpu_vram_kb" =~ ^[0-9]+$ || "$gpu_vram_kb" -le 0 ]]; then
gpu_vram_kb=0
fi
# If nvidia-smi didn't work, try to detect from lspci (AMD GPUs)
if [[ $gpu_vram_kb -eq 0 ]]; then
echo " nvidia-smi failed, checking lspci for AMD GPU VRAM..."
local total_vram_mb=0
while IFS= read -r line; do
local size_num
size_num=$(echo "$line" | grep -oE '[0-9]+(M|G|K)' | head -1 || true)
if [[ -n "$size_num" ]]; then
local size_val
size_val=$(echo "$size_num" | grep -oE '[0-9]+')
local size_unit
size_unit=$(echo "$size_num" | grep -oE '(M|G|K)')
if [[ -n "$size_val" && -n "$size_unit" ]]; then
case "$size_unit" in
M) total_vram_mb=$((total_vram_mb + size_val)) ;;
G) total_vram_mb=$((total_vram_mb + size_val * 1024)) ;;
K) total_vram_mb=$((total_vram_mb + size_val / 1024)) ;;
esac
fi
fi
done < <(lspci -vnn 2>/dev/null | grep -i -A 15 -E 'VGA|3D|Display' | grep -i 'Memory at')
if [[ $total_vram_mb -gt 0 ]]; then
gpu_vram_kb=$((total_vram_mb * 1024))
gpu_vram_gb=$((total_vram_mb / 1024))
echo " AMD GPU VRAM from lspci: ${gpu_vram_gb}GB ($total_vram_mb MB)"
else
echo " No VRAM found from lspci"
fi
fi
echo ""
# Detect RAM
echo "--- RAM ---"
detect_ram
local ram_kb
ram_kb=$(grep MemTotal /proc/meminfo 2>/dev/null | awk '{print $2}' || echo 0)
local ram_gb=$((ram_kb / 1024 / 1024))
echo ""
# Detect model context window
echo "--- Model Context Window ---"
detect_model_context "$model_name"
local model_context_kb
model_context_kb=$(detect_model_context "$model_name" | tail -1)
echo ""
# Calculate framework overhead
echo "--- Framework Overhead ---"
calculate_overhead "$project_dir"
local overhead_tokens
overhead_tokens=$(calculate_overhead "$project_dir" | tail -1)
echo ""
# Recommend context window (also outputs headroom_pct and recommended_kb)
echo "--- Recommendation ---"
local recommendation_output
recommendation_output=$(recommend_context "$gpu_vram_gb" "$ram_gb" "$model_context_kb" "$overhead_tokens")
local line_count
line_count=$(echo "$recommendation_output" | wc -l)
local recommended_kb
local headroom_pct
local max_peak_kb
if [[ $line_count -ge 3 ]]; then
# Manual mode: outputs headroom_pct, recommended_kb, max_peak_kb
headroom_pct=$(echo "$recommendation_output" | head -1 | tr -d '[:space:]')
recommended_kb=$(echo "$recommendation_output" | sed -n '2p' | tr -d '[:space:]')
max_peak_kb=$(echo "$recommendation_output" | tail -1 | tr -d '[:space:]')
else
# Auto-detect mode: outputs headroom_pct, recommended_kb
headroom_pct=$(echo "$recommendation_output" | head -1 | tr -d '[:space:]')
recommended_kb=$(echo "$recommendation_output" | tail -1 | tr -d '[:space:]')
# Calculate max peak context based on headroom
max_peak_kb=$((recommended_kb * (100 - headroom_pct) / 100))
fi
# Convert to human-readable
local recommended_k
if [[ $recommended_kb -gt 0 ]]; then
recommended_k=$((recommended_kb / 1000))
else
recommended_k=8 # Default fallback
fi
echo ""
echo "=== Recommended Configuration ==="
echo "Target context: ${recommended_k}k tokens"
echo "Headroom: ${headroom_pct}%"
echo "Max peak context per sub-task: $((recommended_k * (100 - headroom_pct) / 100))k tokens"
# Output as JSON for programmatic use
echo ""
echo "=== JSON Output ==="
cat <<EOF
{
"gpu_vram_gb": $gpu_vram_gb,
"ram_gb": $ram_gb,
"model_context_kb": $model_context_kb,
"framework_overhead_tokens": $overhead_tokens,
"recommended_kb": $recommended_kb,
"recommended_k": $recommended_k,
"headroom": 0.25,
"max_peak_context_kb": $max_peak_kb
}
EOF
}
main "$@"