From a74eadfb861854519c838778da2e353d860b4f39 Mon Sep 17 00:00:00 2001 From: laptran Date: Thu, 11 Jun 2026 20:25:27 -0400 Subject: [PATCH] Split config: VRAM/model settings to config.md, agent behavior to AGENT.md --- AGENT.md | 5 +- README.md | 80 ++++- config.md | 61 ++++ contracts/vram_config.md | 11 + install.sh | 46 +++ prompts/adversarial_bug_find.md | 13 +- prompts/bug_finder.md | 2 + prompts/decompose.md | 233 ++++++++++++++ prompts/doc_review.md | 7 +- prompts/implement.md | 14 + prompts/onboarding.md | 28 ++ prompts/orchestrate.md | 234 +++++++++++++- prompts/referee.md | 3 + prompts/workflow.md | 3 +- scripts/vram_detect.sh | 549 ++++++++++++++++++++++++++++++++ 15 files changed, 1276 insertions(+), 13 deletions(-) create mode 100644 config.md create mode 100644 contracts/vram_config.md mode change 100644 => 100755 install.sh create mode 100644 prompts/decompose.md create mode 100755 scripts/vram_detect.sh diff --git a/AGENT.md b/AGENT.md index e527d74..d03ff98 100644 --- a/AGENT.md +++ b/AGENT.md @@ -3,6 +3,8 @@ ## Autopilot Autopilot: Enabled +> Note: For global framework settings (VRAM, model, system requirements), see `~/.agent-framework/config.md`. + ## Routing IF task type = research → load prompts/research.md + RULES.md @@ -13,7 +15,8 @@ IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code IF task type = adversarial_bug_find → load prompts/adversarial_bug_find.md + SPEC.md + code IF task type = referee → load prompts/referee.md + SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md IF task type = doc_review → load prompts/doc_review.md + DESIGN.md +IF task type = decompose → load prompts/decompose.md + SPEC.md IF task type = orchestrate → load prompts/orchestrate.md + project structure IF task type = compaction → load prompts/compaction.md -Always start by reading this file to determine mode. \ No newline at end of file +Always start by reading this file to determine mode. diff --git a/README.md b/README.md index 668ff40..fb70e31 100644 --- a/README.md +++ b/README.md @@ -66,7 +66,8 @@ The framework features an **Autopilot** mode that allows the agent to drive a pr ### Lifecycle of a Task 1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. **Interactive** — agent grills you for requirements and gets sign-off. -2. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off. +2. **Decomposition** (optional): Break the task into sub-tasks sized for your VRAM. **Interactive** — agent analyzes the spec and proposes sub-tasks with token budget estimates, gets sign-off. See [VRAM Configuration](#vram-configuration) below. +3. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off. 3. **Test Design** (optional): Produce a `TEST_PLAN.md` (Test specification). No code allowed. **Interactive** — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off. 4. **Implement**: Write code and tests based *only* on the `SPEC.md`, `DESIGN.md` (if present), and `TEST_PLAN.md` (if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows. 5. **Bug Find**: Aggressive search for bugs and spec deviations. @@ -79,11 +80,69 @@ The framework features an **Autopilot** mode that allows the agent to drive a pr #### Autopilot mode (default) The default mode is **Autopilot: Enabled**. The Orchestrator automatically drives tasks through all phases. Just say: - **Start a new task**: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion) +- **Decompose a task**: "Decompose the add-user-auth task" (the Orchestrator breaks it into sub-tasks sized for your VRAM) - **Continue**: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it) +#### VRAM Configuration +For low-VRAM systems (8GB, 16GB), set your VRAM limits in `~/.agent-framework/config.md`: + +```markdown +## VRAM Configuration +- **Auto-detect**: Yes # Let the agent detect your VRAM automatically +- **Target VRAM context**: 16k tokens # Override auto-detect if needed +- **Headroom**: 25% +- **Max peak context per sub-task**: 12k tokens +``` + +**Auto-detection**: When `Auto-detect: Yes`, the Orchestrator runs `scripts/vram_detect.sh` to probe: +- GPU VRAM (via `nvidia-smi`) +- System RAM (via `free`) +- Model context window (from config.md or API config files) +- Framework overhead (by counting token load in loaded prompts) + +**Manual override**: When `Auto-detect: No`, use the manually specified values: + +```markdown +## VRAM Configuration +- **Auto-detect**: No +- **Target VRAM context**: 8k +- **Headroom**: 30% +- **Max peak context per sub-task**: 5.6k +``` + +#### Model Configuration +When using a local LLM or a specific API model, set the model in `~/.agent-framework/config.md`: + +```markdown +## Model Configuration +- **Model**: auto # Use auto-detection from API config files +- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k) +``` + +**Auto-detection**: When `Model: auto`, the framework detects the model name from API config files (`.env`, `config.yaml`, etc.) and looks up its context window. + +**Manual override**: When you know your model name, specify it: + +```markdown +## Model Configuration +- **Model**: gpt-4o +- **Override context window**: 128k +``` + +When you run "Decompose the X task", the Orchestrator will: +1. Analyze the task's SPEC.md +2. Detect VRAM limits (auto or manual) +3. Break it into sub-tasks, each sized to fit within your VRAM limit +4. Estimate the token budget for each sub-task +5. Create sub-task folders under `tasks/{parent-task}/subtasks/{sub-task}/` +6. Propagate VRAM config to each sub-task + +Sub-tasks run independently through the full lifecycle. The parent task is complete only when ALL sub-tasks pass. + #### Manual mode (opt-in) Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like: - "Research add user authentication" — starts a new task +- "Decompose the add-user-auth task" — breaks into VRAM-sized sub-tasks (optional) - "Design the add-user-auth task" — designs the architecture (optional) - "Design tests for the add-user-auth task" — designs test cases (optional) - "Implement the add-user-auth task" — implements the task with tests @@ -93,12 +152,27 @@ Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you p - "Review the add-user-auth task" — referee evaluates - "orchestrate" — asks the Orchestrator what to do next +## Sub-Task Management + +When a task is decomposed, the Orchestrator creates sub-tasks under `tasks/{parent-task}/subtasks/`. Each sub-task: + +- Has its own full lifecycle (Research → Decomposition → Design → ... → Complete) +- Starts at the Research phase with an empty folder +- Receives a `PARENT_SPEC.md` with the parent task's context +- Receives a `VRAM_CONFIG.md` with the VRAM constraints +- Runs independently — sub-tasks in the same wave can run in parallel +- The parent task is NOT complete until ALL sub-tasks pass + ## Key Components -- `AGENT.md`: Project-specific configuration and mode selection. +- `AGENT.md`: Project-specific agent behavior (Autopilot mode, routing rules). +- `config.md`: Global framework settings (VRAM, model, system requirements). - `RULES.md`: Living document of project constraints and past failure modes. -- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, etc.). +- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, Decompose, etc.). - `workflow.md`: The state machine governing the Autopilot lifecycle. - `test_design.md`: Produces a TEST_PLAN.md — an explicit test specification before implementation. +- `decompose.md`: Breaks a task into VRAM-sized sub-tasks. +- `scripts/vram_detect.sh`: Auto-detects GPU VRAM, RAM, model context window, and framework overhead. +- `contracts/vram_config.md`: Contract for VRAM-aware task decomposition. ## Contact & Support [Insert Contact Info] diff --git a/config.md b/config.md new file mode 100644 index 0000000..b367d75 --- /dev/null +++ b/config.md @@ -0,0 +1,61 @@ +# Framework Configuration + +This file contains global framework settings that apply across all projects. + +## VRAM Configuration + +Settings for task decomposition based on available VRAM. + +- **Auto-detect**: Yes # Detect GPU VRAM, RAM, and model context window automatically +- **Target context**: 16k tokens # Override auto-detect if needed +- **Headroom**: 25% # Leave headroom for code, context, and reasoning +- **Max peak context per sub-task**: 12k tokens # Max context for any single sub-task + +### Auto-detection + +When `Auto-detect: Yes`, the framework probes your system to detect: +- GPU VRAM (via `nvidia-smi` or `lspci`) +- System RAM (via `free`) +- Model context window (via API config or model name lookup) +- Framework overhead (by reading all loaded prompt files) + +To disable auto-detection and use manual values: + +``` +## VRAM Configuration +- **Auto-detect**: No +- **Target context**: 8k +- **Headroom**: 30% +- **Max peak context per sub-task**: 5.6k +``` + +## Model Configuration + +Settings for the LLM model being used. + +- **Model**: auto # Use auto-detection from API config, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet) +- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k) + +### Auto-detection + +When `Model: auto`, the framework detects the model name from: +1. `AGENT.md` in the project (if specified there) +2. API config files (`.env`, `config.yaml`, `config.json`, etc.) +3. Model name lookup by name (e.g., gpt-4o → 128k, claude-3-5-sonnet → 200k) + +To disable auto-detection and use manual values: + +``` +## Model Configuration +- **Model**: gpt-4o +- **Override context window**: 128k +``` + +## System Requirements + +Requirements for the environment the framework runs in. + +- **nvidia-smi**: Required if NVIDIA GPU (for VRAM detection) +- **lspci**: Fallback if NVIDIA GPU not available (for AMD GPU VRAM detection) +- **/proc/meminfo**: Required for RAM detection (Linux) +- **sysctl**: Fallback for RAM detection (macOS) diff --git a/contracts/vram_config.md b/contracts/vram_config.md new file mode 100644 index 0000000..1802e78 --- /dev/null +++ b/contracts/vram_config.md @@ -0,0 +1,11 @@ +## VRAM Configuration Contract + +### Acceptance Criteria +- [ ] Target VRAM context is specified +- [ ] Headroom is specified (recommended: 25-40%) +- [ ] Max peak context per sub-task is calculated +- [ ] Sub-tasks are sized to fit within the max peak context +- [ ] VRAM_CONFIG.md is propagated to each sub-task + +### Stop Condition +When all checkboxes are checked, output "CONTRACT_MET" and stop. diff --git a/install.sh b/install.sh old mode 100644 new mode 100755 index c83c447..4443291 --- a/install.sh +++ b/install.sh @@ -12,5 +12,51 @@ fi echo "Cloning agent-framework to $FRAMEWORK_DIR..." git clone https://gitea.yourdomain.com/you/agent-framework.git "$FRAMEWORK_DIR" +echo "" +echo "=== VRAM / Context Detection ===" +echo "Detecting your system's VRAM to recommend task decomposition settings..." +echo "" + +# Run VRAM detection script if it exists +if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.sh" ]; then + # Run in project-dir context so it can read framework overhead + detection_output=$(cd "$FRAMEWORK_DIR" && bash "$FRAMEWORK_DIR/scripts/vram_detect.sh" 2>&1) + + # Extract JSON output (last section after "=== JSON Output ===") + json_output=$(echo "$detection_output" | sed -n '/=== JSON Output ===/,/EOF/p' | grep -v '=== JSON Output ===' | grep -v '^EOF$') + + if [ -n "$json_output" ]; then + echo "$detection_output" + + # Extract key values from JSON for display + recommended_k=$(echo "$json_output" | grep '"recommended_k"' | grep -oP '\d+') + max_peak_kb=$(echo "$json_output" | grep '"max_peak_context_kb"' | grep -oP '\d+') + headroom=$(echo "$json_output" | grep '"headroom"' | grep -oP '\d+\.\d+') + gpu_vram=$(echo "$json_output" | grep '"gpu_vram_gb"' | grep -oP '\d+') + ram_gb=$(echo "$json_output" | grep '"ram_gb"' | grep -oP '\d+') + model_context=$(echo "$json_output" | grep '"model_context_kb"' | grep -oP '\d+') + + echo "" + echo "=== Recommended VRAM Configuration ===" + echo "For low-VRAM systems (8GB, 16GB VRAM), add this to ~/.agent-framework/config.md:" + echo "" + echo "## VRAM Configuration" + echo "- **Auto-detect**: Yes # Let the agent detect automatically" + echo "- **Target context**: ${recommended_k}k tokens # Override auto-detect if needed" + echo "- **Headroom**: ${headroom}%" + echo "- **Max peak context per sub-task**: $((max_peak_kb / 1000))k tokens" + echo "" + echo "This ensures tasks are decomposed into sub-tasks that fit within your" + echo "available VRAM. For more information, see the README." + else + echo "Could not detect VRAM. You can manually set your VRAM configuration in ~/.agent-framework/config.md." + echo "See the README for details." + fi +else + echo "VRAM detection script not found. You can manually set your VRAM configuration in ~/.agent-framework/config.md." + echo "See the README for details." +fi + +echo "" echo "Installation complete." echo "Next step: cd into a project and run the onboarding prompt." \ No newline at end of file diff --git a/prompts/adversarial_bug_find.md b/prompts/adversarial_bug_find.md index 18a1e19..64dbb5e 100644 --- a/prompts/adversarial_bug_find.md +++ b/prompts/adversarial_bug_find.md @@ -1,9 +1,20 @@ You are the Adversarial Bug Finder. -Read the SPEC.md and the code. +## Read These Files + +1. {project}/tasks/{task-name}/SPEC.md +2. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists) +3. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists) +4. The code Your job is to find bugs that are difficult to spot, such as complex logic errors, race conditions, and performance bottlenecks. Be more aggressive and exhaustive than a standard bug finder. +If VRAM_CONFIG.md exists, also check for: +- Memory leaks (loading large files into context that could cause OOM) +- N+1 query patterns that could cause memory exhaustion +- Infinite loops that could run out of context +- Unbounded recursion that could cause stack overflow + Output your findings in ADVERSARIAL_BUG_REPORT.md. When finished, output "ADVERSARIAL_BUG_FIND_COMPLETE". \ No newline at end of file diff --git a/prompts/bug_finder.md b/prompts/bug_finder.md index 1969770..f4c6fb3 100644 --- a/prompts/bug_finder.md +++ b/prompts/bug_finder.md @@ -5,6 +5,8 @@ You are the Bug Finder. Your job is to find every bug, deviation from spec, and 1. {project}/tasks/{task-name}/SPEC.md 2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists) 3. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists) +4. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists) +5. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists) ## Task diff --git a/prompts/decompose.md b/prompts/decompose.md new file mode 100644 index 0000000..c795106 --- /dev/null +++ b/prompts/decompose.md @@ -0,0 +1,233 @@ +You are in decomposition mode. + +Your only job is to take a completed SPEC.md and break it into the smallest possible, independently verifiable sub-tasks. Each sub-task should be small enough to complete in one session and should have clear, testable acceptance criteria. + +## Read These Files + +1. {project}/tasks/{task-name}/SPEC.md +2. {project}/.agent-framework/RULES.md +3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings) +4. {project}/.agent-framework/AGENT.md (if exists — for model override) +5. {project}/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection) + +## Task + +{task-description} + +## Decomposition Rules + +### Rule 1: Smallest Possible Unit +Break every feature into the smallest possible units that are still independently testable. If a sub-task can be done in one session, it should be a sub-task. + +### Rule 2: Each Sub-Task Must Be Self-Contained +Each sub-task must have: +- Its own goal statement (one sentence) +- Clear acceptance criteria (at least 2-3) +- Dependencies on other sub-tasks (if any) +- Its own contract (SPEC.md) that references the parent task + +### Rule 3: Dependencies Must Be Explicit +If sub-task B depends on sub-task A, state it clearly. A sub-task with no dependencies can run in parallel. Sub-tasks that depend on parallel sub-tasks must wait for both. + +### Rule 4: Define Execution Order +After listing all sub-tasks, define the execution order considering dependencies. Group independent sub-tasks into "waves" that can be done in parallel. + +### Rule 5: Do Not Create Sub-Sub-Tasks +Decomposition produces **one level** of sub-tasks only. If a sub-task is still too large, the implementer should break it down further during implementation, but do NOT nest sub-tasks. + +### Rule 6: Token Budget Per Sub-Task +Each sub-task must fit within the target VRAM context window. Estimate the total token budget for each sub-task's full lifecycle and break it down further if it exceeds the limit. + +**How to estimate token budget for a sub-task:** + +During a sub-task's lifecycle, the following files are loaded into context at various phases: + +- **Research phase**: RULES.md + AGENT.md + task description +- **Design phase**: SPEC.md + RULES.md +- **Implement phase**: SPEC.md + DESIGN.md + TEST_PLAN.md + RULES.md + AGENT.md + CONTRACT.md +- **Bug Find phase**: SPEC.md + code (limited scope) +- **Adversarial Bug Find phase**: SPEC.md + code (limited scope) +- **Doc Review phase**: DESIGN.md +- **Referee phase**: SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md + +The **peak context** is during the Implement phase, where all files are loaded together. Estimate the token count of the combined files for the sub-task. + +**Guidelines:** +- **8k VRAM**: Peak context for Implement phase should be ≤ 6k tokens (leave 2k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 6k tokens. +- **16k VRAM**: Peak context for Implement phase should be ≤ 12k tokens (leave 4k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 12k tokens. +- **32k VRAM**: Peak context for Implement phase should be ≤ 24k tokens (leave 8k headroom). +- **64k VRAM**: Peak context for Implement phase should be ≤ 48k tokens (leave 16k headroom). + +If a sub-task's estimated context exceeds the limit, break it into smaller sub-tasks. + +**How to estimate token count:** +- Roughly 1 token = 4 characters (for English text) +- A SPEC.md with 5 requirements, each with 2-3 acceptance criteria, is typically 500-1000 tokens +- A DESIGN.md with 3 sections and 5-10 bullet points is typically 1000-3000 tokens +- A TEST_PLAN.md with 5-10 test cases is typically 1000-2000 tokens +- RULES.md is typically 200-1000 tokens (varies per project) +- AGENT.md is typically 300-1000 tokens +- A CONTRACT.md is typically 200-500 tokens + +**Quick estimate formula:** +``` +Peak context ≈ SPEC.md tokens + DESIGN.md tokens + TEST_PLAN.md tokens + RULES.md tokens + AGENT.md tokens + CONTRACT.md tokens +``` + +### Rule 7: Sub-Task Size Targets +Aim for sub-tasks that are: +- **Small** (8k VRAM): ~200-800 tokens of combined spec/design/test files +- **Small** (16k VRAM): ~200-1500 tokens of combined spec/design/test files +- **Medium** (8k VRAM): ~800-2000 tokens of combined spec/design/test files +- **Medium** (16k VRAM): ~1500-4000 tokens of combined spec/design/test files +- **Large** (8k VRAM): ~2000-4000 tokens of combined spec/design/test files +- **Large** (16k VRAM): ~4000-8000 tokens of combined spec/design/test files + +If a sub-task exceeds the "Large" target for the target VRAM, break it further. + +## Decomposition Protocol (Interactive) + +### Phase 1: Analysis + +Before decomposing, analyze the SPEC.md: +1. Identify all distinct features/requirements +2. Identify data models that need to be created +3. Identify API endpoints or interfaces +4. Identify infrastructure changes +5. Identify configuration changes +6. Estimate the token budget for the full task (sum of all requirements' SPEC + DESIGN + TEST files) +7. **Detect VRAM limits**: + - Check `~/.agent-framework/config.md` for VRAM Configuration section + - If `Auto-detect: Yes`, run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window + - If `Auto-detect: No`, use the manually specified values from config.md + - Report the detected VRAM limits +8. **Detect model context window**: + - Check `~/.agent-framework/config.md` for Model Configuration section + - If `Model: auto`, run the detection script to detect the model name and its context window + - If `Override context window: auto`, use the detected context window + - If both are specified, use the specified values + - If model detection fails, use 128k tokens as default +9. Determine the target VRAM context window based on the detection results + +### Phase 2: Propose Decomposition + +Present a draft decomposition to the user. Format: + +**Target VRAM**: {8k/16k/32k/64k} tokens + +**Waves:** +- **Wave 1**: Sub-task 1, Sub-task 2, Sub-task 3 (can run in parallel) +- **Wave 2**: Sub-task 4, Sub-task 5 (depends on Wave 1) +- **Wave 3**: Sub-task 6 (depends on Wave 2) + +**Sub-task Details:** +1. **{sub-task-name}** + - Goal: {one sentence} + - Dependencies: {list of sub-task names, or "None"} + - Acceptance criteria: + - [ ] {criterion 1} + - [ ] {criterion 2} + - Estimated scope: {small/medium/large} + - **Estimated token budget**: ~{estimate} tokens (peak: ~{peak} tokens during Implement phase) + - **Fits within VRAM**: Yes/No (if No, explain why and suggest how to split further) + +### Phase 3: Review and Refine + +Present the draft decomposition to the user and ask: +- "Are there any sub-tasks that are too large?" +- "Are there any sub-tasks that should be combined?" +- "Are the dependencies correct?" +- "Are there any sub-tasks I missed?" +- "Is the execution order optimal?" +- "Do the token budget estimates look reasonable for your VRAM?" +- "Are there any sub-tasks that exceed your VRAM limit?" + +Incorporate the user's feedback and revise the decomposition accordingly. + +### Phase 4: Get Sign-Off + +Before writing the DECOMPOSITION.md, you MUST get explicit sign-off from the user. Say: + +> "Based on our discussion, here is the final decomposition: +> [brief summary of waves, sub-tasks, and token budgets] +> Does this cover everything? Please confirm with 'APPROVED' before I finalize." + +Only produce the DECOMPOSITION.md after the user says "APPROVED" or equivalent. + +## Output + +Produce a file called DECOMPOSITION.md at {project}/tasks/{task-name}/DECOMPOSITION.md that contains: + +```markdown +# Task Decomposition + +## Parent Task +{parent-task-name} + +## VRAM Configuration +- **Target VRAM**: {8k/16k/32k/64k} tokens +- **Headroom**: {25-40%} (leaves headroom for code, context, and reasoning) +- **Max peak context per sub-task**: {estimate} tokens + +## Waves + +### Wave 1: {wave-name} +- {sub-task-name-1} +- {sub-task-name-2} +- {sub-task-name-3} + +### Wave 2: {wave-name} +- {sub-task-name-4} +- {sub-task-name-5} + +## Sub-Task Details + +### 1. {sub-task-name-1} +- **Goal**: {one sentence} +- **Dependencies**: None (or list sub-task names) +- **Acceptance Criteria**: + - [ ] {criterion 1} + - [ ] {criterion 2} +- **Estimated Scope**: {small/medium/large} +- **Estimated token budget**: ~{estimate} tokens + - SPEC.md: ~{x} tokens + - DESIGN.md: ~{x} tokens + - TEST_PLAN.md: ~{x} tokens + - RULES.md: ~{x} tokens + - AGENT.md: ~{x} tokens + - CONTRACT.md: ~{x} tokens +- **Peak context (Implement phase)**: ~{peak} tokens +- **Fits within VRAM**: Yes + +### 2. {sub-task-name-2} +- **Goal**: {one sentence} +- **Dependencies**: {list of sub-task names} +- **Acceptance Criteria**: + - [ ] {criterion 1} + - [ ] {criterion 2} +- **Estimated Scope**: {small/medium/large} +- **Estimated token budget**: ~{estimate} tokens + - SPEC.md: ~{x} tokens + - DESIGN.md: ~{x} tokens + - TEST_PLAN.md: ~{x} tokens + - RULES.md: ~{x} tokens + - AGENT.md: ~{x} tokens + - CONTRACT.md: ~{x} tokens +- **Peak context (Implement phase)**: ~{peak} tokens +- **Fits within VRAM**: Yes + +... etc ... + +## Execution Order +1. Complete Wave 1 (all sub-tasks can run in parallel) +2. Complete Wave 2 (depends on Wave 1) +3. ... etc ... +``` + +When the decomposition is complete, output "CONTRACT_MET" and stop. + +Do not create sub-task folders or files. The Orchestrator will handle creating sub-task folders based on this DECOMPOSITION.md. + +## Stop Condition (MANDATORY) +You are not allowed to end this session until you have produced the DECOMPOSITION.md file AND output the exact phrase "CONTRACT_MET". +Until then, continue working or ask clarifying questions. diff --git a/prompts/doc_review.md b/prompts/doc_review.md index d6cbd57..420a41d 100644 --- a/prompts/doc_review.md +++ b/prompts/doc_review.md @@ -4,8 +4,10 @@ You are in Documentation Review mode. 1. {project}/tasks/{task-name}/DESIGN.md — Look for the "Documentation Plan" section 2. {project}/tasks/{task-name}/SPEC.md — Check what the spec requires -3. Any existing documentation files mentioned in the DESIGN.md Documentation Plan -4. The code that was implemented (implementation artifacts) +3. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists) +4. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists) +5. Any existing documentation files mentioned in the DESIGN.md Documentation Plan +6. The code that was implemented (implementation artifacts) ## Task @@ -36,6 +38,7 @@ You are in Documentation Review mode. ### Documentation Gaps - Are there any areas where the documentation is thin or missing? - Are there any complex flows or non-obvious logic that should be documented? +- If VRAM_CONFIG.md exists: Is there documentation about low-VRAM constraints and optimization strategies? ## Output diff --git a/prompts/implement.md b/prompts/implement.md index 53cbf64..d2bdd6e 100644 --- a/prompts/implement.md +++ b/prompts/implement.md @@ -8,6 +8,8 @@ You are in implementation mode. 4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists) 5. {project}/tasks/{task-name}/DESIGN.md (if exists) 6. {project}/tasks/{task-name}/TEST_PLAN.md (if exists) +7. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists — for low-VRAM systems) +8. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists — for sub-tasks) ## Task @@ -26,6 +28,18 @@ You are in implementation mode. - Keep functions small and focused. - Use existing patterns in the codebase. +### VRAM-Aware Implementation (if VRAM_CONFIG.md exists) + +If `VRAM_CONFIG.md` is present, the implementer is working in a low-VRAM environment and must: + +- **Keep functions small**: Each function should be ≤ 50 lines. Break large functions into smaller ones. +- **Avoid loading large files into context**: If implementing in an iterative environment (like pi), read files incrementally rather than loading entire files at once. +- **Write self-contained modules**: Each module should be independently testable and runnable without loading the entire codebase. +- **Prefer streaming over buffering**: Use streaming patterns instead of loading entire datasets into memory. +- **Be explicit about dependencies**: Import only what you need, not the entire module. +- **Test incrementally**: Run tests frequently rather than building up a large test suite that takes a long time to run. +- **Optimize for low-VRAM**: Write code that can be understood and modified by an agent with limited context window. + ### End State - You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met - All tests must pass. diff --git a/prompts/onboarding.md b/prompts/onboarding.md index 8319ca8..ac7b370 100644 --- a/prompts/onboarding.md +++ b/prompts/onboarding.md @@ -8,6 +8,7 @@ Your only job is to set up the minimal agent framework structure in the target p 2. ~/.agent-framework/ONBOARDING.md — human reference for drop-in vs from-scratch scenarios 3. {project}/.agent-framework/AGENT.md (if it exists) 4. {project}/.agent-framework/RULES.md (if it exists) +5. ~/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection) ## Task @@ -36,6 +37,32 @@ Before proceeding, check if the project is running an older version of the frame 5. Explore the project root at a high level (ls, key directories, README if present). 6. Produce a short onboarding report. +### Step 2: VRAM Configuration + +Check if VRAM configuration is available in `~/.agent-framework/config.md`: +1. Read `~/.agent-framework/config.md` to check for VRAM Configuration section. +2. If VRAM Configuration section exists, note the values. +3. If VRAM Configuration section does NOT exist, check if VRAM detection is available: `~/.agent-framework/scripts/vram_detect.sh`. +4. If available, run it to get VRAM recommendations: + ``` + cd ~/.agent-framework && bash ~/.agent-framework/scripts/vram_detect.sh + ``` +5. Parse the JSON output for `recommended_k`, `max_peak_context_kb`, and `headroom`. +6. Add a VRAM Configuration section to `~/.agent-framework/config.md`: + ``` + ## VRAM Configuration + - **Auto-detect**: Yes + - **Target context**: {recommended_k}k tokens + - **Headroom**: 25% + - **Max peak context per sub-task**: {max_peak_kb/1000}k tokens + ``` +7. If VRAM detection failed or is not available, add a minimal section: + ``` + ## VRAM Configuration + - **Auto-detect**: Yes + ``` +8. Report the VRAM configuration status in the onboarding report. + ## Output Create or update the following inside {project}/.agent-framework/: @@ -50,6 +77,7 @@ Then produce a file called ONBOARDING_REPORT.md at {project}/tasks/onboarding/ON - Key observations from the project structure - Any missing pieces the human should provide next - **Upgrade status**: Whether the project's framework files are up to date with the global framework +- **VRAM Configuration**: Whether VRAM config was auto-detected and applied (if available), or needs manual setup When the ritual is complete, output "CONTRACT_MET" and stop. diff --git a/prompts/orchestrate.md b/prompts/orchestrate.md index 97e1b42..5a57d70 100644 --- a/prompts/orchestrate.md +++ b/prompts/orchestrate.md @@ -4,8 +4,113 @@ You are the Orchestrator Driver. Your job is to act as a **state machine** for t 1. {project}/.agent-framework/AGENT.md — Check if Autopilot is enabled 2. {project}/.agent-framework/RULES.md -3. {project}/.agent-framework/prompts/workflow.md — The State Machine -4. Any existing files under {project}/tasks/ +3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings) +4. {project}/.agent-framework/AGENT.md (if exists — for model override) +5. {project}/.agent-framework/prompts/workflow.md — The State Machine +6. Any existing files under {project}/tasks/ +7. {project}/.agent-framework/scripts/vram_detect.sh — VRAM detection script (if exists) + +## VRAM Detection + +When VRAM configuration is needed (during task decomposition, sub-task creation, etc.), the Orchestrator MUST attempt to detect VRAM/VRAM limits dynamically. + +### Detection Priority + +1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`. +2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window. +3. **Manual override**: Check if `~/.agent-framework/config.md` has `Auto-detect: No` under VRAM Configuration. If so, use the manually specified values. +4. **Fallback**: Use 8k tokens as default, with 25% headroom. + +### How to Read VRAM Config from config.md + +```markdown +## VRAM Configuration +- **Auto-detect**: Yes/No +- **Target context**: {value}k tokens (override if Auto-detect: No) +- **Headroom**: {value}% (override if Auto-detect: No) +- **Max peak context per sub-task**: {value}k tokens (override if Auto-detect: No) +``` + +- If `Auto-detect: Yes`, run the detection script and use its output. +- If `Auto-detect: No`, use the manually specified values. +- If config.md has no VRAM Configuration section, default to `Auto-detect: Yes`. + +### Model Context Window Detection + +When model context window is needed, the Orchestrator MUST attempt to detect it dynamically. + +#### Detection Priority + +1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to detect the model name and its context window. Parse the JSON output for `model_context_kb`. +2. **Auto-detect via config**: Check `~/.agent-framework/config.md` for the model name and override context window. +3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window. +4. **Fallback**: Use 128k tokens as default (common for modern models). + +#### How to Read Model Config from config.md + +```markdown +## Model Configuration +- **Model**: auto # Use auto-detection, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet) +- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k) +``` + +- If `Model: auto`, detect the model name from API config files or AGENT.md. +- If `Override context window: auto`, use the detected context window. +- If both are specified, use the specified values. + +#### Model Name Lookup + +When the model name is detected, look up its context window: +- gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4: 128k tokens +- claude-3-5-sonnet, claude-3-5-haiku, claude-3-opus, claude-3-sonnet, claude-3-haiku, claude-2: 200k tokens + +### Detection Script Output (JSON) + +The detection script outputs JSON like: +```json +{ + "gpu_vram_gb": 8, + "ram_gb": 16, + "model_context_kb": 128000, + "framework_overhead_tokens": 4000, + "recommended_kb": 16000, + "recommended_k": 16, + "headroom": 0.25, + "max_peak_context_kb": 12000 +} +``` + +Use `recommended_kb` for the target context, `headroom` for headroom, and `max_peak_context_kb` for max peak context per sub-task. + +### Auto-Detect When to Run Detection + +The Orchestrator should run VRAM detection in the following scenarios: + +1. **When a new task is created** — to set the VRAM config for the new task. +2. **When Decomposition is triggered** — to ensure sub-tasks are sized correctly. +3. **When sub-tasks are created** — to propagate VRAM config to sub-task folders. +4. **When a sub-task's VRAM_CONFIG.md is missing** — to create one with auto-detected values. + +### Reporting Detection Results + +When auto-detecting VRAM, the Orchestrator should report: +- GPU VRAM detected (if any) +- System RAM detected +- Model context window detected (if any) +- Framework overhead estimated +- Recommended VRAM context window +- Whether auto-detection was used or manual override + +Example: +``` +VRAM Detection Results: +- GPU VRAM: 8GB (nvidia-smi) +- RAM: 16GB +- Model context window: 128k (API-based) +- Framework overhead: ~4k tokens +- **Recommended: 16k tokens** (GPU VRAM-based, 25% headroom) +- **Using: 16k tokens** (auto-detected) +``` ## Task @@ -22,7 +127,8 @@ Each task is a state machine. The Orchestrator determines the current state and | State | Condition | Next State (Autopilot) | |-------|-----------|----------------------| | **New** | No artifacts in task folder | Research | -| **Research** | Has `SPEC.md` | Design (optional) or Implement | +| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement | +| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research | | **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | | **Test Design** | Has `TEST_PLAN.md` | Implement | | **Implement** | Has `IMPLEMENTATION.md` | Bug Find | @@ -113,8 +219,9 @@ Examine the tasks/ directory and determine the state of each task folder. Check 9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement** 10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design) 11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design) -12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design) -13. No artifacts → **New** +12. Has `SPEC.md` and `DECOMPOSITION.md` → **Decomposition** (sub-tasks will be created) +13. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design) +14. No artifacts → **New** ## Output Format @@ -172,3 +279,120 @@ In Autopilot mode, after outputting the task statuses, the Orchestrator MUST aut When finished, output "ORCHESTRATION_COMPLETE". Only recommend one phase at a time. Do not suggest running multiple phases in parallel. + +## Sub-Task Management + +When a parent task has a `DECOMPOSITION.md` file, the Orchestrator must create sub-task folders for each sub-task listed. + +### Sub-Task Folder Structure + +When decomposing, the Orchestrator creates sub-tasks under the parent task's `subtasks/` folder: + +``` +tasks/parent-task/ → Parent task (Research → Decomposition → complete) + SPEC.md + DECOMPOSITION.md + subtasks/ + subtask-a/ → Sub-task (full lifecycle independently) + SPEC.md + DESIGN.md + IMPLEMENTATION.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md + VERDICT.md + subtask-b/ + SPEC.md + ... +``` + +### Sub-Task Creation Rules + +When a parent task reaches the **Decomposition** phase (has `SPEC.md` and `DECOMPOSITION.md`): + +1. **Read `~/.agent-framework/config.md`** to get the VRAM configuration and check if auto-detect is enabled. +2. **If Auto-detect: Yes**, run `{project}/.agent-framework/scripts/vram_detect.sh` to detect VRAM limits. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`. Report the detection results. +3. **If Auto-detect: No**, use the manually specified values from config.md. +4. **Read `DECOMPOSITION.md`** to extract all sub-task names, dependencies, and their estimated token budgets. +5. **Verify VRAM constraints**: + - For each sub-task, check if its estimated peak context (from DECOMPOSITION.md) is within the max peak context. + - If a sub-task exceeds the VRAM limit, create a warning and suggest splitting it further. +6. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`: + - For each sub-task, create an empty folder (no files yet — the sub-task starts at the Research phase). + - The sub-task name should be a kebab-case version of the sub-task description from DECOMPOSITION.md. +7. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration. +8. **Create a PARENT_SPEC.md** file for each sub-task with: + - The parent task's SPEC.md content + - A reference to the parent task name + - The VRAM configuration (auto-detected or manual) + - Any context the sub-task needs from the parent +9. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently. +10. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention). + +### Sub-Task Lifecycle + +Each sub-task follows the full lifecycle independently: +- Starts at the **Research** phase (no artifacts in the sub-task folder) +- Goes through Research → Decomposition (optional) → Design (optional) → Test Design (optional) → Implement → Bug Find → Adversarial Bug Find → Doc Review → Referee +- Ends at **Complete** (VERDICT.md with PASS) or **Human Intervention** (VERDICT.md with FAIL/NEEDS_REVIEW) + +### VRAM-Aware Sub-Task Splitting + +If a sub-task's estimated peak context exceeds the VRAM limit from AGENT.md: +1. Split the sub-task into smaller sub-tasks. +2. Each new sub-task should fit within the VRAM limit. +3. Update the DECOMPOSITION.md to reflect the new sub-tasks. +4. Create the new sub-task folders. +5. Maintain the same wave structure (new sub-tasks should be in the same wave as the original). + +### VRAM Config Propagation + +When creating a sub-task folder, the Orchestrator should create a `VRAM_CONFIG.md` file with: +```markdown +# VRAM Configuration for this sub-task + +- **Auto-detect**: Yes/No +- **Target VRAM context**: {from detection script or config.md override} +- **Headroom**: {from detection script or config.md override} +- **Max peak context per sub-task**: {from detection script or config.md override} +- **GPU VRAM detected**: {value}GB (or "None") +- **RAM detected**: {value}GB +- **Model context window**: {value}k tokens (or "Unknown") +- **Framework overhead**: ~{value} tokens +- **This sub-task's estimated peak context**: {from DECOMPOSITION.md} +- **Fits within VRAM**: Yes/No +``` + +This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly. + +### Sub-Task Parent Specification + +When creating a sub-task folder, the Orchestrator should also create a `PARENT_SPEC.md` file that contains: +- The parent task's SPEC.md content +- A reference to the parent task name +- Any context the sub-task needs from the parent + +This ensures sub-tasks have all the information they need to implement their portion of the parent spec. + +### Sub-Task Dependencies and Wave Management + +Sub-tasks that are in the same wave (can run in parallel) should be driven through the lifecycle in parallel. Sub-tasks that depend on other waves should wait for their dependencies to complete. + +### Sub-Task Completion and Parent Task + +When a sub-task reaches a terminal state: +- If the sub-task **PASSes**: The parent task can move to the next wave of sub-tasks. +- If a sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator pauses and reports human intervention is required. The parent task is stuck until the failing sub-task is resolved. + +When ALL sub-tasks are in terminal state: +- If ALL sub-tasks **PASS**: The parent task is **Complete**. +- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks and reports human intervention is required. + +### Sub-Task Verdict Reporting + +When a sub-task reaches the Referee phase, the VERDICT.md should include: +- The sub-task's own verdict (PASS/FAIL/NEEDS_REVIEW) +- A reference to the parent task name +- Any findings that affect the parent task + +The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status. diff --git a/prompts/referee.md b/prompts/referee.md index 6232a07..7b23661 100644 --- a/prompts/referee.md +++ b/prompts/referee.md @@ -10,6 +10,8 @@ You are the Referee. Your job is to objectively evaluate whether the implementat 6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists) 7. {project}/tasks/{task-name}/DESIGN.md (if exists) 8. {project}/tasks/{task-name}/TEST_PLAN.md (if exists) +9. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists) +10. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists) ## Task @@ -38,6 +40,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat - Are functions small and focused? - Is there proper error handling? - Are there any obvious performance issues? +- If VRAM_CONFIG.md exists: Is the code optimized for low-VRAM (small functions, streaming patterns, no large file loading)? ### Testing - Do all tests pass? diff --git a/prompts/workflow.md b/prompts/workflow.md index 5a357b9..85223fc 100644 --- a/prompts/workflow.md +++ b/prompts/workflow.md @@ -7,7 +7,8 @@ This file defines the linear progression of a task in the agent-framework. The O | Current State | Signal (Artifact) | Next Phase | Action | | :--- | :--- | :--- | :--- | | **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` | -| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code | +| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement | Generate `DECOMPOSITION.md` or `DESIGN.md` or code | +| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research | Orchestrator creates sub-task folders | | **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code | | **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests | | **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` | diff --git a/scripts/vram_detect.sh b/scripts/vram_detect.sh new file mode 100755 index 0000000..0181d88 --- /dev/null +++ b/scripts/vram_detect.sh @@ -0,0 +1,549 @@ +#!/usr/bin/env bash +# VRAM/Context Detection Script +# Detects GPU VRAM, system RAM, and model context window to recommend +# a safe VRAM context window for task decomposition. +# +# Usage: ./vram_detect.sh [model_name] +# - If model_name is provided, looks up its context window +# - Otherwise, tries to detect from API config or config.md + +set -uo pipefail # Don't exit on error - we want to continue even if detection fails + +# ─── GPU VRAM Detection ─── +detect_gpu_vram() { + local total_vram_kb=0 + local vram_per_gpu_kb=0 + local num_gpus=0 + + # Try nvidia-smi first (NVIDIA GPUs) + if command -v nvidia-smi &>/dev/null; then + local vram_kb + # Use timeout to avoid hanging on nvidia-smi (e.g., driver not loaded) + vram_kb=$(timeout 5 nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ' || true) + # Validate that vram_kb is a positive number + if [[ -n "$vram_kb" && "$vram_kb" =~ ^[0-9]+$ && "$vram_kb" -gt 0 ]]; then + total_vram_kb=$((vram_kb * 1024)) # MB → KB + vram_per_gpu_kb=$((total_vram_kb / (num_gpus+1))) + num_gpus=1 + echo "GPU: NVIDIA (nvidia-smi available)" + echo "VRAM per GPU: $((vram_kb / 1024))GB ($vram_kb MB)" + else + echo "GPU: NVIDIA (nvidia-smi available but driver not responding)" + fi + fi + + # Fallback: lspci + if [[ $total_vram_kb -eq 0 && $num_gpus -eq 0 ]]; then + local gpu_info + gpu_info=$(lspci 2>/dev/null | grep -i -E 'VGA|3D|Display' | head -5) + if [[ -n "$gpu_info" ]]; then + echo "GPU detected: $gpu_info" + # Try to get VRAM from lspci -vnn memory regions + # GPUs show VRAM as Memory regions in lspci + # Parse patterns like: Memory at f800000000 (64-bit, prefetchable) [size=256M] + local total_vram_mb=0 + while IFS= read -r line; do + # Extract the size value from [size=256M] pattern + local size_num + size_num=$(echo "$line" | grep -oE '[0-9]+(M|G|K)' | head -1 || true) + if [[ -n "$size_num" ]]; then + local size_val + size_val=$(echo "$size_num" | grep -oE '[0-9]+') + local size_unit + size_unit=$(echo "$size_num" | grep -oE '(M|G|K)') + if [[ -n "$size_val" && -n "$size_unit" ]]; then + case "$size_unit" in + M) total_vram_mb=$((total_vram_mb + size_val)) ;; + G) total_vram_mb=$((total_vram_mb + size_val * 1024)) ;; + K) total_vram_mb=$((total_vram_mb + size_val / 1024)) ;; + esac + fi + fi + done < <(lspci -vnn 2>/dev/null | grep -i -A 15 -E 'VGA|3D|Display' | grep -i 'Memory at') + if [[ $total_vram_mb -gt 0 ]]; then + local total_vram_gb=$((total_vram_mb / 1024)) + local total_vram_mb_remain=$((total_vram_mb % 1024)) + echo "VRAM: $total_vram_mb MB ($total_vram_gb GB $total_vram_mb_remain MB)" + else + echo "VRAM: Could not determine from lspci" + fi + # Check for AMD GPU via amdgpu sysfs + if lspci -vnn 2>/dev/null | grep -qi 'amd\|ati'; then + local amdgpu_info + amdgpu_info=$(ls /sys/kernel/debug/amdgpu/ 2>/dev/null | head -1) + if [[ -n "$amdgpu_info" ]]; then + local vram_total + vram_total=$(cat /sys/kernel/debug/amdgpu/${amdgpu_info}/vram_total 2>/dev/null || echo 0) + if [[ "$vram_total" -gt 0 ]]; then + local vram_gb=$((vram_total / 1024 / 1024 / 1024)) + local vram_mb=$((vram_total / 1024 / 1024)) + echo "AMD GPU VRAM: ${vram_gb}GB (${vram_mb}MB)" + fi + fi + fi + fi + fi + + echo "Total VRAM: $((total_vram_kb / 1024 / 1024))GB" + echo "VRAM per GPU: $((vram_per_gpu_kb / 1024 / 1024))GB" + echo "Num GPUs: $num_gpus" +} + +# ─── System RAM Detection ─── +detect_ram() { + local total_kb=0 + local available_kb=0 + + if [[ -f /proc/meminfo ]]; then + total_kb=$(grep MemTotal /proc/meminfo | awk '{print $2}') + available_kb=$(grep MemAvailable /proc/meminfo | awk '{print $2}') + if [[ $total_kb -gt 0 ]]; then + echo "RAM: $((total_kb / 1024 / 1024))GB total, $((available_kb / 1024 / 1024))GB available" + echo "$available_kb $total_kb" + fi + elif command -v sysctl &>/dev/null; then + total_kb=$(sysctl -n hw.memsize 2>/dev/null | awk '{print $1 / 1024}') + if [[ -n "$total_kb" && "$total_kb" -gt 0 ]]; then + echo "RAM: $((total_kb / 1024))GB total" + echo "$total_kb $total_kb" # Assume all available + fi + else + echo "RAM: Could not detect" + echo "0 0" + fi +} + +# ─── Model Context Window Detection ─── +detect_model_context() { + local model_name="$1" + local context_kb=0 + + # If model name provided, look it up + if [[ -n "$model_name" ]]; then + case "$model_name" in + gpt-4o|gpt-4o-2024-05-13|gpt-4o-2024-08-06) + context_kb=128000; echo "Model: $model_name" + echo "Context window: 128k tokens" + ;; + gpt-4o-mini|gpt-4o-mini-2024-07-18) + context_kb=128000; echo "Model: $model_name" + echo "Context window: 128k tokens" + ;; + gpt-4-turbo|gpt-4-turbo-2024-04-09) + context_kb=128000; echo "Model: $model_name" + echo "Context window: 128k tokens" + ;; + gpt-4|gpt-4-0125-preview|gpt-4-1106-preview) + context_kb=128000; echo "Model: $model_name" + echo "Context window: 128k tokens" + ;; + claude-3-5-sonnet|claude-3-5-sonnet-20241022) + context_kb=200000; echo "Model: $model_name" + echo "Context window: 200k tokens" + ;; + claude-3-5-haiku|claude-3-5-haiku-20241022) + context_kb=200000; echo "Model: $model_name" + echo "Context window: 200k tokens" + ;; + claude-3-opus|claude-3-opus-20240229) + context_kb=200000; echo "Model: $model_name" + echo "Context window: 200k tokens" + ;; + claude-3-sonnet|claude-3-sonnet-20240229) + context_kb=200000; echo "Model: $model_name" + echo "Context window: 200k tokens" + ;; + claude-3-haiku|claude-3-haiku-20240307) + context_kb=200000; echo "Model: $model_name" + echo "Context window: 200k tokens" + ;; + claude-2|claude-2.1) + context_kb=200000; echo "Model: $model_name" + echo "Context window: 200k tokens" + ;; + *) + echo "Model: $model_name (unknown context window)" + echo "0" + ;; + esac + echo "$context_kb" + return + fi + + # Try to detect from config.md (global framework model settings) + local project_dir="${1:-.}" + local config_md="${HOME}/.agent-framework/config.md" + local model_from_config="" + local override_context="" + if [[ -f "$config_md" ]]; then + # Check for model name + model_from_config=$(grep -i "model:" "$config_md" 2>/dev/null | grep -v "#" | grep -v "model_context" | grep -v "override" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]') + # Check for override context window + override_context=$(grep -i "override context" "$config_md" 2>/dev/null | grep -v "#" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]') + fi + + # If config.md specifies a model, use it + if [[ -n "$model_from_config" ]]; then + echo "Found model in config.md: $model_from_config" + # Use override context window if specified in config.md + if [[ -n "$override_context" && "$override_context" != "auto" ]]; then + echo "Using override context window from config.md: $override_context" + local override_kb + override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true) + if [[ -n "$override_kb" ]]; then + echo "Context window: ${override_context} tokens (override)" + echo "$((override_kb * 1000))" + return + fi + fi + detect_model_context "$model_from_config" + return + fi + + # Try to detect from AGENT.md (project-level model override) + local agent_md="${project_dir}/.agent-framework/AGENT.md" + if [[ -f "$agent_md" ]]; then + local model_line + model_line=$(grep -i "model" "$agent_md" 2>/dev/null | grep -v "#" | grep -v "target" | grep -v "headroom" | grep -v "peak" | grep -v "Auto-detect" | head -1) + if [[ -n "$model_line" ]]; then + echo "Found model in AGENT.md: $model_line" + # Extract model name from the line + local model + model=$(echo "$model_line" | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]') + if [[ -n "$model" ]]; then + # Use override context window if specified in config.md + if [[ -n "$override_context" && "$override_context" != "auto" ]]; then + echo "Using override context window from config.md: $override_context" + # Convert override_context to kb (e.g., 128k -> 128000, 200k -> 200000) + local override_kb + override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true) + if [[ -n "$override_kb" ]]; then + echo "Context window: ${override_context} tokens (override)" + echo "$((override_kb * 1000))" + return + fi + fi + detect_model_context "$model" + return + fi + fi + fi + + # Try to detect from common API config files + local config_files=( + ".env" + ".env.local" + "config.yaml" + "config.yml" + "config.json" + "settings.yaml" + ".agent-framework/config.yaml" + ".agent-framework/config.json" + ) + + for config_file in "${config_files[@]}"; do + local abs_file="" + for candidate in "${project_dir}/${config_file}" "${project_dir}/.agent-framework/${config_file}"; do + if [[ -f "$candidate" ]]; then + abs_file="$candidate" + break + fi + done + + if [[ -n "$abs_file" ]]; then + local model + model=$(grep -i "model" "$abs_file" 2>/dev/null | grep -v "#" | grep -v "context" | grep -v "max_tokens" | grep -v "temperature" | grep -v "stream" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]') + if [[ -n "$model" ]]; then + echo "Found model in $abs_file: $model" + # Use override context window if specified in config.md + if [[ -n "$override_context" && "$override_context" != "auto" ]]; then + echo "Using override context window from config.md: $override_context" + local override_kb + override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true) + if [[ -n "$override_kb" ]]; then + echo "Context window: ${override_context} tokens (override)" + echo "$((override_kb * 1000))" + return + fi + fi + detect_model_context "$model" + return + fi + fi + done + + echo "Model: Unknown (could not detect from AGENT.md or config files)" + echo "0" +} + +# ─── Agent Framework Overhead Calculation ─── +calculate_overhead() { + local project_dir="${1:-.}" + local overhead_tokens=0 + + # Count tokens for the framework files that are loaded during orchestration + # These are the files loaded during the most common phase (orchestration): + # AGENT.md + RULES.md + workflow.md + orchestrate.md + # Note: other phase files (decompose.md, implement.md, etc.) are only loaded during + # their specific phases, so they don't contribute to the peak context during orchestration. + local framework_files=( + "${project_dir}/.agent-framework/AGENT.md" + "${project_dir}/.agent-framework/RULES.md" + "${project_dir}/.agent-framework/prompts/workflow.md" + "${project_dir}/.agent-framework/prompts/orchestrate.md" + ) + # Fallback: check home directory if project dir doesn't have framework + if [[ ! -f "${project_dir}/.agent-framework/AGENT.md" ]]; then + framework_files=( + "${HOME}/.agent-framework/AGENT.md" + "${HOME}/.agent-framework/RULES.md" + "${HOME}/.agent-framework/prompts/workflow.md" + "${HOME}/.agent-framework/prompts/orchestrate.md" + ) + fi + + for file in "${framework_files[@]}"; do + if [[ -f "$file" ]]; then + # Rough estimate: 1 token ≈ 4 characters (English text) + local chars + chars=$(wc -c < "$file" 2>/dev/null || echo 0) + local tokens=$((chars / 4)) + overhead_tokens=$((overhead_tokens + tokens)) + echo " ${file##*/}: ~${tokens} tokens" + fi + done + + echo "Framework overhead: ~${overhead_tokens} tokens" + echo "$overhead_tokens" +} + +# ─── Recommendation Engine ─── +recommend_context() { + local gpu_vram_gb="$1" + local ram_gb="$2" + local model_context_kb="$3" + local overhead_tokens="$4" + + # Read VRAM config from config.md if it exists + local config_md="${HOME}/.agent-framework/config.md" + local auto_detect="Yes" + local target_context_kb=0 + local override_headroom=25 + local override_max_peak_kb=0 + if [[ -f "$config_md" ]]; then + auto_detect=$(grep -i "auto-detect:" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]' || true) + target_context_kb=$(grep -i "target.*context" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true) + override_headroom=$(grep -i "headroom" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true) + override_max_peak_kb=$(grep -i "max peak" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true) + fi + + # If auto-detect is disabled, use the manually specified values + if [[ -n "$auto_detect" && "$auto_detect" == "No" ]]; then + if [[ -n "$target_context_kb" ]]; then + local max_peak_kb=${override_max_peak_kb:-0} + if [[ $max_peak_kb -eq 0 && $headroom_pct -gt 0 ]]; then + max_peak_kb=$((target_context_kb * (100 - headroom_pct) / 100)) + fi + echo "$headroom_pct" + echo "$target_context_kb" + echo "$max_peak_kb" + return + fi + fi + + local recommended_kb=0 + local headroom_pct=${override_headroom:-25} # Use override from config.md, or default to 25% + + # Output headroom_pct first (for parent to read) + # Then output recommended_kb + # Then output max_peak_kb (only for manual mode) + echo "$headroom_pct" + + # Recommendation logic: + # 1. If GPU VRAM >= 4GB: use VRAM (practical for local inference) + # 2. If model context window is available: use it (for API inference) + # 3. If GPU VRAM < 4GB but > 0: use RAM (VRAM too small for local inference) + # 4. If no GPU VRAM and no model: use RAM as fallback + + # If GPU VRAM >= 4GB, base it on VRAM + if [[ $gpu_vram_gb -ge 4 ]]; then + # Rule of thumb: 1GB VRAM ≈ 4k tokens for local LLMs + # But we need to leave room for the model itself + # For a model, each ~8k context tokens takes about ~3-5MB of GPU VRAM + # So VRAM available for context = VRAM - model size - agent overhead + # Conservative: 1GB VRAM ≈ 2k context tokens + local vram_context_kb=$((gpu_vram_gb * 2000)) + + # Leave headroom for the model itself and agent overhead + recommended_kb=$((vram_context_kb * (100 - headroom_pct) / 100)) + # If model context window is available, use it (for API inference) + elif [[ $model_context_kb -gt 0 ]]; then + # For API-based, we're limited by the model's context window + # But we don't want to use the full window due to overhead + recommended_kb=$((model_context_kb * (100 - headroom_pct) / 100)) + # Fallback: use RAM to estimate + else + # Moderate estimate for low-VRAM systems where VRAM is too small for local inference + # but RAM is available. Use 0.75k tokens per GB of RAM as a moderate estimate. + # This balances between being too conservative (0.5k/GB) and too generous (1k/GB). + local ram_context_kb=$((ram_gb * 750)) + recommended_kb=$((ram_context_kb * (100 - headroom_pct) / 100)) + fi + + # Subtract framework overhead + local net_kb=$((recommended_kb - overhead_tokens)) + if [[ $net_kb -lt 0 ]]; then + net_kb=0 + fi + + # Output: headroom_pct, recommended_kb + echo "$net_kb" +} + +# ─── Main ─── +main() { + local model_name="" + local project_dir="." + + # Parse arguments + while [[ $# -gt 0 ]]; do + case "$1" in + --model|-m) + model_name="$2" + shift 2 + ;; + --project|-p) + project_dir="$2" + shift 2 + ;; + *) + # Could be model name as first argument + if [[ -z "$model_name" ]]; then + model_name="$1" + fi + shift + ;; + esac + done + + echo "=== VRAM / Context Detection ===" + echo "" + + # Detect GPU VRAM + echo "--- GPU VRAM ---" + detect_gpu_vram + local gpu_vram_kb=0 + local gpu_vram_gb=0 + # Use timeout to avoid hanging + gpu_vram_kb=$(timeout 5 nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ' || true) + # Validate that vram_kb is a positive number + if [[ -z "$gpu_vram_kb" || ! "$gpu_vram_kb" =~ ^[0-9]+$ || "$gpu_vram_kb" -le 0 ]]; then + gpu_vram_kb=0 + fi + # If nvidia-smi didn't work, try to detect from lspci (AMD GPUs) + if [[ $gpu_vram_kb -eq 0 ]]; then + echo " nvidia-smi failed, checking lspci for AMD GPU VRAM..." + local total_vram_mb=0 + while IFS= read -r line; do + local size_num + size_num=$(echo "$line" | grep -oE '[0-9]+(M|G|K)' | head -1 || true) + if [[ -n "$size_num" ]]; then + local size_val + size_val=$(echo "$size_num" | grep -oE '[0-9]+') + local size_unit + size_unit=$(echo "$size_num" | grep -oE '(M|G|K)') + if [[ -n "$size_val" && -n "$size_unit" ]]; then + case "$size_unit" in + M) total_vram_mb=$((total_vram_mb + size_val)) ;; + G) total_vram_mb=$((total_vram_mb + size_val * 1024)) ;; + K) total_vram_mb=$((total_vram_mb + size_val / 1024)) ;; + esac + fi + fi + done < <(lspci -vnn 2>/dev/null | grep -i -A 15 -E 'VGA|3D|Display' | grep -i 'Memory at') + if [[ $total_vram_mb -gt 0 ]]; then + gpu_vram_kb=$((total_vram_mb * 1024)) + gpu_vram_gb=$((total_vram_mb / 1024)) + echo " AMD GPU VRAM from lspci: ${gpu_vram_gb}GB ($total_vram_mb MB)" + else + echo " No VRAM found from lspci" + fi + fi + echo "" + + # Detect RAM + echo "--- RAM ---" + detect_ram + local ram_kb + ram_kb=$(grep MemTotal /proc/meminfo 2>/dev/null | awk '{print $2}' || echo 0) + local ram_gb=$((ram_kb / 1024 / 1024)) + echo "" + + # Detect model context window + echo "--- Model Context Window ---" + detect_model_context "$model_name" + local model_context_kb + model_context_kb=$(detect_model_context "$model_name" | tail -1) + echo "" + + # Calculate framework overhead + echo "--- Framework Overhead ---" + calculate_overhead "$project_dir" + local overhead_tokens + overhead_tokens=$(calculate_overhead "$project_dir" | tail -1) + echo "" + + # Recommend context window (also outputs headroom_pct and recommended_kb) + echo "--- Recommendation ---" + local recommendation_output + recommendation_output=$(recommend_context "$gpu_vram_gb" "$ram_gb" "$model_context_kb" "$overhead_tokens") + local line_count + line_count=$(echo "$recommendation_output" | wc -l) + local recommended_kb + local headroom_pct + local max_peak_kb + if [[ $line_count -ge 3 ]]; then + # Manual mode: outputs headroom_pct, recommended_kb, max_peak_kb + headroom_pct=$(echo "$recommendation_output" | head -1 | tr -d '[:space:]') + recommended_kb=$(echo "$recommendation_output" | sed -n '2p' | tr -d '[:space:]') + max_peak_kb=$(echo "$recommendation_output" | tail -1 | tr -d '[:space:]') + else + # Auto-detect mode: outputs headroom_pct, recommended_kb + headroom_pct=$(echo "$recommendation_output" | head -1 | tr -d '[:space:]') + recommended_kb=$(echo "$recommendation_output" | tail -1 | tr -d '[:space:]') + # Calculate max peak context based on headroom + max_peak_kb=$((recommended_kb * (100 - headroom_pct) / 100)) + fi + + # Convert to human-readable + local recommended_k + if [[ $recommended_kb -gt 0 ]]; then + recommended_k=$((recommended_kb / 1000)) + else + recommended_k=8 # Default fallback + fi + + echo "" + echo "=== Recommended Configuration ===" + echo "Target context: ${recommended_k}k tokens" + echo "Headroom: ${headroom_pct}%" + echo "Max peak context per sub-task: $((recommended_k * (100 - headroom_pct) / 100))k tokens" + + # Output as JSON for programmatic use + echo "" + echo "=== JSON Output ===" + cat <