Split config: VRAM/model settings to config.md, agent behavior to AGENT.md
This commit is contained in:
@@ -3,6 +3,8 @@
|
||||
## Autopilot
|
||||
Autopilot: Enabled
|
||||
|
||||
> Note: For global framework settings (VRAM, model, system requirements), see `~/.agent-framework/config.md`.
|
||||
|
||||
## Routing
|
||||
|
||||
IF task type = research → load prompts/research.md + RULES.md
|
||||
@@ -13,7 +15,8 @@ IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code
|
||||
IF task type = adversarial_bug_find → load prompts/adversarial_bug_find.md + SPEC.md + code
|
||||
IF task type = referee → load prompts/referee.md + SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md
|
||||
IF task type = doc_review → load prompts/doc_review.md + DESIGN.md
|
||||
IF task type = decompose → load prompts/decompose.md + SPEC.md
|
||||
IF task type = orchestrate → load prompts/orchestrate.md + project structure
|
||||
IF task type = compaction → load prompts/compaction.md
|
||||
|
||||
Always start by reading this file to determine mode.
|
||||
Always start by reading this file to determine mode.
|
||||
|
||||
@@ -66,7 +66,8 @@ The framework features an **Autopilot** mode that allows the agent to drive a pr
|
||||
|
||||
### Lifecycle of a Task
|
||||
1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. **Interactive** — agent grills you for requirements and gets sign-off.
|
||||
2. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off.
|
||||
2. **Decomposition** (optional): Break the task into sub-tasks sized for your VRAM. **Interactive** — agent analyzes the spec and proposes sub-tasks with token budget estimates, gets sign-off. See [VRAM Configuration](#vram-configuration) below.
|
||||
3. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off.
|
||||
3. **Test Design** (optional): Produce a `TEST_PLAN.md` (Test specification). No code allowed. **Interactive** — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off.
|
||||
4. **Implement**: Write code and tests based *only* on the `SPEC.md`, `DESIGN.md` (if present), and `TEST_PLAN.md` (if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows.
|
||||
5. **Bug Find**: Aggressive search for bugs and spec deviations.
|
||||
@@ -79,11 +80,69 @@ The framework features an **Autopilot** mode that allows the agent to drive a pr
|
||||
#### Autopilot mode (default)
|
||||
The default mode is **Autopilot: Enabled**. The Orchestrator automatically drives tasks through all phases. Just say:
|
||||
- **Start a new task**: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion)
|
||||
- **Decompose a task**: "Decompose the add-user-auth task" (the Orchestrator breaks it into sub-tasks sized for your VRAM)
|
||||
- **Continue**: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it)
|
||||
|
||||
#### VRAM Configuration
|
||||
For low-VRAM systems (8GB, 16GB), set your VRAM limits in `~/.agent-framework/config.md`:
|
||||
|
||||
```markdown
|
||||
## VRAM Configuration
|
||||
- **Auto-detect**: Yes # Let the agent detect your VRAM automatically
|
||||
- **Target VRAM context**: 16k tokens # Override auto-detect if needed
|
||||
- **Headroom**: 25%
|
||||
- **Max peak context per sub-task**: 12k tokens
|
||||
```
|
||||
|
||||
**Auto-detection**: When `Auto-detect: Yes`, the Orchestrator runs `scripts/vram_detect.sh` to probe:
|
||||
- GPU VRAM (via `nvidia-smi`)
|
||||
- System RAM (via `free`)
|
||||
- Model context window (from config.md or API config files)
|
||||
- Framework overhead (by counting token load in loaded prompts)
|
||||
|
||||
**Manual override**: When `Auto-detect: No`, use the manually specified values:
|
||||
|
||||
```markdown
|
||||
## VRAM Configuration
|
||||
- **Auto-detect**: No
|
||||
- **Target VRAM context**: 8k
|
||||
- **Headroom**: 30%
|
||||
- **Max peak context per sub-task**: 5.6k
|
||||
```
|
||||
|
||||
#### Model Configuration
|
||||
When using a local LLM or a specific API model, set the model in `~/.agent-framework/config.md`:
|
||||
|
||||
```markdown
|
||||
## Model Configuration
|
||||
- **Model**: auto # Use auto-detection from API config files
|
||||
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
|
||||
```
|
||||
|
||||
**Auto-detection**: When `Model: auto`, the framework detects the model name from API config files (`.env`, `config.yaml`, etc.) and looks up its context window.
|
||||
|
||||
**Manual override**: When you know your model name, specify it:
|
||||
|
||||
```markdown
|
||||
## Model Configuration
|
||||
- **Model**: gpt-4o
|
||||
- **Override context window**: 128k
|
||||
```
|
||||
|
||||
When you run "Decompose the X task", the Orchestrator will:
|
||||
1. Analyze the task's SPEC.md
|
||||
2. Detect VRAM limits (auto or manual)
|
||||
3. Break it into sub-tasks, each sized to fit within your VRAM limit
|
||||
4. Estimate the token budget for each sub-task
|
||||
5. Create sub-task folders under `tasks/{parent-task}/subtasks/{sub-task}/`
|
||||
6. Propagate VRAM config to each sub-task
|
||||
|
||||
Sub-tasks run independently through the full lifecycle. The parent task is complete only when ALL sub-tasks pass.
|
||||
|
||||
#### Manual mode (opt-in)
|
||||
Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like:
|
||||
- "Research add user authentication" — starts a new task
|
||||
- "Decompose the add-user-auth task" — breaks into VRAM-sized sub-tasks (optional)
|
||||
- "Design the add-user-auth task" — designs the architecture (optional)
|
||||
- "Design tests for the add-user-auth task" — designs test cases (optional)
|
||||
- "Implement the add-user-auth task" — implements the task with tests
|
||||
@@ -93,12 +152,27 @@ Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you p
|
||||
- "Review the add-user-auth task" — referee evaluates
|
||||
- "orchestrate" — asks the Orchestrator what to do next
|
||||
|
||||
## Sub-Task Management
|
||||
|
||||
When a task is decomposed, the Orchestrator creates sub-tasks under `tasks/{parent-task}/subtasks/`. Each sub-task:
|
||||
|
||||
- Has its own full lifecycle (Research → Decomposition → Design → ... → Complete)
|
||||
- Starts at the Research phase with an empty folder
|
||||
- Receives a `PARENT_SPEC.md` with the parent task's context
|
||||
- Receives a `VRAM_CONFIG.md` with the VRAM constraints
|
||||
- Runs independently — sub-tasks in the same wave can run in parallel
|
||||
- The parent task is NOT complete until ALL sub-tasks pass
|
||||
|
||||
## Key Components
|
||||
- `AGENT.md`: Project-specific configuration and mode selection.
|
||||
- `AGENT.md`: Project-specific agent behavior (Autopilot mode, routing rules).
|
||||
- `config.md`: Global framework settings (VRAM, model, system requirements).
|
||||
- `RULES.md`: Living document of project constraints and past failure modes.
|
||||
- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, etc.).
|
||||
- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, Decompose, etc.).
|
||||
- `workflow.md`: The state machine governing the Autopilot lifecycle.
|
||||
- `test_design.md`: Produces a TEST_PLAN.md — an explicit test specification before implementation.
|
||||
- `decompose.md`: Breaks a task into VRAM-sized sub-tasks.
|
||||
- `scripts/vram_detect.sh`: Auto-detects GPU VRAM, RAM, model context window, and framework overhead.
|
||||
- `contracts/vram_config.md`: Contract for VRAM-aware task decomposition.
|
||||
|
||||
## Contact & Support
|
||||
[Insert Contact Info]
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
# Framework Configuration
|
||||
|
||||
This file contains global framework settings that apply across all projects.
|
||||
|
||||
## VRAM Configuration
|
||||
|
||||
Settings for task decomposition based on available VRAM.
|
||||
|
||||
- **Auto-detect**: Yes # Detect GPU VRAM, RAM, and model context window automatically
|
||||
- **Target context**: 16k tokens # Override auto-detect if needed
|
||||
- **Headroom**: 25% # Leave headroom for code, context, and reasoning
|
||||
- **Max peak context per sub-task**: 12k tokens # Max context for any single sub-task
|
||||
|
||||
### Auto-detection
|
||||
|
||||
When `Auto-detect: Yes`, the framework probes your system to detect:
|
||||
- GPU VRAM (via `nvidia-smi` or `lspci`)
|
||||
- System RAM (via `free`)
|
||||
- Model context window (via API config or model name lookup)
|
||||
- Framework overhead (by reading all loaded prompt files)
|
||||
|
||||
To disable auto-detection and use manual values:
|
||||
|
||||
```
|
||||
## VRAM Configuration
|
||||
- **Auto-detect**: No
|
||||
- **Target context**: 8k
|
||||
- **Headroom**: 30%
|
||||
- **Max peak context per sub-task**: 5.6k
|
||||
```
|
||||
|
||||
## Model Configuration
|
||||
|
||||
Settings for the LLM model being used.
|
||||
|
||||
- **Model**: auto # Use auto-detection from API config, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
|
||||
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
|
||||
|
||||
### Auto-detection
|
||||
|
||||
When `Model: auto`, the framework detects the model name from:
|
||||
1. `AGENT.md` in the project (if specified there)
|
||||
2. API config files (`.env`, `config.yaml`, `config.json`, etc.)
|
||||
3. Model name lookup by name (e.g., gpt-4o → 128k, claude-3-5-sonnet → 200k)
|
||||
|
||||
To disable auto-detection and use manual values:
|
||||
|
||||
```
|
||||
## Model Configuration
|
||||
- **Model**: gpt-4o
|
||||
- **Override context window**: 128k
|
||||
```
|
||||
|
||||
## System Requirements
|
||||
|
||||
Requirements for the environment the framework runs in.
|
||||
|
||||
- **nvidia-smi**: Required if NVIDIA GPU (for VRAM detection)
|
||||
- **lspci**: Fallback if NVIDIA GPU not available (for AMD GPU VRAM detection)
|
||||
- **/proc/meminfo**: Required for RAM detection (Linux)
|
||||
- **sysctl**: Fallback for RAM detection (macOS)
|
||||
@@ -0,0 +1,11 @@
|
||||
## VRAM Configuration Contract
|
||||
|
||||
### Acceptance Criteria
|
||||
- [ ] Target VRAM context is specified
|
||||
- [ ] Headroom is specified (recommended: 25-40%)
|
||||
- [ ] Max peak context per sub-task is calculated
|
||||
- [ ] Sub-tasks are sized to fit within the max peak context
|
||||
- [ ] VRAM_CONFIG.md is propagated to each sub-task
|
||||
|
||||
### Stop Condition
|
||||
When all checkboxes are checked, output "CONTRACT_MET" and stop.
|
||||
Regular → Executable
+46
@@ -12,5 +12,51 @@ fi
|
||||
echo "Cloning agent-framework to $FRAMEWORK_DIR..."
|
||||
git clone https://gitea.yourdomain.com/you/agent-framework.git "$FRAMEWORK_DIR"
|
||||
|
||||
echo ""
|
||||
echo "=== VRAM / Context Detection ==="
|
||||
echo "Detecting your system's VRAM to recommend task decomposition settings..."
|
||||
echo ""
|
||||
|
||||
# Run VRAM detection script if it exists
|
||||
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.sh" ]; then
|
||||
# Run in project-dir context so it can read framework overhead
|
||||
detection_output=$(cd "$FRAMEWORK_DIR" && bash "$FRAMEWORK_DIR/scripts/vram_detect.sh" 2>&1)
|
||||
|
||||
# Extract JSON output (last section after "=== JSON Output ===")
|
||||
json_output=$(echo "$detection_output" | sed -n '/=== JSON Output ===/,/EOF/p' | grep -v '=== JSON Output ===' | grep -v '^EOF$')
|
||||
|
||||
if [ -n "$json_output" ]; then
|
||||
echo "$detection_output"
|
||||
|
||||
# Extract key values from JSON for display
|
||||
recommended_k=$(echo "$json_output" | grep '"recommended_k"' | grep -oP '\d+')
|
||||
max_peak_kb=$(echo "$json_output" | grep '"max_peak_context_kb"' | grep -oP '\d+')
|
||||
headroom=$(echo "$json_output" | grep '"headroom"' | grep -oP '\d+\.\d+')
|
||||
gpu_vram=$(echo "$json_output" | grep '"gpu_vram_gb"' | grep -oP '\d+')
|
||||
ram_gb=$(echo "$json_output" | grep '"ram_gb"' | grep -oP '\d+')
|
||||
model_context=$(echo "$json_output" | grep '"model_context_kb"' | grep -oP '\d+')
|
||||
|
||||
echo ""
|
||||
echo "=== Recommended VRAM Configuration ==="
|
||||
echo "For low-VRAM systems (8GB, 16GB VRAM), add this to ~/.agent-framework/config.md:"
|
||||
echo ""
|
||||
echo "## VRAM Configuration"
|
||||
echo "- **Auto-detect**: Yes # Let the agent detect automatically"
|
||||
echo "- **Target context**: ${recommended_k}k tokens # Override auto-detect if needed"
|
||||
echo "- **Headroom**: ${headroom}%"
|
||||
echo "- **Max peak context per sub-task**: $((max_peak_kb / 1000))k tokens"
|
||||
echo ""
|
||||
echo "This ensures tasks are decomposed into sub-tasks that fit within your"
|
||||
echo "available VRAM. For more information, see the README."
|
||||
else
|
||||
echo "Could not detect VRAM. You can manually set your VRAM configuration in ~/.agent-framework/config.md."
|
||||
echo "See the README for details."
|
||||
fi
|
||||
else
|
||||
echo "VRAM detection script not found. You can manually set your VRAM configuration in ~/.agent-framework/config.md."
|
||||
echo "See the README for details."
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "Installation complete."
|
||||
echo "Next step: cd into a project and run the onboarding prompt."
|
||||
@@ -1,9 +1,20 @@
|
||||
You are the Adversarial Bug Finder.
|
||||
|
||||
Read the SPEC.md and the code.
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md
|
||||
2. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
|
||||
3. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
|
||||
4. The code
|
||||
|
||||
Your job is to find bugs that are difficult to spot, such as complex logic errors, race conditions, and performance bottlenecks. Be more aggressive and exhaustive than a standard bug finder.
|
||||
|
||||
If VRAM_CONFIG.md exists, also check for:
|
||||
- Memory leaks (loading large files into context that could cause OOM)
|
||||
- N+1 query patterns that could cause memory exhaustion
|
||||
- Infinite loops that could run out of context
|
||||
- Unbounded recursion that could cause stack overflow
|
||||
|
||||
Output your findings in ADVERSARIAL_BUG_REPORT.md.
|
||||
|
||||
When finished, output "ADVERSARIAL_BUG_FIND_COMPLETE".
|
||||
@@ -5,6 +5,8 @@ You are the Bug Finder. Your job is to find every bug, deviation from spec, and
|
||||
1. {project}/tasks/{task-name}/SPEC.md
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
|
||||
3. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
|
||||
4. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
|
||||
5. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
|
||||
@@ -0,0 +1,233 @@
|
||||
You are in decomposition mode.
|
||||
|
||||
Your only job is to take a completed SPEC.md and break it into the smallest possible, independently verifiable sub-tasks. Each sub-task should be small enough to complete in one session and should have clear, testable acceptance criteria.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md
|
||||
2. {project}/.agent-framework/RULES.md
|
||||
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
|
||||
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
|
||||
5. {project}/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Decomposition Rules
|
||||
|
||||
### Rule 1: Smallest Possible Unit
|
||||
Break every feature into the smallest possible units that are still independently testable. If a sub-task can be done in one session, it should be a sub-task.
|
||||
|
||||
### Rule 2: Each Sub-Task Must Be Self-Contained
|
||||
Each sub-task must have:
|
||||
- Its own goal statement (one sentence)
|
||||
- Clear acceptance criteria (at least 2-3)
|
||||
- Dependencies on other sub-tasks (if any)
|
||||
- Its own contract (SPEC.md) that references the parent task
|
||||
|
||||
### Rule 3: Dependencies Must Be Explicit
|
||||
If sub-task B depends on sub-task A, state it clearly. A sub-task with no dependencies can run in parallel. Sub-tasks that depend on parallel sub-tasks must wait for both.
|
||||
|
||||
### Rule 4: Define Execution Order
|
||||
After listing all sub-tasks, define the execution order considering dependencies. Group independent sub-tasks into "waves" that can be done in parallel.
|
||||
|
||||
### Rule 5: Do Not Create Sub-Sub-Tasks
|
||||
Decomposition produces **one level** of sub-tasks only. If a sub-task is still too large, the implementer should break it down further during implementation, but do NOT nest sub-tasks.
|
||||
|
||||
### Rule 6: Token Budget Per Sub-Task
|
||||
Each sub-task must fit within the target VRAM context window. Estimate the total token budget for each sub-task's full lifecycle and break it down further if it exceeds the limit.
|
||||
|
||||
**How to estimate token budget for a sub-task:**
|
||||
|
||||
During a sub-task's lifecycle, the following files are loaded into context at various phases:
|
||||
|
||||
- **Research phase**: RULES.md + AGENT.md + task description
|
||||
- **Design phase**: SPEC.md + RULES.md
|
||||
- **Implement phase**: SPEC.md + DESIGN.md + TEST_PLAN.md + RULES.md + AGENT.md + CONTRACT.md
|
||||
- **Bug Find phase**: SPEC.md + code (limited scope)
|
||||
- **Adversarial Bug Find phase**: SPEC.md + code (limited scope)
|
||||
- **Doc Review phase**: DESIGN.md
|
||||
- **Referee phase**: SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md
|
||||
|
||||
The **peak context** is during the Implement phase, where all files are loaded together. Estimate the token count of the combined files for the sub-task.
|
||||
|
||||
**Guidelines:**
|
||||
- **8k VRAM**: Peak context for Implement phase should be ≤ 6k tokens (leave 2k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 6k tokens.
|
||||
- **16k VRAM**: Peak context for Implement phase should be ≤ 12k tokens (leave 4k headroom). This means the sub-task's SPEC.md + DESIGN.md + TEST_PLAN.md combined should be ≤ 12k tokens.
|
||||
- **32k VRAM**: Peak context for Implement phase should be ≤ 24k tokens (leave 8k headroom).
|
||||
- **64k VRAM**: Peak context for Implement phase should be ≤ 48k tokens (leave 16k headroom).
|
||||
|
||||
If a sub-task's estimated context exceeds the limit, break it into smaller sub-tasks.
|
||||
|
||||
**How to estimate token count:**
|
||||
- Roughly 1 token = 4 characters (for English text)
|
||||
- A SPEC.md with 5 requirements, each with 2-3 acceptance criteria, is typically 500-1000 tokens
|
||||
- A DESIGN.md with 3 sections and 5-10 bullet points is typically 1000-3000 tokens
|
||||
- A TEST_PLAN.md with 5-10 test cases is typically 1000-2000 tokens
|
||||
- RULES.md is typically 200-1000 tokens (varies per project)
|
||||
- AGENT.md is typically 300-1000 tokens
|
||||
- A CONTRACT.md is typically 200-500 tokens
|
||||
|
||||
**Quick estimate formula:**
|
||||
```
|
||||
Peak context ≈ SPEC.md tokens + DESIGN.md tokens + TEST_PLAN.md tokens + RULES.md tokens + AGENT.md tokens + CONTRACT.md tokens
|
||||
```
|
||||
|
||||
### Rule 7: Sub-Task Size Targets
|
||||
Aim for sub-tasks that are:
|
||||
- **Small** (8k VRAM): ~200-800 tokens of combined spec/design/test files
|
||||
- **Small** (16k VRAM): ~200-1500 tokens of combined spec/design/test files
|
||||
- **Medium** (8k VRAM): ~800-2000 tokens of combined spec/design/test files
|
||||
- **Medium** (16k VRAM): ~1500-4000 tokens of combined spec/design/test files
|
||||
- **Large** (8k VRAM): ~2000-4000 tokens of combined spec/design/test files
|
||||
- **Large** (16k VRAM): ~4000-8000 tokens of combined spec/design/test files
|
||||
|
||||
If a sub-task exceeds the "Large" target for the target VRAM, break it further.
|
||||
|
||||
## Decomposition Protocol (Interactive)
|
||||
|
||||
### Phase 1: Analysis
|
||||
|
||||
Before decomposing, analyze the SPEC.md:
|
||||
1. Identify all distinct features/requirements
|
||||
2. Identify data models that need to be created
|
||||
3. Identify API endpoints or interfaces
|
||||
4. Identify infrastructure changes
|
||||
5. Identify configuration changes
|
||||
6. Estimate the token budget for the full task (sum of all requirements' SPEC + DESIGN + TEST files)
|
||||
7. **Detect VRAM limits**:
|
||||
- Check `~/.agent-framework/config.md` for VRAM Configuration section
|
||||
- If `Auto-detect: Yes`, run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window
|
||||
- If `Auto-detect: No`, use the manually specified values from config.md
|
||||
- Report the detected VRAM limits
|
||||
8. **Detect model context window**:
|
||||
- Check `~/.agent-framework/config.md` for Model Configuration section
|
||||
- If `Model: auto`, run the detection script to detect the model name and its context window
|
||||
- If `Override context window: auto`, use the detected context window
|
||||
- If both are specified, use the specified values
|
||||
- If model detection fails, use 128k tokens as default
|
||||
9. Determine the target VRAM context window based on the detection results
|
||||
|
||||
### Phase 2: Propose Decomposition
|
||||
|
||||
Present a draft decomposition to the user. Format:
|
||||
|
||||
**Target VRAM**: {8k/16k/32k/64k} tokens
|
||||
|
||||
**Waves:**
|
||||
- **Wave 1**: Sub-task 1, Sub-task 2, Sub-task 3 (can run in parallel)
|
||||
- **Wave 2**: Sub-task 4, Sub-task 5 (depends on Wave 1)
|
||||
- **Wave 3**: Sub-task 6 (depends on Wave 2)
|
||||
|
||||
**Sub-task Details:**
|
||||
1. **{sub-task-name}**
|
||||
- Goal: {one sentence}
|
||||
- Dependencies: {list of sub-task names, or "None"}
|
||||
- Acceptance criteria:
|
||||
- [ ] {criterion 1}
|
||||
- [ ] {criterion 2}
|
||||
- Estimated scope: {small/medium/large}
|
||||
- **Estimated token budget**: ~{estimate} tokens (peak: ~{peak} tokens during Implement phase)
|
||||
- **Fits within VRAM**: Yes/No (if No, explain why and suggest how to split further)
|
||||
|
||||
### Phase 3: Review and Refine
|
||||
|
||||
Present the draft decomposition to the user and ask:
|
||||
- "Are there any sub-tasks that are too large?"
|
||||
- "Are there any sub-tasks that should be combined?"
|
||||
- "Are the dependencies correct?"
|
||||
- "Are there any sub-tasks I missed?"
|
||||
- "Is the execution order optimal?"
|
||||
- "Do the token budget estimates look reasonable for your VRAM?"
|
||||
- "Are there any sub-tasks that exceed your VRAM limit?"
|
||||
|
||||
Incorporate the user's feedback and revise the decomposition accordingly.
|
||||
|
||||
### Phase 4: Get Sign-Off
|
||||
|
||||
Before writing the DECOMPOSITION.md, you MUST get explicit sign-off from the user. Say:
|
||||
|
||||
> "Based on our discussion, here is the final decomposition:
|
||||
> [brief summary of waves, sub-tasks, and token budgets]
|
||||
> Does this cover everything? Please confirm with 'APPROVED' before I finalize."
|
||||
|
||||
Only produce the DECOMPOSITION.md after the user says "APPROVED" or equivalent.
|
||||
|
||||
## Output
|
||||
|
||||
Produce a file called DECOMPOSITION.md at {project}/tasks/{task-name}/DECOMPOSITION.md that contains:
|
||||
|
||||
```markdown
|
||||
# Task Decomposition
|
||||
|
||||
## Parent Task
|
||||
{parent-task-name}
|
||||
|
||||
## VRAM Configuration
|
||||
- **Target VRAM**: {8k/16k/32k/64k} tokens
|
||||
- **Headroom**: {25-40%} (leaves headroom for code, context, and reasoning)
|
||||
- **Max peak context per sub-task**: {estimate} tokens
|
||||
|
||||
## Waves
|
||||
|
||||
### Wave 1: {wave-name}
|
||||
- {sub-task-name-1}
|
||||
- {sub-task-name-2}
|
||||
- {sub-task-name-3}
|
||||
|
||||
### Wave 2: {wave-name}
|
||||
- {sub-task-name-4}
|
||||
- {sub-task-name-5}
|
||||
|
||||
## Sub-Task Details
|
||||
|
||||
### 1. {sub-task-name-1}
|
||||
- **Goal**: {one sentence}
|
||||
- **Dependencies**: None (or list sub-task names)
|
||||
- **Acceptance Criteria**:
|
||||
- [ ] {criterion 1}
|
||||
- [ ] {criterion 2}
|
||||
- **Estimated Scope**: {small/medium/large}
|
||||
- **Estimated token budget**: ~{estimate} tokens
|
||||
- SPEC.md: ~{x} tokens
|
||||
- DESIGN.md: ~{x} tokens
|
||||
- TEST_PLAN.md: ~{x} tokens
|
||||
- RULES.md: ~{x} tokens
|
||||
- AGENT.md: ~{x} tokens
|
||||
- CONTRACT.md: ~{x} tokens
|
||||
- **Peak context (Implement phase)**: ~{peak} tokens
|
||||
- **Fits within VRAM**: Yes
|
||||
|
||||
### 2. {sub-task-name-2}
|
||||
- **Goal**: {one sentence}
|
||||
- **Dependencies**: {list of sub-task names}
|
||||
- **Acceptance Criteria**:
|
||||
- [ ] {criterion 1}
|
||||
- [ ] {criterion 2}
|
||||
- **Estimated Scope**: {small/medium/large}
|
||||
- **Estimated token budget**: ~{estimate} tokens
|
||||
- SPEC.md: ~{x} tokens
|
||||
- DESIGN.md: ~{x} tokens
|
||||
- TEST_PLAN.md: ~{x} tokens
|
||||
- RULES.md: ~{x} tokens
|
||||
- AGENT.md: ~{x} tokens
|
||||
- CONTRACT.md: ~{x} tokens
|
||||
- **Peak context (Implement phase)**: ~{peak} tokens
|
||||
- **Fits within VRAM**: Yes
|
||||
|
||||
... etc ...
|
||||
|
||||
## Execution Order
|
||||
1. Complete Wave 1 (all sub-tasks can run in parallel)
|
||||
2. Complete Wave 2 (depends on Wave 1)
|
||||
3. ... etc ...
|
||||
```
|
||||
|
||||
When the decomposition is complete, output "CONTRACT_MET" and stop.
|
||||
|
||||
Do not create sub-task folders or files. The Orchestrator will handle creating sub-task folders based on this DECOMPOSITION.md.
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced the DECOMPOSITION.md file AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
@@ -4,8 +4,10 @@ You are in Documentation Review mode.
|
||||
|
||||
1. {project}/tasks/{task-name}/DESIGN.md — Look for the "Documentation Plan" section
|
||||
2. {project}/tasks/{task-name}/SPEC.md — Check what the spec requires
|
||||
3. Any existing documentation files mentioned in the DESIGN.md Documentation Plan
|
||||
4. The code that was implemented (implementation artifacts)
|
||||
3. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
|
||||
4. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
|
||||
5. Any existing documentation files mentioned in the DESIGN.md Documentation Plan
|
||||
6. The code that was implemented (implementation artifacts)
|
||||
|
||||
## Task
|
||||
|
||||
@@ -36,6 +38,7 @@ You are in Documentation Review mode.
|
||||
### Documentation Gaps
|
||||
- Are there any areas where the documentation is thin or missing?
|
||||
- Are there any complex flows or non-obvious logic that should be documented?
|
||||
- If VRAM_CONFIG.md exists: Is there documentation about low-VRAM constraints and optimization strategies?
|
||||
|
||||
## Output
|
||||
|
||||
|
||||
@@ -8,6 +8,8 @@ You are in implementation mode.
|
||||
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
|
||||
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
|
||||
6. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
|
||||
7. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists — for low-VRAM systems)
|
||||
8. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists — for sub-tasks)
|
||||
|
||||
## Task
|
||||
|
||||
@@ -26,6 +28,18 @@ You are in implementation mode.
|
||||
- Keep functions small and focused.
|
||||
- Use existing patterns in the codebase.
|
||||
|
||||
### VRAM-Aware Implementation (if VRAM_CONFIG.md exists)
|
||||
|
||||
If `VRAM_CONFIG.md` is present, the implementer is working in a low-VRAM environment and must:
|
||||
|
||||
- **Keep functions small**: Each function should be ≤ 50 lines. Break large functions into smaller ones.
|
||||
- **Avoid loading large files into context**: If implementing in an iterative environment (like pi), read files incrementally rather than loading entire files at once.
|
||||
- **Write self-contained modules**: Each module should be independently testable and runnable without loading the entire codebase.
|
||||
- **Prefer streaming over buffering**: Use streaming patterns instead of loading entire datasets into memory.
|
||||
- **Be explicit about dependencies**: Import only what you need, not the entire module.
|
||||
- **Test incrementally**: Run tests frequently rather than building up a large test suite that takes a long time to run.
|
||||
- **Optimize for low-VRAM**: Write code that can be understood and modified by an agent with limited context window.
|
||||
|
||||
### End State
|
||||
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
|
||||
- All tests must pass.
|
||||
|
||||
@@ -8,6 +8,7 @@ Your only job is to set up the minimal agent framework structure in the target p
|
||||
2. ~/.agent-framework/ONBOARDING.md — human reference for drop-in vs from-scratch scenarios
|
||||
3. {project}/.agent-framework/AGENT.md (if it exists)
|
||||
4. {project}/.agent-framework/RULES.md (if it exists)
|
||||
5. ~/.agent-framework/scripts/vram_detect.sh (if exists — for VRAM detection)
|
||||
|
||||
## Task
|
||||
|
||||
@@ -36,6 +37,32 @@ Before proceeding, check if the project is running an older version of the frame
|
||||
5. Explore the project root at a high level (ls, key directories, README if present).
|
||||
6. Produce a short onboarding report.
|
||||
|
||||
### Step 2: VRAM Configuration
|
||||
|
||||
Check if VRAM configuration is available in `~/.agent-framework/config.md`:
|
||||
1. Read `~/.agent-framework/config.md` to check for VRAM Configuration section.
|
||||
2. If VRAM Configuration section exists, note the values.
|
||||
3. If VRAM Configuration section does NOT exist, check if VRAM detection is available: `~/.agent-framework/scripts/vram_detect.sh`.
|
||||
4. If available, run it to get VRAM recommendations:
|
||||
```
|
||||
cd ~/.agent-framework && bash ~/.agent-framework/scripts/vram_detect.sh
|
||||
```
|
||||
5. Parse the JSON output for `recommended_k`, `max_peak_context_kb`, and `headroom`.
|
||||
6. Add a VRAM Configuration section to `~/.agent-framework/config.md`:
|
||||
```
|
||||
## VRAM Configuration
|
||||
- **Auto-detect**: Yes
|
||||
- **Target context**: {recommended_k}k tokens
|
||||
- **Headroom**: 25%
|
||||
- **Max peak context per sub-task**: {max_peak_kb/1000}k tokens
|
||||
```
|
||||
7. If VRAM detection failed or is not available, add a minimal section:
|
||||
```
|
||||
## VRAM Configuration
|
||||
- **Auto-detect**: Yes
|
||||
```
|
||||
8. Report the VRAM configuration status in the onboarding report.
|
||||
|
||||
## Output
|
||||
|
||||
Create or update the following inside {project}/.agent-framework/:
|
||||
@@ -50,6 +77,7 @@ Then produce a file called ONBOARDING_REPORT.md at {project}/tasks/onboarding/ON
|
||||
- Key observations from the project structure
|
||||
- Any missing pieces the human should provide next
|
||||
- **Upgrade status**: Whether the project's framework files are up to date with the global framework
|
||||
- **VRAM Configuration**: Whether VRAM config was auto-detected and applied (if available), or needs manual setup
|
||||
|
||||
When the ritual is complete, output "CONTRACT_MET" and stop.
|
||||
|
||||
|
||||
+229
-5
@@ -4,8 +4,113 @@ You are the Orchestrator Driver. Your job is to act as a **state machine** for t
|
||||
|
||||
1. {project}/.agent-framework/AGENT.md — Check if Autopilot is enabled
|
||||
2. {project}/.agent-framework/RULES.md
|
||||
3. {project}/.agent-framework/prompts/workflow.md — The State Machine
|
||||
4. Any existing files under {project}/tasks/
|
||||
3. ~/.agent-framework/config.md — Global framework configuration (VRAM, model settings)
|
||||
4. {project}/.agent-framework/AGENT.md (if exists — for model override)
|
||||
5. {project}/.agent-framework/prompts/workflow.md — The State Machine
|
||||
6. Any existing files under {project}/tasks/
|
||||
7. {project}/.agent-framework/scripts/vram_detect.sh — VRAM detection script (if exists)
|
||||
|
||||
## VRAM Detection
|
||||
|
||||
When VRAM configuration is needed (during task decomposition, sub-task creation, etc.), the Orchestrator MUST attempt to detect VRAM/VRAM limits dynamically.
|
||||
|
||||
### Detection Priority
|
||||
|
||||
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to probe GPU VRAM, RAM, and model context window. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`.
|
||||
2. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
|
||||
3. **Manual override**: Check if `~/.agent-framework/config.md` has `Auto-detect: No` under VRAM Configuration. If so, use the manually specified values.
|
||||
4. **Fallback**: Use 8k tokens as default, with 25% headroom.
|
||||
|
||||
### How to Read VRAM Config from config.md
|
||||
|
||||
```markdown
|
||||
## VRAM Configuration
|
||||
- **Auto-detect**: Yes/No
|
||||
- **Target context**: {value}k tokens (override if Auto-detect: No)
|
||||
- **Headroom**: {value}% (override if Auto-detect: No)
|
||||
- **Max peak context per sub-task**: {value}k tokens (override if Auto-detect: No)
|
||||
```
|
||||
|
||||
- If `Auto-detect: Yes`, run the detection script and use its output.
|
||||
- If `Auto-detect: No`, use the manually specified values.
|
||||
- If config.md has no VRAM Configuration section, default to `Auto-detect: Yes`.
|
||||
|
||||
### Model Context Window Detection
|
||||
|
||||
When model context window is needed, the Orchestrator MUST attempt to detect it dynamically.
|
||||
|
||||
#### Detection Priority
|
||||
|
||||
1. **Auto-detect via script**: Run `{project}/.agent-framework/scripts/vram_detect.sh` to detect the model name and its context window. Parse the JSON output for `model_context_kb`.
|
||||
2. **Auto-detect via config**: Check `~/.agent-framework/config.md` for the model name and override context window.
|
||||
3. **Auto-detect via API config**: If the script is not available, try to detect the model name from `AGENT.md` or config files (`.env`, `config.yaml`, etc.) and look up its context window.
|
||||
4. **Fallback**: Use 128k tokens as default (common for modern models).
|
||||
|
||||
#### How to Read Model Config from config.md
|
||||
|
||||
```markdown
|
||||
## Model Configuration
|
||||
- **Model**: auto # Use auto-detection, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
|
||||
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
|
||||
```
|
||||
|
||||
- If `Model: auto`, detect the model name from API config files or AGENT.md.
|
||||
- If `Override context window: auto`, use the detected context window.
|
||||
- If both are specified, use the specified values.
|
||||
|
||||
#### Model Name Lookup
|
||||
|
||||
When the model name is detected, look up its context window:
|
||||
- gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-4: 128k tokens
|
||||
- claude-3-5-sonnet, claude-3-5-haiku, claude-3-opus, claude-3-sonnet, claude-3-haiku, claude-2: 200k tokens
|
||||
|
||||
### Detection Script Output (JSON)
|
||||
|
||||
The detection script outputs JSON like:
|
||||
```json
|
||||
{
|
||||
"gpu_vram_gb": 8,
|
||||
"ram_gb": 16,
|
||||
"model_context_kb": 128000,
|
||||
"framework_overhead_tokens": 4000,
|
||||
"recommended_kb": 16000,
|
||||
"recommended_k": 16,
|
||||
"headroom": 0.25,
|
||||
"max_peak_context_kb": 12000
|
||||
}
|
||||
```
|
||||
|
||||
Use `recommended_kb` for the target context, `headroom` for headroom, and `max_peak_context_kb` for max peak context per sub-task.
|
||||
|
||||
### Auto-Detect When to Run Detection
|
||||
|
||||
The Orchestrator should run VRAM detection in the following scenarios:
|
||||
|
||||
1. **When a new task is created** — to set the VRAM config for the new task.
|
||||
2. **When Decomposition is triggered** — to ensure sub-tasks are sized correctly.
|
||||
3. **When sub-tasks are created** — to propagate VRAM config to sub-task folders.
|
||||
4. **When a sub-task's VRAM_CONFIG.md is missing** — to create one with auto-detected values.
|
||||
|
||||
### Reporting Detection Results
|
||||
|
||||
When auto-detecting VRAM, the Orchestrator should report:
|
||||
- GPU VRAM detected (if any)
|
||||
- System RAM detected
|
||||
- Model context window detected (if any)
|
||||
- Framework overhead estimated
|
||||
- Recommended VRAM context window
|
||||
- Whether auto-detection was used or manual override
|
||||
|
||||
Example:
|
||||
```
|
||||
VRAM Detection Results:
|
||||
- GPU VRAM: 8GB (nvidia-smi)
|
||||
- RAM: 16GB
|
||||
- Model context window: 128k (API-based)
|
||||
- Framework overhead: ~4k tokens
|
||||
- **Recommended: 16k tokens** (GPU VRAM-based, 25% headroom)
|
||||
- **Using: 16k tokens** (auto-detected)
|
||||
```
|
||||
|
||||
## Task
|
||||
|
||||
@@ -22,7 +127,8 @@ Each task is a state machine. The Orchestrator determines the current state and
|
||||
| State | Condition | Next State (Autopilot) |
|
||||
|-------|-----------|----------------------|
|
||||
| **New** | No artifacts in task folder | Research |
|
||||
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
|
||||
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement |
|
||||
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research |
|
||||
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
|
||||
| **Test Design** | Has `TEST_PLAN.md` | Implement |
|
||||
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
|
||||
@@ -113,8 +219,9 @@ Examine the tasks/ directory and determine the state of each task folder. Check
|
||||
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
|
||||
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
|
||||
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
|
||||
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
|
||||
13. No artifacts → **New**
|
||||
12. Has `SPEC.md` and `DECOMPOSITION.md` → **Decomposition** (sub-tasks will be created)
|
||||
13. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
|
||||
14. No artifacts → **New**
|
||||
|
||||
## Output Format
|
||||
|
||||
@@ -172,3 +279,120 @@ In Autopilot mode, after outputting the task statuses, the Orchestrator MUST aut
|
||||
When finished, output "ORCHESTRATION_COMPLETE".
|
||||
|
||||
Only recommend one phase at a time. Do not suggest running multiple phases in parallel.
|
||||
|
||||
## Sub-Task Management
|
||||
|
||||
When a parent task has a `DECOMPOSITION.md` file, the Orchestrator must create sub-task folders for each sub-task listed.
|
||||
|
||||
### Sub-Task Folder Structure
|
||||
|
||||
When decomposing, the Orchestrator creates sub-tasks under the parent task's `subtasks/` folder:
|
||||
|
||||
```
|
||||
tasks/parent-task/ → Parent task (Research → Decomposition → complete)
|
||||
SPEC.md
|
||||
DECOMPOSITION.md
|
||||
subtasks/
|
||||
subtask-a/ → Sub-task (full lifecycle independently)
|
||||
SPEC.md
|
||||
DESIGN.md
|
||||
IMPLEMENTATION.md
|
||||
BUG_REPORT.md
|
||||
ADVERSARIAL_BUG_REPORT.md
|
||||
DOC_REVIEW.md
|
||||
VERDICT.md
|
||||
subtask-b/
|
||||
SPEC.md
|
||||
...
|
||||
```
|
||||
|
||||
### Sub-Task Creation Rules
|
||||
|
||||
When a parent task reaches the **Decomposition** phase (has `SPEC.md` and `DECOMPOSITION.md`):
|
||||
|
||||
1. **Read `~/.agent-framework/config.md`** to get the VRAM configuration and check if auto-detect is enabled.
|
||||
2. **If Auto-detect: Yes**, run `{project}/.agent-framework/scripts/vram_detect.sh` to detect VRAM limits. Parse the JSON output for `recommended_kb`, `headroom`, and `max_peak_context_kb`. Report the detection results.
|
||||
3. **If Auto-detect: No**, use the manually specified values from config.md.
|
||||
4. **Read `DECOMPOSITION.md`** to extract all sub-task names, dependencies, and their estimated token budgets.
|
||||
5. **Verify VRAM constraints**:
|
||||
- For each sub-task, check if its estimated peak context (from DECOMPOSITION.md) is within the max peak context.
|
||||
- If a sub-task exceeds the VRAM limit, create a warning and suggest splitting it further.
|
||||
6. **Create sub-task folders** under `tasks/{parent-task-name}/subtasks/{sub-task-name}/`:
|
||||
- For each sub-task, create an empty folder (no files yet — the sub-task starts at the Research phase).
|
||||
- The sub-task name should be a kebab-case version of the sub-task description from DECOMPOSITION.md.
|
||||
7. **Create a VRAM_CONFIG.md** file for each sub-task with the VRAM configuration.
|
||||
8. **Create a PARENT_SPEC.md** file for each sub-task with:
|
||||
- The parent task's SPEC.md content
|
||||
- A reference to the parent task name
|
||||
- The VRAM configuration (auto-detected or manual)
|
||||
- Any context the sub-task needs from the parent
|
||||
9. **Do NOT drive sub-tasks through the lifecycle** — sub-tasks are driven through their full lifecycle (Research → Decomposition (optional) → ... → Complete) independently.
|
||||
10. **The parent task is NOT complete** until ALL sub-tasks are in terminal state (Complete or Human Intervention).
|
||||
|
||||
### Sub-Task Lifecycle
|
||||
|
||||
Each sub-task follows the full lifecycle independently:
|
||||
- Starts at the **Research** phase (no artifacts in the sub-task folder)
|
||||
- Goes through Research → Decomposition (optional) → Design (optional) → Test Design (optional) → Implement → Bug Find → Adversarial Bug Find → Doc Review → Referee
|
||||
- Ends at **Complete** (VERDICT.md with PASS) or **Human Intervention** (VERDICT.md with FAIL/NEEDS_REVIEW)
|
||||
|
||||
### VRAM-Aware Sub-Task Splitting
|
||||
|
||||
If a sub-task's estimated peak context exceeds the VRAM limit from AGENT.md:
|
||||
1. Split the sub-task into smaller sub-tasks.
|
||||
2. Each new sub-task should fit within the VRAM limit.
|
||||
3. Update the DECOMPOSITION.md to reflect the new sub-tasks.
|
||||
4. Create the new sub-task folders.
|
||||
5. Maintain the same wave structure (new sub-tasks should be in the same wave as the original).
|
||||
|
||||
### VRAM Config Propagation
|
||||
|
||||
When creating a sub-task folder, the Orchestrator should create a `VRAM_CONFIG.md` file with:
|
||||
```markdown
|
||||
# VRAM Configuration for this sub-task
|
||||
|
||||
- **Auto-detect**: Yes/No
|
||||
- **Target VRAM context**: {from detection script or config.md override}
|
||||
- **Headroom**: {from detection script or config.md override}
|
||||
- **Max peak context per sub-task**: {from detection script or config.md override}
|
||||
- **GPU VRAM detected**: {value}GB (or "None")
|
||||
- **RAM detected**: {value}GB
|
||||
- **Model context window**: {value}k tokens (or "Unknown")
|
||||
- **Framework overhead**: ~{value} tokens
|
||||
- **This sub-task's estimated peak context**: {from DECOMPOSITION.md}
|
||||
- **Fits within VRAM**: Yes/No
|
||||
```
|
||||
|
||||
This ensures sub-tasks are aware of the VRAM constraints and can optimize their implementation accordingly.
|
||||
|
||||
### Sub-Task Parent Specification
|
||||
|
||||
When creating a sub-task folder, the Orchestrator should also create a `PARENT_SPEC.md` file that contains:
|
||||
- The parent task's SPEC.md content
|
||||
- A reference to the parent task name
|
||||
- Any context the sub-task needs from the parent
|
||||
|
||||
This ensures sub-tasks have all the information they need to implement their portion of the parent spec.
|
||||
|
||||
### Sub-Task Dependencies and Wave Management
|
||||
|
||||
Sub-tasks that are in the same wave (can run in parallel) should be driven through the lifecycle in parallel. Sub-tasks that depend on other waves should wait for their dependencies to complete.
|
||||
|
||||
### Sub-Task Completion and Parent Task
|
||||
|
||||
When a sub-task reaches a terminal state:
|
||||
- If the sub-task **PASSes**: The parent task can move to the next wave of sub-tasks.
|
||||
- If a sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator pauses and reports human intervention is required. The parent task is stuck until the failing sub-task is resolved.
|
||||
|
||||
When ALL sub-tasks are in terminal state:
|
||||
- If ALL sub-tasks **PASS**: The parent task is **Complete**.
|
||||
- If ANY sub-task **FAILs or NEEDS_REVIEW**: The Orchestrator creates fix/review tasks for the failing sub-tasks and reports human intervention is required.
|
||||
|
||||
### Sub-Task Verdict Reporting
|
||||
|
||||
When a sub-task reaches the Referee phase, the VERDICT.md should include:
|
||||
- The sub-task's own verdict (PASS/FAIL/NEEDS_REVIEW)
|
||||
- A reference to the parent task name
|
||||
- Any findings that affect the parent task
|
||||
|
||||
The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status.
|
||||
|
||||
@@ -10,6 +10,8 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
|
||||
7. {project}/tasks/{task-name}/DESIGN.md (if exists)
|
||||
8. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
|
||||
9. {project}/tasks/{task-name}/VRAM_CONFIG.md (if exists)
|
||||
10. {project}/tasks/{task-name}/PARENT_SPEC.md (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
@@ -38,6 +40,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
- Are functions small and focused?
|
||||
- Is there proper error handling?
|
||||
- Are there any obvious performance issues?
|
||||
- If VRAM_CONFIG.md exists: Is the code optimized for low-VRAM (small functions, streaming patterns, no large file loading)?
|
||||
|
||||
### Testing
|
||||
- Do all tests pass?
|
||||
|
||||
+2
-1
@@ -7,7 +7,8 @@ This file defines the linear progression of a task in the agent-framework. The O
|
||||
| Current State | Signal (Artifact) | Next Phase | Action |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` |
|
||||
| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code |
|
||||
| **Research** | Has `SPEC.md` | Decomposition (optional) or Design (optional) or Implement | Generate `DECOMPOSITION.md` or `DESIGN.md` or code |
|
||||
| **Decomposition** | Has `SPEC.md` and `DECOMPOSITION.md` | Sub-task Research | Orchestrator creates sub-task folders |
|
||||
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
|
||||
| **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests |
|
||||
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |
|
||||
|
||||
Executable
+549
@@ -0,0 +1,549 @@
|
||||
#!/usr/bin/env bash
|
||||
# VRAM/Context Detection Script
|
||||
# Detects GPU VRAM, system RAM, and model context window to recommend
|
||||
# a safe VRAM context window for task decomposition.
|
||||
#
|
||||
# Usage: ./vram_detect.sh [model_name]
|
||||
# - If model_name is provided, looks up its context window
|
||||
# - Otherwise, tries to detect from API config or config.md
|
||||
|
||||
set -uo pipefail # Don't exit on error - we want to continue even if detection fails
|
||||
|
||||
# ─── GPU VRAM Detection ───
|
||||
detect_gpu_vram() {
|
||||
local total_vram_kb=0
|
||||
local vram_per_gpu_kb=0
|
||||
local num_gpus=0
|
||||
|
||||
# Try nvidia-smi first (NVIDIA GPUs)
|
||||
if command -v nvidia-smi &>/dev/null; then
|
||||
local vram_kb
|
||||
# Use timeout to avoid hanging on nvidia-smi (e.g., driver not loaded)
|
||||
vram_kb=$(timeout 5 nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ' || true)
|
||||
# Validate that vram_kb is a positive number
|
||||
if [[ -n "$vram_kb" && "$vram_kb" =~ ^[0-9]+$ && "$vram_kb" -gt 0 ]]; then
|
||||
total_vram_kb=$((vram_kb * 1024)) # MB → KB
|
||||
vram_per_gpu_kb=$((total_vram_kb / (num_gpus+1)))
|
||||
num_gpus=1
|
||||
echo "GPU: NVIDIA (nvidia-smi available)"
|
||||
echo "VRAM per GPU: $((vram_kb / 1024))GB ($vram_kb MB)"
|
||||
else
|
||||
echo "GPU: NVIDIA (nvidia-smi available but driver not responding)"
|
||||
fi
|
||||
fi
|
||||
|
||||
# Fallback: lspci
|
||||
if [[ $total_vram_kb -eq 0 && $num_gpus -eq 0 ]]; then
|
||||
local gpu_info
|
||||
gpu_info=$(lspci 2>/dev/null | grep -i -E 'VGA|3D|Display' | head -5)
|
||||
if [[ -n "$gpu_info" ]]; then
|
||||
echo "GPU detected: $gpu_info"
|
||||
# Try to get VRAM from lspci -vnn memory regions
|
||||
# GPUs show VRAM as Memory regions in lspci
|
||||
# Parse patterns like: Memory at f800000000 (64-bit, prefetchable) [size=256M]
|
||||
local total_vram_mb=0
|
||||
while IFS= read -r line; do
|
||||
# Extract the size value from [size=256M] pattern
|
||||
local size_num
|
||||
size_num=$(echo "$line" | grep -oE '[0-9]+(M|G|K)' | head -1 || true)
|
||||
if [[ -n "$size_num" ]]; then
|
||||
local size_val
|
||||
size_val=$(echo "$size_num" | grep -oE '[0-9]+')
|
||||
local size_unit
|
||||
size_unit=$(echo "$size_num" | grep -oE '(M|G|K)')
|
||||
if [[ -n "$size_val" && -n "$size_unit" ]]; then
|
||||
case "$size_unit" in
|
||||
M) total_vram_mb=$((total_vram_mb + size_val)) ;;
|
||||
G) total_vram_mb=$((total_vram_mb + size_val * 1024)) ;;
|
||||
K) total_vram_mb=$((total_vram_mb + size_val / 1024)) ;;
|
||||
esac
|
||||
fi
|
||||
fi
|
||||
done < <(lspci -vnn 2>/dev/null | grep -i -A 15 -E 'VGA|3D|Display' | grep -i 'Memory at')
|
||||
if [[ $total_vram_mb -gt 0 ]]; then
|
||||
local total_vram_gb=$((total_vram_mb / 1024))
|
||||
local total_vram_mb_remain=$((total_vram_mb % 1024))
|
||||
echo "VRAM: $total_vram_mb MB ($total_vram_gb GB $total_vram_mb_remain MB)"
|
||||
else
|
||||
echo "VRAM: Could not determine from lspci"
|
||||
fi
|
||||
# Check for AMD GPU via amdgpu sysfs
|
||||
if lspci -vnn 2>/dev/null | grep -qi 'amd\|ati'; then
|
||||
local amdgpu_info
|
||||
amdgpu_info=$(ls /sys/kernel/debug/amdgpu/ 2>/dev/null | head -1)
|
||||
if [[ -n "$amdgpu_info" ]]; then
|
||||
local vram_total
|
||||
vram_total=$(cat /sys/kernel/debug/amdgpu/${amdgpu_info}/vram_total 2>/dev/null || echo 0)
|
||||
if [[ "$vram_total" -gt 0 ]]; then
|
||||
local vram_gb=$((vram_total / 1024 / 1024 / 1024))
|
||||
local vram_mb=$((vram_total / 1024 / 1024))
|
||||
echo "AMD GPU VRAM: ${vram_gb}GB (${vram_mb}MB)"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
echo "Total VRAM: $((total_vram_kb / 1024 / 1024))GB"
|
||||
echo "VRAM per GPU: $((vram_per_gpu_kb / 1024 / 1024))GB"
|
||||
echo "Num GPUs: $num_gpus"
|
||||
}
|
||||
|
||||
# ─── System RAM Detection ───
|
||||
detect_ram() {
|
||||
local total_kb=0
|
||||
local available_kb=0
|
||||
|
||||
if [[ -f /proc/meminfo ]]; then
|
||||
total_kb=$(grep MemTotal /proc/meminfo | awk '{print $2}')
|
||||
available_kb=$(grep MemAvailable /proc/meminfo | awk '{print $2}')
|
||||
if [[ $total_kb -gt 0 ]]; then
|
||||
echo "RAM: $((total_kb / 1024 / 1024))GB total, $((available_kb / 1024 / 1024))GB available"
|
||||
echo "$available_kb $total_kb"
|
||||
fi
|
||||
elif command -v sysctl &>/dev/null; then
|
||||
total_kb=$(sysctl -n hw.memsize 2>/dev/null | awk '{print $1 / 1024}')
|
||||
if [[ -n "$total_kb" && "$total_kb" -gt 0 ]]; then
|
||||
echo "RAM: $((total_kb / 1024))GB total"
|
||||
echo "$total_kb $total_kb" # Assume all available
|
||||
fi
|
||||
else
|
||||
echo "RAM: Could not detect"
|
||||
echo "0 0"
|
||||
fi
|
||||
}
|
||||
|
||||
# ─── Model Context Window Detection ───
|
||||
detect_model_context() {
|
||||
local model_name="$1"
|
||||
local context_kb=0
|
||||
|
||||
# If model name provided, look it up
|
||||
if [[ -n "$model_name" ]]; then
|
||||
case "$model_name" in
|
||||
gpt-4o|gpt-4o-2024-05-13|gpt-4o-2024-08-06)
|
||||
context_kb=128000; echo "Model: $model_name"
|
||||
echo "Context window: 128k tokens"
|
||||
;;
|
||||
gpt-4o-mini|gpt-4o-mini-2024-07-18)
|
||||
context_kb=128000; echo "Model: $model_name"
|
||||
echo "Context window: 128k tokens"
|
||||
;;
|
||||
gpt-4-turbo|gpt-4-turbo-2024-04-09)
|
||||
context_kb=128000; echo "Model: $model_name"
|
||||
echo "Context window: 128k tokens"
|
||||
;;
|
||||
gpt-4|gpt-4-0125-preview|gpt-4-1106-preview)
|
||||
context_kb=128000; echo "Model: $model_name"
|
||||
echo "Context window: 128k tokens"
|
||||
;;
|
||||
claude-3-5-sonnet|claude-3-5-sonnet-20241022)
|
||||
context_kb=200000; echo "Model: $model_name"
|
||||
echo "Context window: 200k tokens"
|
||||
;;
|
||||
claude-3-5-haiku|claude-3-5-haiku-20241022)
|
||||
context_kb=200000; echo "Model: $model_name"
|
||||
echo "Context window: 200k tokens"
|
||||
;;
|
||||
claude-3-opus|claude-3-opus-20240229)
|
||||
context_kb=200000; echo "Model: $model_name"
|
||||
echo "Context window: 200k tokens"
|
||||
;;
|
||||
claude-3-sonnet|claude-3-sonnet-20240229)
|
||||
context_kb=200000; echo "Model: $model_name"
|
||||
echo "Context window: 200k tokens"
|
||||
;;
|
||||
claude-3-haiku|claude-3-haiku-20240307)
|
||||
context_kb=200000; echo "Model: $model_name"
|
||||
echo "Context window: 200k tokens"
|
||||
;;
|
||||
claude-2|claude-2.1)
|
||||
context_kb=200000; echo "Model: $model_name"
|
||||
echo "Context window: 200k tokens"
|
||||
;;
|
||||
*)
|
||||
echo "Model: $model_name (unknown context window)"
|
||||
echo "0"
|
||||
;;
|
||||
esac
|
||||
echo "$context_kb"
|
||||
return
|
||||
fi
|
||||
|
||||
# Try to detect from config.md (global framework model settings)
|
||||
local project_dir="${1:-.}"
|
||||
local config_md="${HOME}/.agent-framework/config.md"
|
||||
local model_from_config=""
|
||||
local override_context=""
|
||||
if [[ -f "$config_md" ]]; then
|
||||
# Check for model name
|
||||
model_from_config=$(grep -i "model:" "$config_md" 2>/dev/null | grep -v "#" | grep -v "model_context" | grep -v "override" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
|
||||
# Check for override context window
|
||||
override_context=$(grep -i "override context" "$config_md" 2>/dev/null | grep -v "#" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
|
||||
fi
|
||||
|
||||
# If config.md specifies a model, use it
|
||||
if [[ -n "$model_from_config" ]]; then
|
||||
echo "Found model in config.md: $model_from_config"
|
||||
# Use override context window if specified in config.md
|
||||
if [[ -n "$override_context" && "$override_context" != "auto" ]]; then
|
||||
echo "Using override context window from config.md: $override_context"
|
||||
local override_kb
|
||||
override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true)
|
||||
if [[ -n "$override_kb" ]]; then
|
||||
echo "Context window: ${override_context} tokens (override)"
|
||||
echo "$((override_kb * 1000))"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
detect_model_context "$model_from_config"
|
||||
return
|
||||
fi
|
||||
|
||||
# Try to detect from AGENT.md (project-level model override)
|
||||
local agent_md="${project_dir}/.agent-framework/AGENT.md"
|
||||
if [[ -f "$agent_md" ]]; then
|
||||
local model_line
|
||||
model_line=$(grep -i "model" "$agent_md" 2>/dev/null | grep -v "#" | grep -v "target" | grep -v "headroom" | grep -v "peak" | grep -v "Auto-detect" | head -1)
|
||||
if [[ -n "$model_line" ]]; then
|
||||
echo "Found model in AGENT.md: $model_line"
|
||||
# Extract model name from the line
|
||||
local model
|
||||
model=$(echo "$model_line" | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
|
||||
if [[ -n "$model" ]]; then
|
||||
# Use override context window if specified in config.md
|
||||
if [[ -n "$override_context" && "$override_context" != "auto" ]]; then
|
||||
echo "Using override context window from config.md: $override_context"
|
||||
# Convert override_context to kb (e.g., 128k -> 128000, 200k -> 200000)
|
||||
local override_kb
|
||||
override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true)
|
||||
if [[ -n "$override_kb" ]]; then
|
||||
echo "Context window: ${override_context} tokens (override)"
|
||||
echo "$((override_kb * 1000))"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
detect_model_context "$model"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# Try to detect from common API config files
|
||||
local config_files=(
|
||||
".env"
|
||||
".env.local"
|
||||
"config.yaml"
|
||||
"config.yml"
|
||||
"config.json"
|
||||
"settings.yaml"
|
||||
".agent-framework/config.yaml"
|
||||
".agent-framework/config.json"
|
||||
)
|
||||
|
||||
for config_file in "${config_files[@]}"; do
|
||||
local abs_file=""
|
||||
for candidate in "${project_dir}/${config_file}" "${project_dir}/.agent-framework/${config_file}"; do
|
||||
if [[ -f "$candidate" ]]; then
|
||||
abs_file="$candidate"
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
if [[ -n "$abs_file" ]]; then
|
||||
local model
|
||||
model=$(grep -i "model" "$abs_file" 2>/dev/null | grep -v "#" | grep -v "context" | grep -v "max_tokens" | grep -v "temperature" | grep -v "stream" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]')
|
||||
if [[ -n "$model" ]]; then
|
||||
echo "Found model in $abs_file: $model"
|
||||
# Use override context window if specified in config.md
|
||||
if [[ -n "$override_context" && "$override_context" != "auto" ]]; then
|
||||
echo "Using override context window from config.md: $override_context"
|
||||
local override_kb
|
||||
override_kb=$(echo "$override_context" | sed 's/[kK]$//' | grep -oP '[0-9]+' || true)
|
||||
if [[ -n "$override_kb" ]]; then
|
||||
echo "Context window: ${override_context} tokens (override)"
|
||||
echo "$((override_kb * 1000))"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
detect_model_context "$model"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
echo "Model: Unknown (could not detect from AGENT.md or config files)"
|
||||
echo "0"
|
||||
}
|
||||
|
||||
# ─── Agent Framework Overhead Calculation ───
|
||||
calculate_overhead() {
|
||||
local project_dir="${1:-.}"
|
||||
local overhead_tokens=0
|
||||
|
||||
# Count tokens for the framework files that are loaded during orchestration
|
||||
# These are the files loaded during the most common phase (orchestration):
|
||||
# AGENT.md + RULES.md + workflow.md + orchestrate.md
|
||||
# Note: other phase files (decompose.md, implement.md, etc.) are only loaded during
|
||||
# their specific phases, so they don't contribute to the peak context during orchestration.
|
||||
local framework_files=(
|
||||
"${project_dir}/.agent-framework/AGENT.md"
|
||||
"${project_dir}/.agent-framework/RULES.md"
|
||||
"${project_dir}/.agent-framework/prompts/workflow.md"
|
||||
"${project_dir}/.agent-framework/prompts/orchestrate.md"
|
||||
)
|
||||
# Fallback: check home directory if project dir doesn't have framework
|
||||
if [[ ! -f "${project_dir}/.agent-framework/AGENT.md" ]]; then
|
||||
framework_files=(
|
||||
"${HOME}/.agent-framework/AGENT.md"
|
||||
"${HOME}/.agent-framework/RULES.md"
|
||||
"${HOME}/.agent-framework/prompts/workflow.md"
|
||||
"${HOME}/.agent-framework/prompts/orchestrate.md"
|
||||
)
|
||||
fi
|
||||
|
||||
for file in "${framework_files[@]}"; do
|
||||
if [[ -f "$file" ]]; then
|
||||
# Rough estimate: 1 token ≈ 4 characters (English text)
|
||||
local chars
|
||||
chars=$(wc -c < "$file" 2>/dev/null || echo 0)
|
||||
local tokens=$((chars / 4))
|
||||
overhead_tokens=$((overhead_tokens + tokens))
|
||||
echo " ${file##*/}: ~${tokens} tokens"
|
||||
fi
|
||||
done
|
||||
|
||||
echo "Framework overhead: ~${overhead_tokens} tokens"
|
||||
echo "$overhead_tokens"
|
||||
}
|
||||
|
||||
# ─── Recommendation Engine ───
|
||||
recommend_context() {
|
||||
local gpu_vram_gb="$1"
|
||||
local ram_gb="$2"
|
||||
local model_context_kb="$3"
|
||||
local overhead_tokens="$4"
|
||||
|
||||
# Read VRAM config from config.md if it exists
|
||||
local config_md="${HOME}/.agent-framework/config.md"
|
||||
local auto_detect="Yes"
|
||||
local target_context_kb=0
|
||||
local override_headroom=25
|
||||
local override_max_peak_kb=0
|
||||
if [[ -f "$config_md" ]]; then
|
||||
auto_detect=$(grep -i "auto-detect:" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | tr -d '[:space:]' || true)
|
||||
target_context_kb=$(grep -i "target.*context" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true)
|
||||
override_headroom=$(grep -i "headroom" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true)
|
||||
override_max_peak_kb=$(grep -i "max peak" "$config_md" 2>/dev/null | grep -v "#" | grep -i "vram" | head -1 | sed -E 's/.*[:=[:space:]]+//i' | grep -oP '\d+' || true)
|
||||
fi
|
||||
|
||||
# If auto-detect is disabled, use the manually specified values
|
||||
if [[ -n "$auto_detect" && "$auto_detect" == "No" ]]; then
|
||||
if [[ -n "$target_context_kb" ]]; then
|
||||
local max_peak_kb=${override_max_peak_kb:-0}
|
||||
if [[ $max_peak_kb -eq 0 && $headroom_pct -gt 0 ]]; then
|
||||
max_peak_kb=$((target_context_kb * (100 - headroom_pct) / 100))
|
||||
fi
|
||||
echo "$headroom_pct"
|
||||
echo "$target_context_kb"
|
||||
echo "$max_peak_kb"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
|
||||
local recommended_kb=0
|
||||
local headroom_pct=${override_headroom:-25} # Use override from config.md, or default to 25%
|
||||
|
||||
# Output headroom_pct first (for parent to read)
|
||||
# Then output recommended_kb
|
||||
# Then output max_peak_kb (only for manual mode)
|
||||
echo "$headroom_pct"
|
||||
|
||||
# Recommendation logic:
|
||||
# 1. If GPU VRAM >= 4GB: use VRAM (practical for local inference)
|
||||
# 2. If model context window is available: use it (for API inference)
|
||||
# 3. If GPU VRAM < 4GB but > 0: use RAM (VRAM too small for local inference)
|
||||
# 4. If no GPU VRAM and no model: use RAM as fallback
|
||||
|
||||
# If GPU VRAM >= 4GB, base it on VRAM
|
||||
if [[ $gpu_vram_gb -ge 4 ]]; then
|
||||
# Rule of thumb: 1GB VRAM ≈ 4k tokens for local LLMs
|
||||
# But we need to leave room for the model itself
|
||||
# For a model, each ~8k context tokens takes about ~3-5MB of GPU VRAM
|
||||
# So VRAM available for context = VRAM - model size - agent overhead
|
||||
# Conservative: 1GB VRAM ≈ 2k context tokens
|
||||
local vram_context_kb=$((gpu_vram_gb * 2000))
|
||||
|
||||
# Leave headroom for the model itself and agent overhead
|
||||
recommended_kb=$((vram_context_kb * (100 - headroom_pct) / 100))
|
||||
# If model context window is available, use it (for API inference)
|
||||
elif [[ $model_context_kb -gt 0 ]]; then
|
||||
# For API-based, we're limited by the model's context window
|
||||
# But we don't want to use the full window due to overhead
|
||||
recommended_kb=$((model_context_kb * (100 - headroom_pct) / 100))
|
||||
# Fallback: use RAM to estimate
|
||||
else
|
||||
# Moderate estimate for low-VRAM systems where VRAM is too small for local inference
|
||||
# but RAM is available. Use 0.75k tokens per GB of RAM as a moderate estimate.
|
||||
# This balances between being too conservative (0.5k/GB) and too generous (1k/GB).
|
||||
local ram_context_kb=$((ram_gb * 750))
|
||||
recommended_kb=$((ram_context_kb * (100 - headroom_pct) / 100))
|
||||
fi
|
||||
|
||||
# Subtract framework overhead
|
||||
local net_kb=$((recommended_kb - overhead_tokens))
|
||||
if [[ $net_kb -lt 0 ]]; then
|
||||
net_kb=0
|
||||
fi
|
||||
|
||||
# Output: headroom_pct, recommended_kb
|
||||
echo "$net_kb"
|
||||
}
|
||||
|
||||
# ─── Main ───
|
||||
main() {
|
||||
local model_name=""
|
||||
local project_dir="."
|
||||
|
||||
# Parse arguments
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--model|-m)
|
||||
model_name="$2"
|
||||
shift 2
|
||||
;;
|
||||
--project|-p)
|
||||
project_dir="$2"
|
||||
shift 2
|
||||
;;
|
||||
*)
|
||||
# Could be model name as first argument
|
||||
if [[ -z "$model_name" ]]; then
|
||||
model_name="$1"
|
||||
fi
|
||||
shift
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
echo "=== VRAM / Context Detection ==="
|
||||
echo ""
|
||||
|
||||
# Detect GPU VRAM
|
||||
echo "--- GPU VRAM ---"
|
||||
detect_gpu_vram
|
||||
local gpu_vram_kb=0
|
||||
local gpu_vram_gb=0
|
||||
# Use timeout to avoid hanging
|
||||
gpu_vram_kb=$(timeout 5 nvidia-smi --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | head -1 | tr -d ' ' || true)
|
||||
# Validate that vram_kb is a positive number
|
||||
if [[ -z "$gpu_vram_kb" || ! "$gpu_vram_kb" =~ ^[0-9]+$ || "$gpu_vram_kb" -le 0 ]]; then
|
||||
gpu_vram_kb=0
|
||||
fi
|
||||
# If nvidia-smi didn't work, try to detect from lspci (AMD GPUs)
|
||||
if [[ $gpu_vram_kb -eq 0 ]]; then
|
||||
echo " nvidia-smi failed, checking lspci for AMD GPU VRAM..."
|
||||
local total_vram_mb=0
|
||||
while IFS= read -r line; do
|
||||
local size_num
|
||||
size_num=$(echo "$line" | grep -oE '[0-9]+(M|G|K)' | head -1 || true)
|
||||
if [[ -n "$size_num" ]]; then
|
||||
local size_val
|
||||
size_val=$(echo "$size_num" | grep -oE '[0-9]+')
|
||||
local size_unit
|
||||
size_unit=$(echo "$size_num" | grep -oE '(M|G|K)')
|
||||
if [[ -n "$size_val" && -n "$size_unit" ]]; then
|
||||
case "$size_unit" in
|
||||
M) total_vram_mb=$((total_vram_mb + size_val)) ;;
|
||||
G) total_vram_mb=$((total_vram_mb + size_val * 1024)) ;;
|
||||
K) total_vram_mb=$((total_vram_mb + size_val / 1024)) ;;
|
||||
esac
|
||||
fi
|
||||
fi
|
||||
done < <(lspci -vnn 2>/dev/null | grep -i -A 15 -E 'VGA|3D|Display' | grep -i 'Memory at')
|
||||
if [[ $total_vram_mb -gt 0 ]]; then
|
||||
gpu_vram_kb=$((total_vram_mb * 1024))
|
||||
gpu_vram_gb=$((total_vram_mb / 1024))
|
||||
echo " AMD GPU VRAM from lspci: ${gpu_vram_gb}GB ($total_vram_mb MB)"
|
||||
else
|
||||
echo " No VRAM found from lspci"
|
||||
fi
|
||||
fi
|
||||
echo ""
|
||||
|
||||
# Detect RAM
|
||||
echo "--- RAM ---"
|
||||
detect_ram
|
||||
local ram_kb
|
||||
ram_kb=$(grep MemTotal /proc/meminfo 2>/dev/null | awk '{print $2}' || echo 0)
|
||||
local ram_gb=$((ram_kb / 1024 / 1024))
|
||||
echo ""
|
||||
|
||||
# Detect model context window
|
||||
echo "--- Model Context Window ---"
|
||||
detect_model_context "$model_name"
|
||||
local model_context_kb
|
||||
model_context_kb=$(detect_model_context "$model_name" | tail -1)
|
||||
echo ""
|
||||
|
||||
# Calculate framework overhead
|
||||
echo "--- Framework Overhead ---"
|
||||
calculate_overhead "$project_dir"
|
||||
local overhead_tokens
|
||||
overhead_tokens=$(calculate_overhead "$project_dir" | tail -1)
|
||||
echo ""
|
||||
|
||||
# Recommend context window (also outputs headroom_pct and recommended_kb)
|
||||
echo "--- Recommendation ---"
|
||||
local recommendation_output
|
||||
recommendation_output=$(recommend_context "$gpu_vram_gb" "$ram_gb" "$model_context_kb" "$overhead_tokens")
|
||||
local line_count
|
||||
line_count=$(echo "$recommendation_output" | wc -l)
|
||||
local recommended_kb
|
||||
local headroom_pct
|
||||
local max_peak_kb
|
||||
if [[ $line_count -ge 3 ]]; then
|
||||
# Manual mode: outputs headroom_pct, recommended_kb, max_peak_kb
|
||||
headroom_pct=$(echo "$recommendation_output" | head -1 | tr -d '[:space:]')
|
||||
recommended_kb=$(echo "$recommendation_output" | sed -n '2p' | tr -d '[:space:]')
|
||||
max_peak_kb=$(echo "$recommendation_output" | tail -1 | tr -d '[:space:]')
|
||||
else
|
||||
# Auto-detect mode: outputs headroom_pct, recommended_kb
|
||||
headroom_pct=$(echo "$recommendation_output" | head -1 | tr -d '[:space:]')
|
||||
recommended_kb=$(echo "$recommendation_output" | tail -1 | tr -d '[:space:]')
|
||||
# Calculate max peak context based on headroom
|
||||
max_peak_kb=$((recommended_kb * (100 - headroom_pct) / 100))
|
||||
fi
|
||||
|
||||
# Convert to human-readable
|
||||
local recommended_k
|
||||
if [[ $recommended_kb -gt 0 ]]; then
|
||||
recommended_k=$((recommended_kb / 1000))
|
||||
else
|
||||
recommended_k=8 # Default fallback
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "=== Recommended Configuration ==="
|
||||
echo "Target context: ${recommended_k}k tokens"
|
||||
echo "Headroom: ${headroom_pct}%"
|
||||
echo "Max peak context per sub-task: $((recommended_k * (100 - headroom_pct) / 100))k tokens"
|
||||
|
||||
# Output as JSON for programmatic use
|
||||
echo ""
|
||||
echo "=== JSON Output ==="
|
||||
cat <<EOF
|
||||
{
|
||||
"gpu_vram_gb": $gpu_vram_gb,
|
||||
"ram_gb": $ram_gb,
|
||||
"model_context_kb": $model_context_kb,
|
||||
"framework_overhead_tokens": $overhead_tokens,
|
||||
"recommended_kb": $recommended_kb,
|
||||
"recommended_k": $recommended_k,
|
||||
"headroom": 0.25,
|
||||
"max_peak_context_kb": $max_peak_kb
|
||||
}
|
||||
EOF
|
||||
}
|
||||
|
||||
main "$@"
|
||||
Reference in New Issue
Block a user