Files
agent-framework/README.md
T

179 lines
8.4 KiB
Markdown

# Agent Framework
A contract-based operating system for LLM agents, designed to enforce disciplined engineering workflows.
## 1. Global Framework Installation (One-time setup)
Before you can use the framework in any project, you must install the core logic into your local environment.
**Run these commands in your terminal:**
```bash
# Clone the framework into the global config directory
git clone [INSERT_FRAMEWORK_REPO_URL_HERE] ~/.agent-framework
# Enter the directory
cd ~/.agent-framework
# Make the installation script executable and run it
chmod +x install.sh
./install.sh
```
*Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory.*
### Updating the Framework
When the framework is updated, you can update your global installation:
```bash
cd ~/.agent-framework
./update.sh
```
This will:
- Fetch the latest changes from the repository
- Check for uncommitted changes and warn you
- Pull the latest updates
**Upgrading existing projects:** When the framework is updated, existing projects may need their framework files upgraded (new phases added, new prompts, etc.). To upgrade an existing project, tell the agent: "Upgrade the agent-framework for this project." The agent will check for missing files and update them.
---
## 2. Project Setup (Per project)
Once the framework is installed globally, you must "onboard" every individual project you work on.
### Option A: The Agent-Driven Way (Recommended)
If you want the agent to handle the configuration for you, navigate to your project root and run:
> *"Onboard this project into the agent-framework."*
The agent will automatically:
1. Detect your project type (New, Existing, or Upgrade).
2. Create the `./.agent-framework/` directory.
3. Generate your `AGENT.md` and `RULES.md` files.
4. Initiate the "Exploration Ritual" to understand your codebase.
### Option B: The Manual Way
If you prefer to set it up manually, create a `.agent-framework/` directory in your project root and add:
- `AGENT.md`: Project-specific configuration (Mode, rules, etc.).
- `RULES.md`: Project-specific constraints and past failure modes.
---
## The Autopilot Workflow
The framework features an **Autopilot** mode that allows the agent to drive a project to completion with minimal intervention.
### Lifecycle of a Task
1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. **Interactive** — agent grills you for requirements and gets sign-off.
2. **Decomposition** (optional): Break the task into sub-tasks sized for your VRAM. **Interactive** — agent analyzes the spec and proposes sub-tasks with token budget estimates, gets sign-off. See [VRAM Configuration](#vram-configuration) below.
3. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off.
3. **Test Design** (optional): Produce a `TEST_PLAN.md` (Test specification). No code allowed. **Interactive** — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off.
4. **Implement**: Write code and tests based *only* on the `SPEC.md`, `DESIGN.md` (if present), and `TEST_PLAN.md` (if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows.
5. **Bug Find**: Aggressive search for bugs and spec deviations.
6. **Adversarial Bug Find**: Deep search for complex logic errors, race conditions, and performance issues.
7. **Doc Review**: Review documentation against DESIGN.md plan and fix missing docs.
8. **Referee**: Objective evaluation of all bugs, docs, and the final verdict.
### How to use Autopilot
#### Autopilot mode (default)
The default mode is **Autopilot: Enabled**. The Orchestrator automatically drives tasks through all phases. Just say:
- **Start a new task**: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion)
- **Decompose a task**: "Decompose the add-user-auth task" (the Orchestrator breaks it into sub-tasks sized for your VRAM)
- **Continue**: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it)
#### VRAM Configuration
For low-VRAM systems (8GB, 16GB), set your VRAM limits in `~/.agent-framework/config.md`:
```markdown
## VRAM Configuration
- **Auto-detect**: Yes # Let the agent detect your VRAM automatically
- **Target VRAM context**: 16k tokens # Override auto-detect if needed
- **Headroom**: 25%
- **Max peak context per sub-task**: 12k tokens
```
**Auto-detection**: When `Auto-detect: Yes`, the Orchestrator runs `scripts/vram_detect.sh` to probe:
- GPU VRAM (via `nvidia-smi`)
- System RAM (via `free`)
- Model context window (from config.md or API config files)
- Framework overhead (by counting token load in loaded prompts)
**Manual override**: When `Auto-detect: No`, use the manually specified values:
```markdown
## VRAM Configuration
- **Auto-detect**: No
- **Target VRAM context**: 8k
- **Headroom**: 30%
- **Max peak context per sub-task**: 5.6k
```
#### Model Configuration
When using a local LLM or a specific API model, set the model in `~/.agent-framework/config.md`:
```markdown
## Model Configuration
- **Model**: auto # Use auto-detection from API config files
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
```
**Auto-detection**: When `Model: auto`, the framework detects the model name from API config files (`.env`, `config.yaml`, etc.) and looks up its context window.
**Manual override**: When you know your model name, specify it:
```markdown
## Model Configuration
- **Model**: gpt-4o
- **Override context window**: 128k
```
When you run "Decompose the X task", the Orchestrator will:
1. Analyze the task's SPEC.md
2. Detect VRAM limits (auto or manual)
3. Break it into sub-tasks, each sized to fit within your VRAM limit
4. Estimate the token budget for each sub-task
5. Create sub-task folders under `tasks/{parent-task}/subtasks/{sub-task}/`
6. Propagate VRAM config to each sub-task
Sub-tasks run independently through the full lifecycle. The parent task is complete only when ALL sub-tasks pass.
#### Manual mode (opt-in)
Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like:
- "Research add user authentication" — starts a new task
- "Decompose the add-user-auth task" — breaks into VRAM-sized sub-tasks (optional)
- "Design the add-user-auth task" — designs the architecture (optional)
- "Design tests for the add-user-auth task" — designs test cases (optional)
- "Implement the add-user-auth task" — implements the task with tests
- "Find bugs in the add-user-auth task" — finds bugs
- "Perform adversarial bug find for add-user-auth" — deep bug search
- "Review docs for the add-user-auth task" — reviews documentation
- "Review the add-user-auth task" — referee evaluates
- "orchestrate" — asks the Orchestrator what to do next
## Sub-Task Management
When a task is decomposed, the Orchestrator creates sub-tasks under `tasks/{parent-task}/subtasks/`. Each sub-task:
- Has its own full lifecycle (Research → Decomposition → Design → ... → Complete)
- Starts at the Research phase with an empty folder
- Receives a `PARENT_SPEC.md` with the parent task's context
- Receives a `VRAM_CONFIG.md` with the VRAM constraints
- Runs independently — sub-tasks in the same wave can run in parallel
- The parent task is NOT complete until ALL sub-tasks pass
## Key Components
- `AGENT.md`: Project-specific agent behavior (Autopilot mode, routing rules).
- `config.md`: Global framework settings (VRAM, model, system requirements).
- `RULES.md`: Living document of project constraints and past failure modes.
- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, Decompose, etc.).
- `workflow.md`: The state machine governing the Autopilot lifecycle.
- `test_design.md`: Produces a TEST_PLAN.md — an explicit test specification before implementation.
- `decompose.md`: Breaks a task into VRAM-sized sub-tasks.
- `scripts/vram_detect.sh`: Auto-detects GPU VRAM, RAM, model context window, and framework overhead.
- `contracts/vram_config.md`: Contract for VRAM-aware task decomposition.
## Contact & Support
[Insert Contact Info]