Integrate TDD and Grill with Docs into Design, Implementation, and Verification workflows
This commit is contained in:
+47
-90
@@ -1,109 +1,66 @@
|
||||
# Onboarding a Project
|
||||
|
||||
## Quick Checklist
|
||||
This document defines the **Agent Protocol** for initializing a new project. When the agent is asked to "Onboard a project," it must follow these steps.
|
||||
|
||||
- [ ] Create `.agent-framework/` directory in project root
|
||||
- [ ] Create `AGENT.md` (project level)
|
||||
- [ ] Create `RULES.md` (project level)
|
||||
- [ ] Run exploration ritual with agent (fresh session)
|
||||
- [ ] Agent reads global + project AGENT.md and RULES.md
|
||||
- [ ] Create first `tasks/{task-name}/` folder
|
||||
- [ ] Start research phase using `prompts/research.md`
|
||||
## The Exploration Ritual
|
||||
|
||||
## Autopilot Mode
|
||||
The agent's first task in any project is to perform an "Initial Exploration" to establish context.
|
||||
|
||||
The framework now includes an **Autopilot** mode driven by the `orchestrate.md` prompt.
|
||||
### Step 1: Discovery
|
||||
The agent must:
|
||||
1. Explore the project root using `ls` and `find`.
|
||||
2. Read `.agent-framework/AGENT.md`
|
||||
3. Read `.agent-framework/RULES.md`
|
||||
4. Read the global `~/.agent-framework/AGENT.md`
|
||||
|
||||
### How Autopilot Works
|
||||
|
||||
1. **State Detection**: The agent scans your `tasks/` directory and detects the current progress of each task based on the artifacts produced (e.g., `SPEC.md`, `IMPLEMENTATION.md`, etc.).
|
||||
2. **Proactive Driving**: Instead of waiting for you to tell it what to do next, the agent will recommend the exact command needed to progress to the next phase.
|
||||
3. **Verification Loop**: Every task follows a strict lifecycle:
|
||||
`Research` $\rightarrow$ `Implement` $\rightarrow$ `Bug Find` $\rightarrow$ `Adversarial Bug Find` $\rightarrow$ `Referee`.
|
||||
|
||||
### Using Autopilot
|
||||
|
||||
To enable Autopilot, add the following to your project's `AGENT.md`:
|
||||
|
||||
```markdown
|
||||
## Mode
|
||||
Autopilot: Enabled.
|
||||
- The agent should use the Orchestrator to drive tasks to completion.
|
||||
- When a task is completed, the agent should automatically scan for the next pending task.
|
||||
```
|
||||
|
||||
When you want the agent to drive the project, use the following command:
|
||||
> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."*
|
||||
### Step 2: Reporting
|
||||
The agent must report back with:
|
||||
- Confirmation that the framework files were found and read.
|
||||
- A summary of the project rules.
|
||||
- The expected workflow for this project.
|
||||
- Key observations from the project structure.
|
||||
|
||||
---
|
||||
|
||||
## Scenario A: Existing Project (Drop-In)
|
||||
## The Lifecycle of a Project
|
||||
|
||||
### 1. Create the Project Framework Directory
|
||||
```bash
|
||||
mkdir -p /path/to/project/.agent-framework
|
||||
```
|
||||
Once onboarded, the project moves through these phases. The agent should use the provided prompts to transition between them.
|
||||
|
||||
### 2. Create the Two Required Files
|
||||
### Phase 1: Research
|
||||
**Template**: `prompts/research.md`
|
||||
**Output**: `SPEC.md`
|
||||
**Trigger**: *"Research {task-description}"*
|
||||
|
||||
#### AGENT.md (project level)
|
||||
```markdown
|
||||
# AGENT.md (project-name)
|
||||
### Phase 2: Implementation
|
||||
**Template**: `prompts/implement.md`
|
||||
**Output**: Code changes + test results
|
||||
**Trigger**: *"Implement the {task-name} task"*
|
||||
|
||||
This project uses the global framework at ~/.agent-framework/.
|
||||
### Phase 3: Bug Finding
|
||||
**Template**: `prompts/bug_finder.md`
|
||||
**Output**: `BUG_REPORT.md`
|
||||
**Trigger**: *"Find bugs in the {task-name} task"*
|
||||
|
||||
Additional project rules are in RULES.md.
|
||||
### Phase 4: Adversarial Verification
|
||||
**Template**: `prompts/adversarial_bug_find.md`
|
||||
**Output**: `ADVERSARIAL_BUG_REPORT.md`
|
||||
**Trigger**: *"Perform adversarial bug find for {task-name}"*
|
||||
|
||||
## Mode
|
||||
Autopilot: Enabled.
|
||||
```
|
||||
### Phase 5: Referee
|
||||
**Template**: `prompts/referee.md`
|
||||
**Output**: `VERDICT.md`
|
||||
**Trigger**: *"Review the {task-name} task"*
|
||||
|
||||
#### RULES.md (project level)
|
||||
Start with any hard constraints you already know for this project.
|
||||
## Prompt Rendering Convention
|
||||
|
||||
### 3. Run the Initial Exploration Ritual
|
||||
Give the agent this prompt in a fresh session:
|
||||
All prompts are stored as template files in `~/.agent-framework/prompts/`. They use `{placeholder}` syntax.
|
||||
|
||||
```
|
||||
You have been given a new project at this path:
|
||||
/path/to/project/
|
||||
### Placeholders
|
||||
- `{project}`: Absolute path to the project root.
|
||||
- `{task-name}`: The task folder name (kebab-case).
|
||||
- `{task-description}`: A brief, clear summary of the current work.
|
||||
|
||||
Your first actions must be:
|
||||
1. Explore the project root using ls and find.
|
||||
2. Read .agent-framework/AGENT.md
|
||||
3. Read .agent-framework/RULES.md
|
||||
4. Read ~/.agent-framework/AGENT.md
|
||||
|
||||
Report back with:
|
||||
- Confirmation the framework files were found and read
|
||||
- Summary of the project rules
|
||||
- What process this project expects
|
||||
- Key observations from the project structure
|
||||
|
||||
Do not start any task yet.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Scenario B: Starting From Scratch
|
||||
|
||||
Follow the same steps as Scenario A, but start by asking the agent to design the initial architecture:
|
||||
|
||||
> *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"*
|
||||
|
||||
## Phase Prompts (Template Files)
|
||||
|
||||
| Phase | Template | Output | Trigger |
|
||||
|---|---|---|---|
|
||||
| **Research** | `research.md` | `SPEC.md` | *"Research {task-description}"* |
|
||||
| **Implementation** | `implement.md` | Code + Tests | *"Implement the {task-name} task"* |
|
||||
| **Bug Find** | `bug_finder.md` | `BUG_REPORT.md` | *"Find bugs in the {task-name} task"* |
|
||||
| **Adversarial** | `adversarial_bug_find.md` | `ADVERSARIAL_BUG_REPORT.md` | *"Perform adversarial bug find for {task-name}"* |
|
||||
| **Referee** | `referee.md` | `VERDICT.md` | *"Review the {task-name} task"* |
|
||||
|
||||
## Core Principles
|
||||
|
||||
- **Context Is Everything**: Separate research from implementation. Use fresh sessions per task.
|
||||
- **Sycophancy Management**: Use the multi-agent validation loop (Bug Finder $\rightarrow$ Adversarial $\rightarrow$ Referee) to ensure objective results.
|
||||
- **Clear End States**: Tests are mandatory. A task is not complete until all acceptance criteria in the `SPEC.md` (or `CONTRACT.md`) are met.
|
||||
- **Rules + Skills**: Keep `RULES.md` specific to this project. Use the global framework for general behavior.
|
||||
When the agent receives a trigger command, it must:
|
||||
1. Read the corresponding template file.
|
||||
2. Replace all `{placeholders}` with the actual project values.
|
||||
3. Execute the rendered prompt.
|
||||
|
||||
@@ -2,8 +2,48 @@
|
||||
|
||||
A contract-based operating system for LLM agents, designed to enforce disciplined engineering workflows.
|
||||
|
||||
## Core Philosophy
|
||||
The framework prevents agents from "hallucinating" features or jumping into code without a plan. It enforces a strict separation between **Planning (Research)** and **Execution (Implementation)**, mediated by a **Verification** loop.
|
||||
## 1. Global Framework Installation (One-time setup)
|
||||
|
||||
Before you can use the framework in any project, you must install the core logic into your local environment.
|
||||
|
||||
**Run these commands in your terminal:**
|
||||
|
||||
```bash
|
||||
# Clone the framework into the global config directory
|
||||
git clone [INSERT_FRAMEWORK_REPO_URL_HERE] ~/.agent-framework
|
||||
|
||||
# Enter the directory
|
||||
cd ~/.agent-framework
|
||||
|
||||
# Make the installation script executable and run it
|
||||
chmod +x install.sh
|
||||
./install.sh
|
||||
```
|
||||
*Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory.*
|
||||
|
||||
---
|
||||
|
||||
## 2. Project Setup (Per project)
|
||||
|
||||
Once the framework is installed globally, you must "onboard" every individual project you work on.
|
||||
|
||||
### Option A: The Agent-Driven Way (Recommended)
|
||||
If you want the agent to handle the configuration for you, navigate to your project root and run:
|
||||
|
||||
> *"Onboard this project into the agent-framework."*
|
||||
|
||||
The agent will automatically:
|
||||
1. Detect your project type (New, Existing, or Upgrade).
|
||||
2. Create the `./.agent-framework/` directory.
|
||||
3. Generate your `AGENT.md` and `RULES.md` files.
|
||||
4. Initiate the "Exploration Ritual" to understand your codebase.
|
||||
|
||||
### Option B: The Manual Way
|
||||
If you prefer to set it up manually, create a `.agent-framework/` directory in your project root and add:
|
||||
- `AGENT.md`: Project-specific configuration (Mode, rules, etc.).
|
||||
- `RULES.md`: Project-specific constraints and past failure modes.
|
||||
|
||||
---
|
||||
|
||||
## The Autopilot Workflow
|
||||
The framework features an **Autopilot** mode that allows the agent to drive a project to completion with minimal intervention.
|
||||
@@ -20,17 +60,11 @@ Add `Autopilot: Enabled` to your project's `.agent-framework/AGENT.md` and use t
|
||||
|
||||
> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."*
|
||||
|
||||
The agent will then:
|
||||
- Scan the `tasks/` directory.
|
||||
- Identify the current phase of every task based on existing artifacts.
|
||||
- Recommend the next command to progress each task.
|
||||
- Automatically move to the next phase once the current one's artifacts are produced.
|
||||
|
||||
## Key Components
|
||||
- `AGENT.md`: Project-specific configuration and mode selection.
|
||||
- `RULES.md`: Living document of project constraints and past failure modes.
|
||||
- `prompts/`: Specialized system prompts for each phase (Research, Implement, Bug Finder, etc.).
|
||||
- `workflow.md`: The state machine governing the Autopilot lifecycle.
|
||||
|
||||
## Installation
|
||||
Clone the repository and set up your project following the instructions in `ONBOARDING.md`.
|
||||
## Contact & Support
|
||||
[Insert Contact Info]
|
||||
|
||||
+17
-64
@@ -1,92 +1,45 @@
|
||||
You are the Bug Finder. Your job is to adversarially test the implementation and find every bug, deviation from spec, and edge case.
|
||||
You are the Bug Finder. Your job is to find every bug, deviation from spec, and edge case.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
|
||||
3. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists)
|
||||
1. {project}/tasks/{task-name}/SPEC.md
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
|
||||
3. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Adversarial Checklist
|
||||
## Checklist
|
||||
|
||||
### Spec Compliance
|
||||
- Does the implementation match the SPEC.md exactly?
|
||||
- Are there missing features or stubs?
|
||||
- Are there features not in the spec (scope creep)?
|
||||
|
||||
### Edge Cases
|
||||
- Empty inputs, null values, zero-length arrays
|
||||
- Large inputs (performance, memory)
|
||||
- Malformed data, unexpected types
|
||||
- Concurrent access, race conditions
|
||||
- Dependency failures (network, database, API)
|
||||
|
||||
### Security
|
||||
- SQL injection, XSS, CSRF
|
||||
- Authentication and authorization gaps
|
||||
- Data exposure (logs, error messages, API responses)
|
||||
- Rate limiting, input validation
|
||||
- File upload, path traversal
|
||||
|
||||
### Data Flow
|
||||
- Trace data from input to output
|
||||
- Are mutations safe?
|
||||
- Is sensitive data exposed?
|
||||
- Can data be lost or corrupted?
|
||||
|
||||
### Concurrency & Race Conditions
|
||||
- Shared state without synchronization
|
||||
- Async operations without error handling
|
||||
- Deadlocks, livelocks
|
||||
- Transaction isolation issues
|
||||
|
||||
### Error Handling
|
||||
- Are all errors caught and logged?
|
||||
- Are silent failures possible?
|
||||
- Are swallowed exceptions present?
|
||||
- Is there graceful degradation?
|
||||
|
||||
### Performance
|
||||
- O(n^2) or worse algorithms
|
||||
- N+1 query patterns
|
||||
- Memory leaks, unbounded caches
|
||||
- Unbounded loops, infinite recursion
|
||||
|
||||
### Testing
|
||||
- Are all edge cases covered by tests?
|
||||
- Are tests actually testing the right things?
|
||||
- Are there false positives (tests that pass but don't verify)?
|
||||
- **Spec**: Does it match the SPEC.md exactly? Are there missing features or scope creep?
|
||||
- **Edges**: Check nulls, empties, large inputs, malformed data, and concurrency.
|
||||
- **Security**: Check for injection, auth gaps, data exposure, and input validation.
|
||||
- **Performance**: Look for O(n^2)+, N+1 queries, memory leaks, and infinite loops.
|
||||
- **Errors**: Are all errors caught, logged, and handled gracefully?
|
||||
|
||||
## Output Format
|
||||
|
||||
Produce a BUG_REPORT.md at {project}/tasks/{task-name}/BUG_REPORT.md with:
|
||||
Produce a BUG_REPORT.md at {project}/tasks/{task-name}/BUG_REPORT.md:
|
||||
|
||||
```markdown
|
||||
# Bug Report: {task-name}
|
||||
|
||||
## Summary
|
||||
{Brief overview of findings}
|
||||
{Brief overview}
|
||||
|
||||
## Bugs Found
|
||||
|
||||
### Bug 1: {Title}
|
||||
- **Severity**: Critical / High / Medium / Low
|
||||
- **Description**: {What's wrong}
|
||||
- **Location**: {File:line}
|
||||
- **Reproduction**: {Steps to reproduce}
|
||||
- **Suggested Fix**: {How to fix}
|
||||
|
||||
### Bug 2: ...
|
||||
- **Description**: {What's wrong}
|
||||
- **Reproduction**: {Steps}
|
||||
- **Suggested Fix**: {Fix}
|
||||
|
||||
## Score
|
||||
{Assign a score: +1 for low, +5 for medium, +10 for critical}
|
||||
```
|
||||
|
||||
## Important
|
||||
|
||||
- Be aggressive. Your job is to find bugs, not to be nice.
|
||||
- If you find nothing, say so explicitly — but double-check everything first.
|
||||
- Do NOT invent bugs. Only report real issues.
|
||||
- Be aggressive. Do NOT invent bugs.
|
||||
- If no bugs found, state it explicitly.
|
||||
|
||||
+14
-1
@@ -11,6 +11,15 @@ Your job is to create a clear, actionable design for the project based on the sp
|
||||
|
||||
{task-description}
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md
|
||||
2. {project}/.agent-framework/RULES.md
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Output
|
||||
|
||||
Produce a file called `DESIGN.md` at {project}/tasks/{task-name}/DESIGN.md containing:
|
||||
@@ -38,7 +47,11 @@ Produce a file called `DESIGN.md` at {project}/tasks/{task-name}/DESIGN.md conta
|
||||
- Biggest technical or product risks
|
||||
- Areas that need exploration or validation first
|
||||
|
||||
### 6. Non-Functional Requirements
|
||||
### 6. Documentation Plan
|
||||
- List all specific documentation that must be updated or created (e.g., README sections, API docs, docstrings).
|
||||
- Define the "source of truth" for each piece of documentation.
|
||||
|
||||
### 7. Non-Functional Requirements
|
||||
- Performance, security, reliability, or scale considerations (if relevant)
|
||||
|
||||
When the design is complete, output "CONTRACT_MET" and stop.
|
||||
|
||||
+16
-8
@@ -2,25 +2,33 @@ You are in implementation mode.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md — what needs to be built
|
||||
2. {project}/.agent-framework/RULES.md — project-specific rules
|
||||
3. {project}/.agent-framework/AGENT.md — project agent config (if exists)
|
||||
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
|
||||
1. {project}/tasks/{task-name}/SPEC.md
|
||||
2. {project}/.agent-framework/RULES.md
|
||||
3. {project}/.agent-framework/AGENT.md (if exists)
|
||||
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
|
||||
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Implementation Rules
|
||||
## Implementation Rules (TDD Mode)
|
||||
|
||||
- Follow the SPEC.md exactly. Do not add features not listed.
|
||||
- Write tests first when a correct seam exists.
|
||||
- Follow the SPEC.md and DESIGN.md exactly.
|
||||
- **Strict TDD Loop**: You are NOT allowed to write a feature in one go. You must follow the Red-Green-Refactor cycle:
|
||||
1. **RED**: Write a failing test for the smallest possible unit of the feature.
|
||||
2. **GREEN**: Write the minimum amount of code required to make that test pass.
|
||||
3. **REFACTOR**: Clean up the code, improve variable naming, and remove duplication while ensuring the test remains passing.
|
||||
- **Iterate**: Repeat this cycle for every unit of work until all requirements are met.
|
||||
- Write tests first.
|
||||
- Run tests and report full output.
|
||||
- Keep functions small and focused.
|
||||
- Use existing patterns in the codebase.
|
||||
|
||||
### End State
|
||||
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
|
||||
- All tests must pass.
|
||||
- **Documentation**: All documentation identified in the DESIGN.md must be updated or created.
|
||||
- Run the full test suite and report results
|
||||
- Do NOT declare victory until tests pass
|
||||
|
||||
@@ -34,4 +42,4 @@ Report back with:
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced all required deliverables AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
Until then, continue working or ask clarifying questions.
|
||||
|
||||
+13
-6
@@ -2,11 +2,12 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
|
||||
3. {project}/tasks/{task-name}/BUG_REPORT.md — bugs found by Bug Finder (if exists)
|
||||
4. {project}/tasks/{task-name}/ADVERSARIAL_BUG_REPORT.md — bugs found by Adversarial Bug Finder (if exists)
|
||||
5. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists)
|
||||
1. {project}/tasks/{task-name}/SPEC.md
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
|
||||
3. {project}/tasks/{task-name}/BUG_REPORT.md (if exists)
|
||||
4. {project}/tasks/{task-name}/ADVERSARIAL_BUG_REPORT.md (if exists)
|
||||
5. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
|
||||
6. {project}/tasks/{task-name}/DESIGN.md (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
@@ -41,6 +42,12 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
- Are edge cases covered?
|
||||
- Are there false positives (tests that pass but don't verify)?
|
||||
|
||||
### Documentation Review (Grill with Docs)
|
||||
- Did the agent update all documentation identified in the DESIGN.md?
|
||||
- Is the documentation accurate and reflects the final implementation?
|
||||
- Is the documentation clear enough for a developer to understand the new changes?
|
||||
- Does the documentation cover any edge cases or non-obvious logic?
|
||||
|
||||
## Verdict
|
||||
|
||||
Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with:
|
||||
@@ -77,4 +84,4 @@ Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with:
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced the VERDICT.md file AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
Until then, continue working or ask clarifying questions.
|
||||
|
||||
Reference in New Issue
Block a user