Integrate TDD and Grill with Docs into Design, Implementation, and Verification workflows

This commit is contained in:
2026-06-10 14:46:11 -04:00
parent 59b6339765
commit 3cdb95e083
6 changed files with 151 additions and 179 deletions
+47 -90
View File
@@ -1,109 +1,66 @@
# Onboarding a Project
## Quick Checklist
This document defines the **Agent Protocol** for initializing a new project. When the agent is asked to "Onboard a project," it must follow these steps.
- [ ] Create `.agent-framework/` directory in project root
- [ ] Create `AGENT.md` (project level)
- [ ] Create `RULES.md` (project level)
- [ ] Run exploration ritual with agent (fresh session)
- [ ] Agent reads global + project AGENT.md and RULES.md
- [ ] Create first `tasks/{task-name}/` folder
- [ ] Start research phase using `prompts/research.md`
## The Exploration Ritual
## Autopilot Mode
The agent's first task in any project is to perform an "Initial Exploration" to establish context.
The framework now includes an **Autopilot** mode driven by the `orchestrate.md` prompt.
### Step 1: Discovery
The agent must:
1. Explore the project root using `ls` and `find`.
2. Read `.agent-framework/AGENT.md`
3. Read `.agent-framework/RULES.md`
4. Read the global `~/.agent-framework/AGENT.md`
### How Autopilot Works
1. **State Detection**: The agent scans your `tasks/` directory and detects the current progress of each task based on the artifacts produced (e.g., `SPEC.md`, `IMPLEMENTATION.md`, etc.).
2. **Proactive Driving**: Instead of waiting for you to tell it what to do next, the agent will recommend the exact command needed to progress to the next phase.
3. **Verification Loop**: Every task follows a strict lifecycle:
`Research` $\rightarrow$ `Implement` $\rightarrow$ `Bug Find` $\rightarrow$ `Adversarial Bug Find` $\rightarrow$ `Referee`.
### Using Autopilot
To enable Autopilot, add the following to your project's `AGENT.md`:
```markdown
## Mode
Autopilot: Enabled.
- The agent should use the Orchestrator to drive tasks to completion.
- When a task is completed, the agent should automatically scan for the next pending task.
```
When you want the agent to drive the project, use the following command:
> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."*
### Step 2: Reporting
The agent must report back with:
- Confirmation that the framework files were found and read.
- A summary of the project rules.
- The expected workflow for this project.
- Key observations from the project structure.
---
## Scenario A: Existing Project (Drop-In)
## The Lifecycle of a Project
### 1. Create the Project Framework Directory
```bash
mkdir -p /path/to/project/.agent-framework
```
Once onboarded, the project moves through these phases. The agent should use the provided prompts to transition between them.
### 2. Create the Two Required Files
### Phase 1: Research
**Template**: `prompts/research.md`
**Output**: `SPEC.md`
**Trigger**: *"Research {task-description}"*
#### AGENT.md (project level)
```markdown
# AGENT.md (project-name)
### Phase 2: Implementation
**Template**: `prompts/implement.md`
**Output**: Code changes + test results
**Trigger**: *"Implement the {task-name} task"*
This project uses the global framework at ~/.agent-framework/.
### Phase 3: Bug Finding
**Template**: `prompts/bug_finder.md`
**Output**: `BUG_REPORT.md`
**Trigger**: *"Find bugs in the {task-name} task"*
Additional project rules are in RULES.md.
### Phase 4: Adversarial Verification
**Template**: `prompts/adversarial_bug_find.md`
**Output**: `ADVERSARIAL_BUG_REPORT.md`
**Trigger**: *"Perform adversarial bug find for {task-name}"*
## Mode
Autopilot: Enabled.
```
### Phase 5: Referee
**Template**: `prompts/referee.md`
**Output**: `VERDICT.md`
**Trigger**: *"Review the {task-name} task"*
#### RULES.md (project level)
Start with any hard constraints you already know for this project.
## Prompt Rendering Convention
### 3. Run the Initial Exploration Ritual
Give the agent this prompt in a fresh session:
All prompts are stored as template files in `~/.agent-framework/prompts/`. They use `{placeholder}` syntax.
```
You have been given a new project at this path:
/path/to/project/
### Placeholders
- `{project}`: Absolute path to the project root.
- `{task-name}`: The task folder name (kebab-case).
- `{task-description}`: A brief, clear summary of the current work.
Your first actions must be:
1. Explore the project root using ls and find.
2. Read .agent-framework/AGENT.md
3. Read .agent-framework/RULES.md
4. Read ~/.agent-framework/AGENT.md
Report back with:
- Confirmation the framework files were found and read
- Summary of the project rules
- What process this project expects
- Key observations from the project structure
Do not start any task yet.
```
---
## Scenario B: Starting From Scratch
Follow the same steps as Scenario A, but start by asking the agent to design the initial architecture:
> *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"*
## Phase Prompts (Template Files)
| Phase | Template | Output | Trigger |
|---|---|---|---|
| **Research** | `research.md` | `SPEC.md` | *"Research {task-description}"* |
| **Implementation** | `implement.md` | Code + Tests | *"Implement the {task-name} task"* |
| **Bug Find** | `bug_finder.md` | `BUG_REPORT.md` | *"Find bugs in the {task-name} task"* |
| **Adversarial** | `adversarial_bug_find.md` | `ADVERSARIAL_BUG_REPORT.md` | *"Perform adversarial bug find for {task-name}"* |
| **Referee** | `referee.md` | `VERDICT.md` | *"Review the {task-name} task"* |
## Core Principles
- **Context Is Everything**: Separate research from implementation. Use fresh sessions per task.
- **Sycophancy Management**: Use the multi-agent validation loop (Bug Finder $\rightarrow$ Adversarial $\rightarrow$ Referee) to ensure objective results.
- **Clear End States**: Tests are mandatory. A task is not complete until all acceptance criteria in the `SPEC.md` (or `CONTRACT.md`) are met.
- **Rules + Skills**: Keep `RULES.md` specific to this project. Use the global framework for general behavior.
When the agent receives a trigger command, it must:
1. Read the corresponding template file.
2. Replace all `{placeholders}` with the actual project values.
3. Execute the rendered prompt.
+44 -10
View File
@@ -2,8 +2,48 @@
A contract-based operating system for LLM agents, designed to enforce disciplined engineering workflows.
## Core Philosophy
The framework prevents agents from "hallucinating" features or jumping into code without a plan. It enforces a strict separation between **Planning (Research)** and **Execution (Implementation)**, mediated by a **Verification** loop.
## 1. Global Framework Installation (One-time setup)
Before you can use the framework in any project, you must install the core logic into your local environment.
**Run these commands in your terminal:**
```bash
# Clone the framework into the global config directory
git clone [INSERT_FRAMEWORK_REPO_URL_HERE] ~/.agent-framework
# Enter the directory
cd ~/.agent-framework
# Make the installation script executable and run it
chmod +x install.sh
./install.sh
```
*Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory.*
---
## 2. Project Setup (Per project)
Once the framework is installed globally, you must "onboard" every individual project you work on.
### Option A: The Agent-Driven Way (Recommended)
If you want the agent to handle the configuration for you, navigate to your project root and run:
> *"Onboard this project into the agent-framework."*
The agent will automatically:
1. Detect your project type (New, Existing, or Upgrade).
2. Create the `./.agent-framework/` directory.
3. Generate your `AGENT.md` and `RULES.md` files.
4. Initiate the "Exploration Ritual" to understand your codebase.
### Option B: The Manual Way
If you prefer to set it up manually, create a `.agent-framework/` directory in your project root and add:
- `AGENT.md`: Project-specific configuration (Mode, rules, etc.).
- `RULES.md`: Project-specific constraints and past failure modes.
---
## The Autopilot Workflow
The framework features an **Autopilot** mode that allows the agent to drive a project to completion with minimal intervention.
@@ -20,17 +60,11 @@ Add `Autopilot: Enabled` to your project's `.agent-framework/AGENT.md` and use t
> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."*
The agent will then:
- Scan the `tasks/` directory.
- Identify the current phase of every task based on existing artifacts.
- Recommend the next command to progress each task.
- Automatically move to the next phase once the current one's artifacts are produced.
## Key Components
- `AGENT.md`: Project-specific configuration and mode selection.
- `RULES.md`: Living document of project constraints and past failure modes.
- `prompts/`: Specialized system prompts for each phase (Research, Implement, Bug Finder, etc.).
- `workflow.md`: The state machine governing the Autopilot lifecycle.
## Installation
Clone the repository and set up your project following the instructions in `ONBOARDING.md`.
## Contact & Support
[Insert Contact Info]
+17 -64
View File
@@ -1,92 +1,45 @@
You are the Bug Finder. Your job is to adversarially test the implementation and find every bug, deviation from spec, and edge case.
You are the Bug Finder. Your job is to find every bug, deviation from spec, and edge case.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
3. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists)
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
3. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
## Task
{task-description}
## Adversarial Checklist
## Checklist
### Spec Compliance
- Does the implementation match the SPEC.md exactly?
- Are there missing features or stubs?
- Are there features not in the spec (scope creep)?
### Edge Cases
- Empty inputs, null values, zero-length arrays
- Large inputs (performance, memory)
- Malformed data, unexpected types
- Concurrent access, race conditions
- Dependency failures (network, database, API)
### Security
- SQL injection, XSS, CSRF
- Authentication and authorization gaps
- Data exposure (logs, error messages, API responses)
- Rate limiting, input validation
- File upload, path traversal
### Data Flow
- Trace data from input to output
- Are mutations safe?
- Is sensitive data exposed?
- Can data be lost or corrupted?
### Concurrency & Race Conditions
- Shared state without synchronization
- Async operations without error handling
- Deadlocks, livelocks
- Transaction isolation issues
### Error Handling
- Are all errors caught and logged?
- Are silent failures possible?
- Are swallowed exceptions present?
- Is there graceful degradation?
### Performance
- O(n^2) or worse algorithms
- N+1 query patterns
- Memory leaks, unbounded caches
- Unbounded loops, infinite recursion
### Testing
- Are all edge cases covered by tests?
- Are tests actually testing the right things?
- Are there false positives (tests that pass but don't verify)?
- **Spec**: Does it match the SPEC.md exactly? Are there missing features or scope creep?
- **Edges**: Check nulls, empties, large inputs, malformed data, and concurrency.
- **Security**: Check for injection, auth gaps, data exposure, and input validation.
- **Performance**: Look for O(n^2)+, N+1 queries, memory leaks, and infinite loops.
- **Errors**: Are all errors caught, logged, and handled gracefully?
## Output Format
Produce a BUG_REPORT.md at {project}/tasks/{task-name}/BUG_REPORT.md with:
Produce a BUG_REPORT.md at {project}/tasks/{task-name}/BUG_REPORT.md:
```markdown
# Bug Report: {task-name}
## Summary
{Brief overview of findings}
{Brief overview}
## Bugs Found
### Bug 1: {Title}
- **Severity**: Critical / High / Medium / Low
- **Description**: {What's wrong}
- **Location**: {File:line}
- **Reproduction**: {Steps to reproduce}
- **Suggested Fix**: {How to fix}
### Bug 2: ...
- **Description**: {What's wrong}
- **Reproduction**: {Steps}
- **Suggested Fix**: {Fix}
## Score
{Assign a score: +1 for low, +5 for medium, +10 for critical}
```
## Important
- Be aggressive. Your job is to find bugs, not to be nice.
- If you find nothing, say so explicitly — but double-check everything first.
- Do NOT invent bugs. Only report real issues.
- Be aggressive. Do NOT invent bugs.
- If no bugs found, state it explicitly.
+14 -1
View File
@@ -11,6 +11,15 @@ Your job is to create a clear, actionable design for the project based on the sp
{task-description}
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/.agent-framework/RULES.md
## Task
{task-description}
## Output
Produce a file called `DESIGN.md` at {project}/tasks/{task-name}/DESIGN.md containing:
@@ -38,7 +47,11 @@ Produce a file called `DESIGN.md` at {project}/tasks/{task-name}/DESIGN.md conta
- Biggest technical or product risks
- Areas that need exploration or validation first
### 6. Non-Functional Requirements
### 6. Documentation Plan
- List all specific documentation that must be updated or created (e.g., README sections, API docs, docstrings).
- Define the "source of truth" for each piece of documentation.
### 7. Non-Functional Requirements
- Performance, security, reliability, or scale considerations (if relevant)
When the design is complete, output "CONTRACT_MET" and stop.
+16 -8
View File
@@ -2,25 +2,33 @@ You are in implementation mode.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md — what needs to be built
2. {project}/.agent-framework/RULES.md — project-specific rules
3. {project}/.agent-framework/AGENT.md — project agent config (if exists)
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/.agent-framework/RULES.md
3. {project}/.agent-framework/AGENT.md (if exists)
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
## Task
{task-description}
## Implementation Rules
## Implementation Rules (TDD Mode)
- Follow the SPEC.md exactly. Do not add features not listed.
- Write tests first when a correct seam exists.
- Follow the SPEC.md and DESIGN.md exactly.
- **Strict TDD Loop**: You are NOT allowed to write a feature in one go. You must follow the Red-Green-Refactor cycle:
1. **RED**: Write a failing test for the smallest possible unit of the feature.
2. **GREEN**: Write the minimum amount of code required to make that test pass.
3. **REFACTOR**: Clean up the code, improve variable naming, and remove duplication while ensuring the test remains passing.
- **Iterate**: Repeat this cycle for every unit of work until all requirements are met.
- Write tests first.
- Run tests and report full output.
- Keep functions small and focused.
- Use existing patterns in the codebase.
### End State
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
- All tests must pass.
- **Documentation**: All documentation identified in the DESIGN.md must be updated or created.
- Run the full test suite and report results
- Do NOT declare victory until tests pass
@@ -34,4 +42,4 @@ Report back with:
## Stop Condition (MANDATORY)
You are not allowed to end this session until you have produced all required deliverables AND output the exact phrase "CONTRACT_MET".
Until then, continue working or ask clarifying questions.
Until then, continue working or ask clarifying questions.
+13 -6
View File
@@ -2,11 +2,12 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
3. {project}/tasks/{task-name}/BUG_REPORT.md — bugs found by Bug Finder (if exists)
4. {project}/tasks/{task-name}/ADVERSARIAL_BUG_REPORT.md — bugs found by Adversarial Bug Finder (if exists)
5. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists)
1. {project}/tasks/{task-name}/SPEC.md
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
3. {project}/tasks/{task-name}/BUG_REPORT.md (if exists)
4. {project}/tasks/{task-name}/ADVERSARIAL_BUG_REPORT.md (if exists)
5. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
6. {project}/tasks/{task-name}/DESIGN.md (if exists)
## Task
@@ -41,6 +42,12 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
- Are edge cases covered?
- Are there false positives (tests that pass but don't verify)?
### Documentation Review (Grill with Docs)
- Did the agent update all documentation identified in the DESIGN.md?
- Is the documentation accurate and reflects the final implementation?
- Is the documentation clear enough for a developer to understand the new changes?
- Does the documentation cover any edge cases or non-obvious logic?
## Verdict
Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with:
@@ -77,4 +84,4 @@ Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with:
## Stop Condition (MANDATORY)
You are not allowed to end this session until you have produced the VERDICT.md file AND output the exact phrase "CONTRACT_MET".
Until then, continue working or ask clarifying questions.
Until then, continue working or ask clarifying questions.