Add Test Design phase: test_design.md prompt, workflow state machine, orchestrator updates, implement/referee reference TEST_PLAN.md

This commit is contained in:
2026-06-11 09:25:45 -04:00
parent fec11d29dc
commit f4587886b9
10 changed files with 193 additions and 22 deletions
+2 -1
View File
@@ -7,7 +7,8 @@ Autopilot: Enabled
IF task type = research → load prompts/research.md + RULES.md
IF task type = design → load prompts/design.md + SPEC.md
IF task type = implement → load prompts/implement.md + SPEC.md + DESIGN.md + CONTRACT.md
IF task type = test_design → load prompts/test_design.md + SPEC.md + DESIGN.md
IF task type = implement → load prompts/implement.md + SPEC.md + DESIGN.md + TEST_PLAN.md + CONTRACT.md
IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code
IF task type = adversarial_bug_find → load prompts/adversarial_bug_find.md + SPEC.md + code
IF task type = referee → load prompts/referee.md + SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md
+13 -2
View File
@@ -2,7 +2,7 @@
## Summary
A comprehensive review of the agent-framework project identified **12 bugs** (after removing the visualize*.html files, which eliminated 8 visualization-related bugs). Remaining issues range from critical (Orchestrator is not a real state machine) to low (unused placeholders).
A comprehensive review of the agent-framework project identified **13 bugs** (after removing the visualize*.html files, which eliminated 8 visualization-related bugs). Remaining issues range from critical (Orchestrator is not a real state machine) to low (unused placeholders).
---
@@ -96,6 +96,16 @@ A comprehensive review of the agent-framework project identified **12 bugs** (af
---
### Bug 14: No test design phase — tests written on the fly during implementation
- **Severity**: Medium
- **Location**: `prompts/implement.md`, `prompts/referee.md`, `prompts/workflow.md`
- **Description**: Tests are written during the implementation phase as part of the TDD loop. There is no formal Test Design phase before implementation starts. The spec says "acceptance criteria are testable" but doesn't require test cases to be defined upfront. This means:
- The implementer writes tests on the fly, so there's no test specification to verify against later
- The referee checks "Are edge cases covered?" and "Do all tests pass?" but doesn't have a test spec to compare against
- There's no way for the user to sign off on test coverage before implementation starts
- **Reproduction**: Start "Implement add user auth". The implementer writes tests during the TDD loop. No one has reviewed or approved the test coverage before code was written.
- **Suggested Fix**: **RESOLVED** — Added a Test Design phase (optional) between Design and Implement. Produces TEST_PLAN.md — an explicit test specification that the implementer follows. The referee checks that all test cases from the TEST_PLAN.md are implemented.
### Bug 11: Orchestrator contradiction — Autopilot mode task creation from verdicts
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Tasks from bug verdicts" section
@@ -129,4 +139,5 @@ A comprehensive review of the agent-framework project identified **12 bugs** (af
| 11 | Medium | +5 | **RESOLVED** — Split Orchestrator verdict task creation into Manual/Autopilot modes |
| 12 | Medium | +5 | **RESOLVED** — Added `update.sh` script, upgrade option in onboarding, update instructions in README |
| 13 | Low | +1 | **RESOLVED** — Autopilot is now Enabled by default (was Disabled) |
| **Total** | | **53** | |
| 14 | Medium | +5 | **RESOLVED** — Added Test Design phase (optional) between Design and Implement |
| **Total** | | **58** | |
+7
View File
@@ -44,10 +44,17 @@ Instead of memorizing trigger phrases for each phase, you can just say **"orches
**Trigger**: *"Design the {task-name} task"* (or just *"orchestrate"* in manual mode)
**Interaction**: Agent will grill you for design decisions, trade-offs, and constraints. Present draft for review. Get your sign-off before finalizing.
### Phase 1c: Test Design (Optional)
**Template**: `prompts/test_design.md`
**Output**: `TEST_PLAN.md`
**Trigger**: *"Design tests for the {task-name} task"* (or just *"orchestrate"* in manual mode)
**Interaction**: Agent will grill you for test coverage, edge cases, and test strategy. Present draft for review. Get your sign-off before finalizing.
### Phase 2: Implementation
**Template**: `prompts/implement.md`
**Output**: Code changes + test results
**Trigger**: *"Implement the {task-name} task"* (or just *"orchestrate"* in manual mode)
**Note**: The implementer follows the TEST_PLAN.md (if present) and implements code with tests using TDD.
### Phase 3: Bug Finding
**Template**: `prompts/bug_finder.md`
+10 -7
View File
@@ -67,11 +67,12 @@ The framework features an **Autopilot** mode that allows the agent to drive a pr
### Lifecycle of a Task
1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. **Interactive** — agent grills you for requirements and gets sign-off.
2. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off.
3. **Implement**: Write code and tests based *only* on the `SPEC.md` (and `DESIGN.md` if present).
4. **Bug Find**: Aggressive search for bugs and spec deviations.
5. **Adversarial Bug Find**: Deep search for complex logic errors, race conditions, and performance issues.
6. **Doc Review**: Review documentation against DESIGN.md plan and fix missing docs.
7. **Referee**: Objective evaluation of all bugs, docs, and the final verdict.
3. **Test Design** (optional): Produce a `TEST_PLAN.md` (Test specification). No code allowed. **Interactive** — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off.
4. **Implement**: Write code and tests based *only* on the `SPEC.md`, `DESIGN.md` (if present), and `TEST_PLAN.md` (if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows.
5. **Bug Find**: Aggressive search for bugs and spec deviations.
6. **Adversarial Bug Find**: Deep search for complex logic errors, race conditions, and performance issues.
7. **Doc Review**: Review documentation against DESIGN.md plan and fix missing docs.
8. **Referee**: Objective evaluation of all bugs, docs, and the final verdict.
### How to use Autopilot
@@ -84,7 +85,8 @@ The default mode is **Autopilot: Enabled**. The Orchestrator automatically drive
Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like:
- "Research add user authentication" — starts a new task
- "Design the add-user-auth task" — designs the architecture (optional)
- "Implement the add-user-auth task" — implements the task
- "Design tests for the add-user-auth task" — designs test cases (optional)
- "Implement the add-user-auth task" — implements the task with tests
- "Find bugs in the add-user-auth task" — finds bugs
- "Perform adversarial bug find for add-user-auth" — deep bug search
- "Review docs for the add-user-auth task" — reviews documentation
@@ -94,8 +96,9 @@ Set `Autopilot: Disabled` in your project's `.agent-framework/AGENT.md` if you p
## Key Components
- `AGENT.md`: Project-specific configuration and mode selection.
- `RULES.md`: Living document of project constraints and past failure modes.
- `prompts/`: Specialized system prompts for each phase (Research, Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, etc.).
- `prompts/`: Specialized system prompts for each phase (Research, Design, Test Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, etc.).
- `workflow.md`: The state machine governing the Autopilot lifecycle.
- `test_design.md`: Produces a TEST_PLAN.md — an explicit test specification before implementation.
## Contact & Support
[Insert Contact Info]
+2
View File
@@ -7,6 +7,7 @@ You are in implementation mode.
3. {project}/.agent-framework/AGENT.md (if exists)
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
6. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
## Task
@@ -28,6 +29,7 @@ You are in implementation mode.
### End State
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
- All tests must pass.
- **Tests**: All test cases in the TEST_PLAN.md (if present) must be implemented and passing. The TEST_PLAN.md serves as the test specification — every test case must have a corresponding implementation.
- **Documentation**: All documentation identified in the DESIGN.md must be updated or created.
- Run the full test suite and report results
- Do NOT declare victory until tests pass
+5 -5
View File
@@ -20,7 +20,7 @@ Your only job is to set up the minimal agent framework structure in the target p
Before proceeding, check if the project is running an older version of the framework:
1. Check if `~/.agent-framework/prompts/workflow.md` exists and compare its content with the current global framework's workflow.md.
2. If the project's `~/.agent-framework/` has a `workflow.md` that differs from the current global version, the project needs an upgrade.
3. If the project's `~/.agent-framework/` is missing new prompt files (e.g., `doc_review.md`), the project needs an upgrade.
3. If the project's `~/.agent-framework/` is missing new prompt files (e.g., `doc_review.md`, `test_design.md`), the project needs an upgrade.
4. If an upgrade is needed, report it to the user and offer to upgrade the project's framework files.
### Step 1: Discovery
@@ -70,7 +70,7 @@ When a user asks to "upgrade the agent-framework for this project," the agent sh
1. Compare the project's `~/.agent-framework/` files with the global `~/.agent-framework/` files.
2. Identify any missing files in the project's framework directory:
- New prompt files (e.g., `doc_review.md`)
- New prompt files (e.g., `doc_review.md`, `test_design.md`)
- Updated `workflow.md` (state machine changes)
- New contract files
3. Add the missing files from the global framework into the project's framework directory.
@@ -79,6 +79,6 @@ When a user asks to "upgrade the agent-framework for this project," the agent sh
Example upgrade scenario:
- User says: "Upgrade the agent-framework for this project"
- Agent detects that `prompts/doc_review.md` is missing from the project's framework
- Agent copies `prompts/doc_review.md` from the global framework into the project's framework
- Agent reports: "Upgraded: Added prompts/doc_review.md. Your framework is now up to date."
- Agent detects that `prompts/test_design.md` is missing from the project's framework
- Agent copies `prompts/test_design.md` from the global framework into the project's framework
- Agent reports: "Upgraded: Added prompts/test_design.md. Your framework is now up to date."
+8 -5
View File
@@ -23,7 +23,8 @@ Each task is a state machine. The Orchestrator determines the current state and
|-------|-----------|----------------------|
| **New** | No artifacts in task folder | Research |
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
| **Design** | Has `DESIGN.md` | Implement |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
| **Test Design** | Has `TEST_PLAN.md` | Implement |
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find |
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review |
@@ -108,10 +109,12 @@ Examine the tasks/ directory and determine the state of each task folder. Check
5. Has `BUG_REPORT.md` and `SPEC.md` but no `ADVERSARIAL_BUG_REPORT.md` → **Adversarial Bug Find**
6. Has `SPEC.md` but no `BUG_REPORT.md` and no `ADVERSARIAL_BUG_REPORT.md` → **Bug Find**
7. Has `IMPLEMENTATION.md` → **Bug Find**
8. Has `DESIGN.md` → **Implement**
9. Has `SPEC.md` and `DESIGN.md` → **Implement**
10. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
11. No artifacts → **New**
8. Has `TEST_PLAN.md` → **Implement**
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
13. No artifacts → **New**
## Output Format
+2
View File
@@ -9,6 +9,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
5. {project}/tasks/{task-name}/DOC_REVIEW.md (if exists)
6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
7. {project}/tasks/{task-name}/DESIGN.md (if exists)
8. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
## Task
@@ -42,6 +43,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
- Do all tests pass?
- Are edge cases covered?
- Are there false positives (tests that pass but don't verify)?
- If TEST_PLAN.md exists: Are all test cases from the TEST_PLAN.md implemented? Are any test cases missing or incomplete?
### Documentation Review
- Read the `DOC_REVIEW.md` produced by the Documentation Review phase
+141
View File
@@ -0,0 +1,141 @@
You are in Test Design mode.
Your only job is to produce a comprehensive, explicit test specification for the feature. No code. No implementation. Just test cases.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md — Requirements and acceptance criteria
2. {project}/tasks/{task-name}/DESIGN.md — Architecture and data model (if exists)
3. {project}/.agent-framework/RULES.md — Project constraints
## Task
{task-description}
## Test Design Protocol
### Phase 1: Discovery Questions
Before writing anything, ask the user questions to understand the testing scope:
**Coverage:**
- What edge cases must be tested? (null inputs, empty lists, large inputs, etc.)
- What error conditions need test coverage?
- Are there any security-sensitive operations that need specific test cases?
- Are there performance requirements that need benchmark tests?
**Test Levels:**
- Should we test at the unit level, integration level, or both?
- Are there any end-to-end scenarios that need test coverage?
- Are there any third-party integrations that need mock tests?
**Non-Functional:**
- Are there any performance benchmarks needed?
- Are there any load or concurrency tests required?
### Phase 2: Present Draft TEST_PLAN.md
After asking questions, present a draft TEST_PLAN.md for review. The draft should contain:
### 1. Unit Tests
- Test cases for each requirement in the SPEC.md
- Edge case tests (null, empty, boundary, etc.)
- Error path tests
### 2. Integration Tests
- Tests for interactions between modules
- Tests for API contracts
- Tests for data flow between components
### 3. End-to-End Tests
- Complete user flow tests
- Critical path scenarios
### 4. Non-Functional Tests (if applicable)
- Performance benchmarks
- Concurrency tests
- Security tests
### Phase 3: Review and Refine
Present the draft TEST_PLAN.md to the user and ask:
- "Does this cover all the requirements? What am I missing?"
- "Are there any edge cases or error conditions I should have included?"
- "Are there any performance or security requirements that need tests?"
Incorporate the user's feedback and revise the TEST_PLAN.md accordingly. Repeat this loop until the user signs off.
### Phase 4: Get Sign-Off
Before writing the TEST_PLAN.md, you MUST get explicit sign-off from the user. Say:
> "Based on our discussion, here is the final test plan:
> [brief summary]
> Does this cover everything? Please confirm with 'APPROVED' before I finalize."
Only produce the TEST_PLAN.md after the user says "APPROVED" or equivalent.
## Output
Produce a file called TEST_PLAN.md at {project}/tasks/{task-name}/TEST_PLAN.md that contains:
```markdown
# Test Plan: {task-name}
## Summary
{Brief overview of test strategy}
## Unit Tests
### Test 1: {Test name}
- **Requirement**: {Which SPEC.md requirement this tests}
- **Scenario**: {What the test verifies}
- **Input**: {Test input data}
- **Expected Output**: {Expected result}
- **Edge Case**: {Any edge case this covers}
### Test 2: {Test name}
- **Requirement**: {Which SPEC.md requirement this tests}
- **Scenario**: {What the test verifies}
- **Input**: {Test input data}
- **Expected Output**: {Expected result}
- **Edge Case**: {Any edge case this covers}
## Integration Tests
### Test 1: {Test name}
- **Scope**: {What modules/components this tests}
- **Scenario**: {What the test verifies}
- **Input**: {Test input data}
- **Expected Output**: {Expected result}
## End-to-End Tests
### Test 1: {Test name}
- **Scenario**: {What the test verifies}
- **Steps**: {Step-by-step scenario}
- **Expected Output**: {Expected result}
## Non-Functional Tests
### Test 1: {Test name}
- **Type**: {Performance / Concurrency / Security}
- **Scenario**: {What the test verifies}
- **Threshold**: {Performance metric / Security requirement}
## Test Coverage Summary
- Total tests: {Count}
- Unit tests: {Count}
- Integration tests: {Count}
- End-to-end tests: {Count}
- Non-functional tests: {Count}
```
## Important
- Be thorough. Every requirement in SPEC.md must have at least one test.
- Every edge case mentioned in the spec must have a test.
- Error conditions must have test cases.
- Do NOT write any code — only define test cases.
- Do NOT write test implementation — only describe what the tests should verify.
When the test plan is complete, output "CONTRACT_MET" and stop.
## Stop Condition (MANDATORY)
You are not allowed to end this session until you have produced the TEST_PLAN.md file AND output the exact phrase "CONTRACT_MET".
Until then, continue working or ask clarifying questions.
+3 -2
View File
@@ -8,7 +8,8 @@ This file defines the linear progression of a task in the agent-framework. The O
| :--- | :--- | :--- | :--- |
| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` |
| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code |
| **Design** | Has `DESIGN.md` | Implement | Generate code and tests |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
| **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests |
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review | Generate `DOC_REVIEW.md` |
@@ -58,7 +59,7 @@ When the Orchestrator detects a `VERDICT.md` with `FAIL` or `NEEDS_REVIEW`, the
## Autopilot Rules
1. **Linear Progression**: Never skip a phase (except Design, which is optional).
1. **Linear Progression**: Never skip a phase (except Design and Test Design, which are optional).
2. **Artifact Check**: A phase is only considered "complete" if its corresponding artifact exists and is non-empty.
3. **Automatic Transition**: Upon completion of an artifact, the Orchestrator must immediately identify and propose the next phase in the lifecycle.
4. **Human Intervention**: If the Referee marks a task as `FAIL`, `NEEDS_REVIEW`, or identifies "Tie-Breaks", the Autopilot pauses and waits for user input — the Orchestrator does NOT auto-create fix/review/tiebreak tasks in Autopilot mode; the user must decide whether to create fix tasks and how to proceed. (In manual mode, the Orchestrator auto-creates these tasks.)