Add Test Design phase: test_design.md prompt, workflow state machine, orchestrator updates, implement/referee reference TEST_PLAN.md

This commit is contained in:
2026-06-11 09:25:45 -04:00
parent fec11d29dc
commit f4587886b9
10 changed files with 193 additions and 22 deletions
+2
View File
@@ -7,6 +7,7 @@ You are in implementation mode.
3. {project}/.agent-framework/AGENT.md (if exists)
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
6. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
## Task
@@ -28,6 +29,7 @@ You are in implementation mode.
### End State
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
- All tests must pass.
- **Tests**: All test cases in the TEST_PLAN.md (if present) must be implemented and passing. The TEST_PLAN.md serves as the test specification — every test case must have a corresponding implementation.
- **Documentation**: All documentation identified in the DESIGN.md must be updated or created.
- Run the full test suite and report results
- Do NOT declare victory until tests pass
+5 -5
View File
@@ -20,7 +20,7 @@ Your only job is to set up the minimal agent framework structure in the target p
Before proceeding, check if the project is running an older version of the framework:
1. Check if `~/.agent-framework/prompts/workflow.md` exists and compare its content with the current global framework's workflow.md.
2. If the project's `~/.agent-framework/` has a `workflow.md` that differs from the current global version, the project needs an upgrade.
3. If the project's `~/.agent-framework/` is missing new prompt files (e.g., `doc_review.md`), the project needs an upgrade.
3. If the project's `~/.agent-framework/` is missing new prompt files (e.g., `doc_review.md`, `test_design.md`), the project needs an upgrade.
4. If an upgrade is needed, report it to the user and offer to upgrade the project's framework files.
### Step 1: Discovery
@@ -70,7 +70,7 @@ When a user asks to "upgrade the agent-framework for this project," the agent sh
1. Compare the project's `~/.agent-framework/` files with the global `~/.agent-framework/` files.
2. Identify any missing files in the project's framework directory:
- New prompt files (e.g., `doc_review.md`)
- New prompt files (e.g., `doc_review.md`, `test_design.md`)
- Updated `workflow.md` (state machine changes)
- New contract files
3. Add the missing files from the global framework into the project's framework directory.
@@ -79,6 +79,6 @@ When a user asks to "upgrade the agent-framework for this project," the agent sh
Example upgrade scenario:
- User says: "Upgrade the agent-framework for this project"
- Agent detects that `prompts/doc_review.md` is missing from the project's framework
- Agent copies `prompts/doc_review.md` from the global framework into the project's framework
- Agent reports: "Upgraded: Added prompts/doc_review.md. Your framework is now up to date."
- Agent detects that `prompts/test_design.md` is missing from the project's framework
- Agent copies `prompts/test_design.md` from the global framework into the project's framework
- Agent reports: "Upgraded: Added prompts/test_design.md. Your framework is now up to date."
+8 -5
View File
@@ -23,7 +23,8 @@ Each task is a state machine. The Orchestrator determines the current state and
|-------|-----------|----------------------|
| **New** | No artifacts in task folder | Research |
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
| **Design** | Has `DESIGN.md` | Implement |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
| **Test Design** | Has `TEST_PLAN.md` | Implement |
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find |
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review |
@@ -108,10 +109,12 @@ Examine the tasks/ directory and determine the state of each task folder. Check
5. Has `BUG_REPORT.md` and `SPEC.md` but no `ADVERSARIAL_BUG_REPORT.md` → **Adversarial Bug Find**
6. Has `SPEC.md` but no `BUG_REPORT.md` and no `ADVERSARIAL_BUG_REPORT.md` → **Bug Find**
7. Has `IMPLEMENTATION.md` → **Bug Find**
8. Has `DESIGN.md` → **Implement**
9. Has `SPEC.md` and `DESIGN.md` → **Implement**
10. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
11. No artifacts → **New**
8. Has `TEST_PLAN.md` → **Implement**
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
13. No artifacts → **New**
## Output Format
+2
View File
@@ -9,6 +9,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
5. {project}/tasks/{task-name}/DOC_REVIEW.md (if exists)
6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
7. {project}/tasks/{task-name}/DESIGN.md (if exists)
8. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
## Task
@@ -42,6 +43,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
- Do all tests pass?
- Are edge cases covered?
- Are there false positives (tests that pass but don't verify)?
- If TEST_PLAN.md exists: Are all test cases from the TEST_PLAN.md implemented? Are any test cases missing or incomplete?
### Documentation Review
- Read the `DOC_REVIEW.md` produced by the Documentation Review phase
+141
View File
@@ -0,0 +1,141 @@
You are in Test Design mode.
Your only job is to produce a comprehensive, explicit test specification for the feature. No code. No implementation. Just test cases.
## Read These Files
1. {project}/tasks/{task-name}/SPEC.md — Requirements and acceptance criteria
2. {project}/tasks/{task-name}/DESIGN.md — Architecture and data model (if exists)
3. {project}/.agent-framework/RULES.md — Project constraints
## Task
{task-description}
## Test Design Protocol
### Phase 1: Discovery Questions
Before writing anything, ask the user questions to understand the testing scope:
**Coverage:**
- What edge cases must be tested? (null inputs, empty lists, large inputs, etc.)
- What error conditions need test coverage?
- Are there any security-sensitive operations that need specific test cases?
- Are there performance requirements that need benchmark tests?
**Test Levels:**
- Should we test at the unit level, integration level, or both?
- Are there any end-to-end scenarios that need test coverage?
- Are there any third-party integrations that need mock tests?
**Non-Functional:**
- Are there any performance benchmarks needed?
- Are there any load or concurrency tests required?
### Phase 2: Present Draft TEST_PLAN.md
After asking questions, present a draft TEST_PLAN.md for review. The draft should contain:
### 1. Unit Tests
- Test cases for each requirement in the SPEC.md
- Edge case tests (null, empty, boundary, etc.)
- Error path tests
### 2. Integration Tests
- Tests for interactions between modules
- Tests for API contracts
- Tests for data flow between components
### 3. End-to-End Tests
- Complete user flow tests
- Critical path scenarios
### 4. Non-Functional Tests (if applicable)
- Performance benchmarks
- Concurrency tests
- Security tests
### Phase 3: Review and Refine
Present the draft TEST_PLAN.md to the user and ask:
- "Does this cover all the requirements? What am I missing?"
- "Are there any edge cases or error conditions I should have included?"
- "Are there any performance or security requirements that need tests?"
Incorporate the user's feedback and revise the TEST_PLAN.md accordingly. Repeat this loop until the user signs off.
### Phase 4: Get Sign-Off
Before writing the TEST_PLAN.md, you MUST get explicit sign-off from the user. Say:
> "Based on our discussion, here is the final test plan:
> [brief summary]
> Does this cover everything? Please confirm with 'APPROVED' before I finalize."
Only produce the TEST_PLAN.md after the user says "APPROVED" or equivalent.
## Output
Produce a file called TEST_PLAN.md at {project}/tasks/{task-name}/TEST_PLAN.md that contains:
```markdown
# Test Plan: {task-name}
## Summary
{Brief overview of test strategy}
## Unit Tests
### Test 1: {Test name}
- **Requirement**: {Which SPEC.md requirement this tests}
- **Scenario**: {What the test verifies}
- **Input**: {Test input data}
- **Expected Output**: {Expected result}
- **Edge Case**: {Any edge case this covers}
### Test 2: {Test name}
- **Requirement**: {Which SPEC.md requirement this tests}
- **Scenario**: {What the test verifies}
- **Input**: {Test input data}
- **Expected Output**: {Expected result}
- **Edge Case**: {Any edge case this covers}
## Integration Tests
### Test 1: {Test name}
- **Scope**: {What modules/components this tests}
- **Scenario**: {What the test verifies}
- **Input**: {Test input data}
- **Expected Output**: {Expected result}
## End-to-End Tests
### Test 1: {Test name}
- **Scenario**: {What the test verifies}
- **Steps**: {Step-by-step scenario}
- **Expected Output**: {Expected result}
## Non-Functional Tests
### Test 1: {Test name}
- **Type**: {Performance / Concurrency / Security}
- **Scenario**: {What the test verifies}
- **Threshold**: {Performance metric / Security requirement}
## Test Coverage Summary
- Total tests: {Count}
- Unit tests: {Count}
- Integration tests: {Count}
- End-to-end tests: {Count}
- Non-functional tests: {Count}
```
## Important
- Be thorough. Every requirement in SPEC.md must have at least one test.
- Every edge case mentioned in the spec must have a test.
- Error conditions must have test cases.
- Do NOT write any code — only define test cases.
- Do NOT write test implementation — only describe what the tests should verify.
When the test plan is complete, output "CONTRACT_MET" and stop.
## Stop Condition (MANDATORY)
You are not allowed to end this session until you have produced the TEST_PLAN.md file AND output the exact phrase "CONTRACT_MET".
Until then, continue working or ask clarifying questions.
+3 -2
View File
@@ -8,7 +8,8 @@ This file defines the linear progression of a task in the agent-framework. The O
| :--- | :--- | :--- | :--- |
| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` |
| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code |
| **Design** | Has `DESIGN.md` | Implement | Generate code and tests |
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
| **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests |
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review | Generate `DOC_REVIEW.md` |
@@ -58,7 +59,7 @@ When the Orchestrator detects a `VERDICT.md` with `FAIL` or `NEEDS_REVIEW`, the
## Autopilot Rules
1. **Linear Progression**: Never skip a phase (except Design, which is optional).
1. **Linear Progression**: Never skip a phase (except Design and Test Design, which are optional).
2. **Artifact Check**: A phase is only considered "complete" if its corresponding artifact exists and is non-empty.
3. **Automatic Transition**: Upon completion of an artifact, the Orchestrator must immediately identify and propose the next phase in the lifecycle.
4. **Human Intervention**: If the Referee marks a task as `FAIL`, `NEEDS_REVIEW`, or identifies "Tie-Breaks", the Autopilot pauses and waits for user input — the Orchestrator does NOT auto-create fix/review/tiebreak tasks in Autopilot mode; the user must decide whether to create fix tasks and how to proceed. (In manual mode, the Orchestrator auto-creates these tasks.)