Add Test Design phase: test_design.md prompt, workflow state machine, orchestrator updates, implement/referee reference TEST_PLAN.md
This commit is contained in:
@@ -7,6 +7,7 @@ You are in implementation mode.
|
||||
3. {project}/.agent-framework/AGENT.md (if exists)
|
||||
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists)
|
||||
5. {project}/tasks/{task-name}/DESIGN.md (if exists)
|
||||
6. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
@@ -28,6 +29,7 @@ You are in implementation mode.
|
||||
### End State
|
||||
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
|
||||
- All tests must pass.
|
||||
- **Tests**: All test cases in the TEST_PLAN.md (if present) must be implemented and passing. The TEST_PLAN.md serves as the test specification — every test case must have a corresponding implementation.
|
||||
- **Documentation**: All documentation identified in the DESIGN.md must be updated or created.
|
||||
- Run the full test suite and report results
|
||||
- Do NOT declare victory until tests pass
|
||||
|
||||
@@ -20,7 +20,7 @@ Your only job is to set up the minimal agent framework structure in the target p
|
||||
Before proceeding, check if the project is running an older version of the framework:
|
||||
1. Check if `~/.agent-framework/prompts/workflow.md` exists and compare its content with the current global framework's workflow.md.
|
||||
2. If the project's `~/.agent-framework/` has a `workflow.md` that differs from the current global version, the project needs an upgrade.
|
||||
3. If the project's `~/.agent-framework/` is missing new prompt files (e.g., `doc_review.md`), the project needs an upgrade.
|
||||
3. If the project's `~/.agent-framework/` is missing new prompt files (e.g., `doc_review.md`, `test_design.md`), the project needs an upgrade.
|
||||
4. If an upgrade is needed, report it to the user and offer to upgrade the project's framework files.
|
||||
|
||||
### Step 1: Discovery
|
||||
@@ -70,7 +70,7 @@ When a user asks to "upgrade the agent-framework for this project," the agent sh
|
||||
|
||||
1. Compare the project's `~/.agent-framework/` files with the global `~/.agent-framework/` files.
|
||||
2. Identify any missing files in the project's framework directory:
|
||||
- New prompt files (e.g., `doc_review.md`)
|
||||
- New prompt files (e.g., `doc_review.md`, `test_design.md`)
|
||||
- Updated `workflow.md` (state machine changes)
|
||||
- New contract files
|
||||
3. Add the missing files from the global framework into the project's framework directory.
|
||||
@@ -79,6 +79,6 @@ When a user asks to "upgrade the agent-framework for this project," the agent sh
|
||||
|
||||
Example upgrade scenario:
|
||||
- User says: "Upgrade the agent-framework for this project"
|
||||
- Agent detects that `prompts/doc_review.md` is missing from the project's framework
|
||||
- Agent copies `prompts/doc_review.md` from the global framework into the project's framework
|
||||
- Agent reports: "Upgraded: Added prompts/doc_review.md. Your framework is now up to date."
|
||||
- Agent detects that `prompts/test_design.md` is missing from the project's framework
|
||||
- Agent copies `prompts/test_design.md` from the global framework into the project's framework
|
||||
- Agent reports: "Upgraded: Added prompts/test_design.md. Your framework is now up to date."
|
||||
@@ -23,7 +23,8 @@ Each task is a state machine. The Orchestrator determines the current state and
|
||||
|-------|-----------|----------------------|
|
||||
| **New** | No artifacts in task folder | Research |
|
||||
| **Research** | Has `SPEC.md` | Design (optional) or Implement |
|
||||
| **Design** | Has `DESIGN.md` | Implement |
|
||||
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement |
|
||||
| **Test Design** | Has `TEST_PLAN.md` | Implement |
|
||||
| **Implement** | Has `IMPLEMENTATION.md` | Bug Find |
|
||||
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find |
|
||||
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review |
|
||||
@@ -108,10 +109,12 @@ Examine the tasks/ directory and determine the state of each task folder. Check
|
||||
5. Has `BUG_REPORT.md` and `SPEC.md` but no `ADVERSARIAL_BUG_REPORT.md` → **Adversarial Bug Find**
|
||||
6. Has `SPEC.md` but no `BUG_REPORT.md` and no `ADVERSARIAL_BUG_REPORT.md` → **Bug Find**
|
||||
7. Has `IMPLEMENTATION.md` → **Bug Find**
|
||||
8. Has `DESIGN.md` → **Implement**
|
||||
9. Has `SPEC.md` and `DESIGN.md` → **Implement**
|
||||
10. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
|
||||
11. No artifacts → **New**
|
||||
8. Has `TEST_PLAN.md` → **Implement**
|
||||
9. Has `SPEC.md` and `TEST_PLAN.md` → **Implement**
|
||||
10. Has `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
|
||||
11. Has `SPEC.md` and `DESIGN.md` → **Test Design** (optional) or **Implement** (if user skips test design)
|
||||
12. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design)
|
||||
13. No artifacts → **New**
|
||||
|
||||
## Output Format
|
||||
|
||||
|
||||
@@ -9,6 +9,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
5. {project}/tasks/{task-name}/DOC_REVIEW.md (if exists)
|
||||
6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists)
|
||||
7. {project}/tasks/{task-name}/DESIGN.md (if exists)
|
||||
8. {project}/tasks/{task-name}/TEST_PLAN.md (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
@@ -42,6 +43,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
- Do all tests pass?
|
||||
- Are edge cases covered?
|
||||
- Are there false positives (tests that pass but don't verify)?
|
||||
- If TEST_PLAN.md exists: Are all test cases from the TEST_PLAN.md implemented? Are any test cases missing or incomplete?
|
||||
|
||||
### Documentation Review
|
||||
- Read the `DOC_REVIEW.md` produced by the Documentation Review phase
|
||||
|
||||
@@ -0,0 +1,141 @@
|
||||
You are in Test Design mode.
|
||||
|
||||
Your only job is to produce a comprehensive, explicit test specification for the feature. No code. No implementation. Just test cases.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md — Requirements and acceptance criteria
|
||||
2. {project}/tasks/{task-name}/DESIGN.md — Architecture and data model (if exists)
|
||||
3. {project}/.agent-framework/RULES.md — Project constraints
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Test Design Protocol
|
||||
|
||||
### Phase 1: Discovery Questions
|
||||
|
||||
Before writing anything, ask the user questions to understand the testing scope:
|
||||
|
||||
**Coverage:**
|
||||
- What edge cases must be tested? (null inputs, empty lists, large inputs, etc.)
|
||||
- What error conditions need test coverage?
|
||||
- Are there any security-sensitive operations that need specific test cases?
|
||||
- Are there performance requirements that need benchmark tests?
|
||||
|
||||
**Test Levels:**
|
||||
- Should we test at the unit level, integration level, or both?
|
||||
- Are there any end-to-end scenarios that need test coverage?
|
||||
- Are there any third-party integrations that need mock tests?
|
||||
|
||||
**Non-Functional:**
|
||||
- Are there any performance benchmarks needed?
|
||||
- Are there any load or concurrency tests required?
|
||||
|
||||
### Phase 2: Present Draft TEST_PLAN.md
|
||||
|
||||
After asking questions, present a draft TEST_PLAN.md for review. The draft should contain:
|
||||
|
||||
### 1. Unit Tests
|
||||
- Test cases for each requirement in the SPEC.md
|
||||
- Edge case tests (null, empty, boundary, etc.)
|
||||
- Error path tests
|
||||
|
||||
### 2. Integration Tests
|
||||
- Tests for interactions between modules
|
||||
- Tests for API contracts
|
||||
- Tests for data flow between components
|
||||
|
||||
### 3. End-to-End Tests
|
||||
- Complete user flow tests
|
||||
- Critical path scenarios
|
||||
|
||||
### 4. Non-Functional Tests (if applicable)
|
||||
- Performance benchmarks
|
||||
- Concurrency tests
|
||||
- Security tests
|
||||
|
||||
### Phase 3: Review and Refine
|
||||
|
||||
Present the draft TEST_PLAN.md to the user and ask:
|
||||
- "Does this cover all the requirements? What am I missing?"
|
||||
- "Are there any edge cases or error conditions I should have included?"
|
||||
- "Are there any performance or security requirements that need tests?"
|
||||
|
||||
Incorporate the user's feedback and revise the TEST_PLAN.md accordingly. Repeat this loop until the user signs off.
|
||||
|
||||
### Phase 4: Get Sign-Off
|
||||
|
||||
Before writing the TEST_PLAN.md, you MUST get explicit sign-off from the user. Say:
|
||||
|
||||
> "Based on our discussion, here is the final test plan:
|
||||
> [brief summary]
|
||||
> Does this cover everything? Please confirm with 'APPROVED' before I finalize."
|
||||
|
||||
Only produce the TEST_PLAN.md after the user says "APPROVED" or equivalent.
|
||||
|
||||
## Output
|
||||
|
||||
Produce a file called TEST_PLAN.md at {project}/tasks/{task-name}/TEST_PLAN.md that contains:
|
||||
|
||||
```markdown
|
||||
# Test Plan: {task-name}
|
||||
|
||||
## Summary
|
||||
{Brief overview of test strategy}
|
||||
|
||||
## Unit Tests
|
||||
### Test 1: {Test name}
|
||||
- **Requirement**: {Which SPEC.md requirement this tests}
|
||||
- **Scenario**: {What the test verifies}
|
||||
- **Input**: {Test input data}
|
||||
- **Expected Output**: {Expected result}
|
||||
- **Edge Case**: {Any edge case this covers}
|
||||
|
||||
### Test 2: {Test name}
|
||||
- **Requirement**: {Which SPEC.md requirement this tests}
|
||||
- **Scenario**: {What the test verifies}
|
||||
- **Input**: {Test input data}
|
||||
- **Expected Output**: {Expected result}
|
||||
- **Edge Case**: {Any edge case this covers}
|
||||
|
||||
## Integration Tests
|
||||
### Test 1: {Test name}
|
||||
- **Scope**: {What modules/components this tests}
|
||||
- **Scenario**: {What the test verifies}
|
||||
- **Input**: {Test input data}
|
||||
- **Expected Output**: {Expected result}
|
||||
|
||||
## End-to-End Tests
|
||||
### Test 1: {Test name}
|
||||
- **Scenario**: {What the test verifies}
|
||||
- **Steps**: {Step-by-step scenario}
|
||||
- **Expected Output**: {Expected result}
|
||||
|
||||
## Non-Functional Tests
|
||||
### Test 1: {Test name}
|
||||
- **Type**: {Performance / Concurrency / Security}
|
||||
- **Scenario**: {What the test verifies}
|
||||
- **Threshold**: {Performance metric / Security requirement}
|
||||
|
||||
## Test Coverage Summary
|
||||
- Total tests: {Count}
|
||||
- Unit tests: {Count}
|
||||
- Integration tests: {Count}
|
||||
- End-to-end tests: {Count}
|
||||
- Non-functional tests: {Count}
|
||||
```
|
||||
|
||||
## Important
|
||||
- Be thorough. Every requirement in SPEC.md must have at least one test.
|
||||
- Every edge case mentioned in the spec must have a test.
|
||||
- Error conditions must have test cases.
|
||||
- Do NOT write any code — only define test cases.
|
||||
- Do NOT write test implementation — only describe what the tests should verify.
|
||||
|
||||
When the test plan is complete, output "CONTRACT_MET" and stop.
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced the TEST_PLAN.md file AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
+3
-2
@@ -8,7 +8,8 @@ This file defines the linear progression of a task in the agent-framework. The O
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` |
|
||||
| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code |
|
||||
| **Design** | Has `DESIGN.md` | Implement | Generate code and tests |
|
||||
| **Design** | Has `DESIGN.md` | Test Design (optional) or Implement | Generate `TEST_PLAN.md` or code |
|
||||
| **Test Design** | Has `TEST_PLAN.md` | Implement | Generate code and tests |
|
||||
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |
|
||||
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
|
||||
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review | Generate `DOC_REVIEW.md` |
|
||||
@@ -58,7 +59,7 @@ When the Orchestrator detects a `VERDICT.md` with `FAIL` or `NEEDS_REVIEW`, the
|
||||
|
||||
## Autopilot Rules
|
||||
|
||||
1. **Linear Progression**: Never skip a phase (except Design, which is optional).
|
||||
1. **Linear Progression**: Never skip a phase (except Design and Test Design, which are optional).
|
||||
2. **Artifact Check**: A phase is only considered "complete" if its corresponding artifact exists and is non-empty.
|
||||
3. **Automatic Transition**: Upon completion of an artifact, the Orchestrator must immediately identify and propose the next phase in the lifecycle.
|
||||
4. **Human Intervention**: If the Referee marks a task as `FAIL`, `NEEDS_REVIEW`, or identifies "Tie-Breaks", the Autopilot pauses and waits for user input — the Orchestrator does NOT auto-create fix/review/tiebreak tasks in Autopilot mode; the user must decide whether to create fix tasks and how to proceed. (In manual mode, the Orchestrator auto-creates these tasks.)
|
||||
|
||||
Reference in New Issue
Block a user