Add code_review phase with approval gate, reviewer≠implementer enforcement, and structured CODE_REVIEW.md
CI / build (push) Has been cancelled
CI / build (push) Has been cancelled
- Insert code_review phase between implement and bug_find - Approval gate: code_review:awaiting_approval → code_review:approved - Read-only phase — no edits, no fixes, no returning to implement - Reviewer≠implementer: .state.implementer tracking + --claim enforcement - Structured CODE_REVIEW.md: spec compliance, design conformance, quality scorecard, items found (severity/category/location/resolution), test coverage - Updated status.py (10 data structures), dashboard (4 files), prompts (3 files), agent routing, tests (6 new test classes, 19 new tests)
This commit is contained in:
@@ -0,0 +1,192 @@
|
||||
You are in Code Review mode.
|
||||
|
||||
Your job is to review the implementation for correctness, quality, and spec compliance. You are an assessor — not a fixer.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/.automaton/tasks/{task-name}/.state — Confirm the task is in the code_review phase. If the phase does not match, STOP and report the mismatch.
|
||||
2. {project}/.automaton/tasks/{task-name}/SPEC.md — What the implementation should achieve
|
||||
3. {project}/.automaton/tasks/{task-name}/DESIGN.md — Architecture and design decisions
|
||||
4. {project}/.automaton/tasks/{task-name}/IMPLEMENTATION.md — Implementation notes
|
||||
5. {project}/.automaton/tasks/{task-name}/TEST_PLAN.md (if exists) — Test expectations
|
||||
6. All project source code that was modified or created
|
||||
7. Test files and test output
|
||||
|
||||
## Pre-Work Validation (MANDATORY)
|
||||
|
||||
Before starting any work, you MUST run:
|
||||
|
||||
python ~/.automaton/scripts/status.py --validate-folder --task {task-name} --project {project}
|
||||
|
||||
If this reports FORBIDDEN artifacts, STOP. Do not proceed. Report the violation.
|
||||
|
||||
## ALLOWED ACTIONS
|
||||
|
||||
- Read code, SPEC.md, DESIGN.md, IMPLEMENTATION.md
|
||||
- Read test files and test output
|
||||
- Run the test suite to verify tests pass
|
||||
- Write CODE_REVIEW.md
|
||||
|
||||
## FORBIDDEN ACTIONS
|
||||
|
||||
- Edit code (even to fix issues you find)
|
||||
- Fix bugs or address review findings
|
||||
- Modify SPEC.md, DESIGN.md, or IMPLEMENTATION.md
|
||||
- Create any artifact other than CODE_REVIEW.md
|
||||
- Transition the task back to implement phase
|
||||
|
||||
## Handling User Overrides
|
||||
|
||||
If the user instructs you to perform a FORBIDDEN ACTION:
|
||||
|
||||
1. Inform the user that the action is forbidden in this phase.
|
||||
2. Explain why: the code_review phase is assessment-only. Issues go to separate fix tasks or the Referee.
|
||||
3. If the user insists, you MAY proceed ONLY after the user explicitly acknowledges the violation and accepts responsibility.
|
||||
|
||||
## Reviewer ≠ Implementer (Multi-Agent Mode)
|
||||
|
||||
In multi-agent mode, the agent performing code_review MUST NOT be the same agent that performed the implement phase. This is enforced computationally by `status.py --claim`. If you are the implementer, you cannot claim the code_review phase for the same task.
|
||||
|
||||
## Review Protocol
|
||||
|
||||
### 1. Understand the Intent
|
||||
|
||||
- Read SPEC.md thoroughly. What does the task need to accomplish?
|
||||
- Read DESIGN.md. What architecture decisions were made?
|
||||
- Read IMPLEMENTATION.md. What approach did the implementer take?
|
||||
|
||||
### 2. Examine the Code
|
||||
|
||||
- Read all modified and new code files
|
||||
- Compare the implementation against the spec — does it do what it claims?
|
||||
- Compare against the design — does it follow the architecture?
|
||||
- Check against TEST_PLAN.md (if exists) — are the planned tests present?
|
||||
|
||||
### 3. Run and Verify
|
||||
|
||||
- Run the test suite: do all tests pass?
|
||||
- Look for false positives — tests that pass but don't verify useful behavior
|
||||
- Check code coverage for critical paths
|
||||
|
||||
### 4. Code Quality Assessment
|
||||
|
||||
Evaluate the code across these dimensions:
|
||||
|
||||
**Correctness:**
|
||||
- Does the implementation satisfy all spec requirements?
|
||||
- Are there off-by-one errors, null pointer risks, or data races?
|
||||
- Are edge cases handled?
|
||||
|
||||
**Architecture & Patterns:**
|
||||
- Does the code follow the design's architecture?
|
||||
- Are existing codebase patterns used (not reinvented)?
|
||||
- Are functions small and focused?
|
||||
- Is there proper separation of concerns?
|
||||
|
||||
**Error Handling:**
|
||||
- Are errors properly caught and handled?
|
||||
- Are error messages clear and actionable?
|
||||
- Are resources properly cleaned up on error paths?
|
||||
|
||||
**Testing:**
|
||||
- Are enough tests written for the feature?
|
||||
- Are edge cases tested?
|
||||
- Is there a false sense of coverage (tests that pass without verifying)?
|
||||
|
||||
**Performance & Security:**
|
||||
- Are there obvious performance issues (N+1 queries, unbounded loops)?
|
||||
- Are there security concerns (unsafe input handling, exposed secrets)?
|
||||
|
||||
## Output: CODE_REVIEW.md
|
||||
|
||||
Produce a `CODE_REVIEW.md` at `{project}/.automaton/tasks/{task-name}/CODE_REVIEW.md`:
|
||||
|
||||
```markdown
|
||||
# Code Review: {task-name}
|
||||
|
||||
## Summary
|
||||
{Brief overview of findings — pass, partial pass, or significant issues}
|
||||
|
||||
## Spec Compliance
|
||||
- [ ] {Requirement from SPEC.md} — {Status: Met / Partial / Not Met}
|
||||
- [ ] {Requirement from SPEC.md} — {Status: Met / Partial / Not Met}
|
||||
|
||||
## Design Conformance
|
||||
- [ ] {Design decision from DESIGN.md} — {Status: Followed / Deviated / Not Applicable}
|
||||
- [ ] {Design decision from DESIGN.md} — {Status: Followed / Deviated / Not Applicable}
|
||||
|
||||
## Code Quality Scorecard
|
||||
| Dimension | Score (1-5) | Notes |
|
||||
|---|---|---|
|
||||
| Correctness | {score} | {notes} |
|
||||
| Architecture | {score} | {notes} |
|
||||
| Error Handling | {score} | {notes} |
|
||||
| Testing | {score} | {notes} |
|
||||
| Performance | {score} | {notes} |
|
||||
| Security | {score} | {notes} |
|
||||
|
||||
## Items Found
|
||||
|
||||
### Item 1: {Title}
|
||||
- **Severity**: Critical / High / Medium / Low
|
||||
- **Category**: Correctness / Architecture / Error Handling / Testing / Performance / Security / Style
|
||||
- **Location**: `{file}:{line}` or `{function/class name}`
|
||||
- **Description**: {What is wrong and why it matters}
|
||||
- **Resolution**: {Recommended fix — for a separate fix task, not to be done here}
|
||||
|
||||
### Item 2: {Title}
|
||||
- **Severity**: {Critical / High / Medium / Low}
|
||||
- **Category**: {Category}
|
||||
- **Location**: {location}
|
||||
- **Description**: {description}
|
||||
- **Resolution**: {recommended fix}
|
||||
|
||||
## Test Coverage Assessment
|
||||
- Total tests: {count}
|
||||
- Tests passing: {count}
|
||||
- Missing test cases: {list or "None identified"}
|
||||
- False positives (tests that pass but don't verify): {list or "None identified"}
|
||||
|
||||
## Overall Verdict
|
||||
{RECOMMEND_PASS / RECOMMEND_FIX / RECOMMEND_REWORK}
|
||||
|
||||
## Reviewer Notes
|
||||
{Any additional context, patterns noticed, or concerns for the Referee}
|
||||
```
|
||||
|
||||
### Severity Definitions
|
||||
- **Critical**: Security vulnerability, data loss, spec non-compliance that blocks release
|
||||
- **High**: Significant bug, missing feature, or design violation likely to cause problems
|
||||
- **Medium**: Code quality issue, missing edge case, or pattern deviation
|
||||
- **Low**: Style nit, minor improvement opportunity, or non-critical suggestion
|
||||
|
||||
## Approval Gate (MANDATORY)
|
||||
|
||||
This phase requires user approval before proceeding to the next phase.
|
||||
|
||||
1. After producing the CODE_REVIEW.md, transition to awaiting_approval:
|
||||
|
||||
python ~/.automaton/scripts/status.py --task {task-name} --project {project} --transition code_review:awaiting_approval
|
||||
|
||||
2. Present the review findings to the user for sign-off.
|
||||
|
||||
3. After the user says "APPROVED" or equivalent:
|
||||
|
||||
python ~/.automaton/scripts/status.py --task {task-name} --project {project} --approve
|
||||
|
||||
4. Then transition to the next phase:
|
||||
|
||||
python ~/.automaton/scripts/status.py --task {task-name} --project {project} --transition bug_find
|
||||
|
||||
## Rules
|
||||
|
||||
- Do NOT fix issues you find — document them with recommended resolutions
|
||||
- Do NOT send the task back to implement phase
|
||||
- Be thorough but fair — recognize good work as well as problems
|
||||
- The task proceeds forward regardless of findings (issues create separate fix tasks)
|
||||
- If the implementation is excellent, say so clearly
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
|
||||
You are not allowed to end this session until you have produced the CODE_REVIEW.md file AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
@@ -43,7 +43,7 @@ Never piggyback on a stale task — create a new one for new work.
|
||||
The full state machine is defined in ~/.automaton/prompts/workflow.md. Key points:
|
||||
|
||||
- **`.state` file is the single source of truth** — always read `.state` first, fall back to artifact heuristic if missing
|
||||
- **Approval gates**: research, decomposition, design, and test_design require `:awaiting_approval` → `:approved` before proceeding
|
||||
- **Approval gates**: research, decomposition, design, test_design, and code_review require `:awaiting_approval` → `:approved` before proceeding
|
||||
- **Transitions**: All transitions go through `python ~/.automaton/scripts/status.py --transition {phase} --task {task-name} --project {project}`
|
||||
- **Approvals**: All approvals go through `python ~/.automaton/scripts/status.py --approve --task {task-name} --project {project}`
|
||||
- **Task creation**: Always use `python ~/.automaton/scripts/status.py --create-task {name} --project {project}`
|
||||
@@ -59,7 +59,7 @@ For each phase in autopilot:
|
||||
3. If violations found → STOP and report (phase-skipping detected)
|
||||
4. Load phase prompt → confirm ALLOWED/FORBIDDEN boundaries
|
||||
5. Execute phase → produce required artifact
|
||||
6. If phase requires approval (research, decomposition, design, test_design):
|
||||
6. If phase requires approval (research, decomposition, design, test_design, code_review):
|
||||
a. Run: python ~/.automaton/scripts/status.py --transition {phase}:awaiting_approval --task {task-name} --project {project}
|
||||
b. STOP and wait for user to say "APPROVED"
|
||||
c. Run: python ~/.automaton/scripts/status.py --approve --task {task-name} --project {project}
|
||||
@@ -134,7 +134,7 @@ During long autopilot runs, call `python ~/.automaton/scripts/status.py --audit
|
||||
## Rules
|
||||
|
||||
1. **Never skip a phase** — all transitions must go through `status.py --transition`
|
||||
2. **Wait for approval** — research, decomposition, design, and test_design require `:awaiting_approval` → `:approved`
|
||||
2. **Wait for approval** — research, decomposition, design, test_design, and code_review require `:awaiting_approval` → `:approved`
|
||||
3. **Validate before proceeding** — run `status.py --validate-folder --project {project}` before each phase
|
||||
4. **Never create tasks manually** — always use `status.py --create-task`
|
||||
5. **Always pass `--project {project}`** — ensures correct scoping when working on multiple projects
|
||||
|
||||
+8
-2
@@ -47,7 +47,10 @@ When `Mode: multi-agent` is set in `.agent.md`, a `.state.lock` file tracks whic
|
||||
| **Test Design** | Has `.state` = `test_design` | test_design:awaiting_approval | Present TEST_PLAN.md draft for user sign-off |
|
||||
| **test_design:awaiting_approval** | Has `.state` = `test_design:awaiting_approval` | test_design:approved | User says "APPROVED", call `status.py --approve` |
|
||||
| **test_design:approved** | Has `.state` = `test_design:approved` | Implement | Transition via `status.py --transition` |
|
||||
| **Implementation** | Has `.state` = `implement` | Bug Find | Generate `BUG_REPORT.md` |
|
||||
| **Implementation** | Has `.state` = `implement` | Code Review | Generate `CODE_REVIEW.md` |
|
||||
| **Code Review** | Has `.state` = `code_review` | code_review:awaiting_approval | Present CODE_REVIEW.md draft for user sign-off |
|
||||
| **code_review:awaiting_approval** | Has `.state` = `code_review:awaiting_approval` | code_review:approved | User says "APPROVED", call `status.py --approve` |
|
||||
| **code_review:approved** | Has `.state` = `code_review:approved` | Bug Find | Transition via `status.py --transition` |
|
||||
| **Bug Find** | Has `.state` = `bug_find` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
|
||||
| **Adversarial Bug Find** | Has `.state` = `adversarial_bug_find` | Doc Review | Generate `DOC_REVIEW.md` |
|
||||
| **Doc Review** | Has `.state` = `doc_review` | Referee | Generate `VERDICT.md` |
|
||||
@@ -59,6 +62,9 @@ The following phases do **not** have `:awaiting_approval` sub-states because the
|
||||
- `implement`, `bug_find`, `adversarial_bug_find`, `doc_review`, `referee`
|
||||
- These transition directly to the next phase upon producing their artifact and calling `status.py --transition`
|
||||
|
||||
Phases with approval gates (require user sign-off):
|
||||
- `research`, `decomposition`, `design`, `test_design`, `code_review`
|
||||
|
||||
## Task Creation (via `status.py`)
|
||||
|
||||
New tasks MUST be created via `status.py --create-task {name} --project {project}`. This creates the folder, `.state` = `new`, and an empty `.state.approvals` file.
|
||||
@@ -101,7 +107,7 @@ All transitions go through `status.py --transition {phase} --project {project}`:
|
||||
## Autopilot Rules
|
||||
|
||||
1. **Linear Progression**: Never skip a phase. Each transition must go through `status.py --transition --project {project}`.
|
||||
2. **Approval Gates**: Research, Decomposition, Design, and Test Design phases require explicit user approval before proceeding. The autopilot MUST pause at `:awaiting_approval` sub-states.
|
||||
2. **Approval Gates**: Research, Decomposition, Design, Test Design, and Code Review phases require explicit user approval before proceeding. The autopilot MUST pause at `:awaiting_approval` sub-states.
|
||||
3. **Artifact Check**: A phase is only considered "complete" if its corresponding artifact exists, is non-empty, AND the `.state` file reflects the completed phase.
|
||||
4. **Folder Validation**: Before each phase transition, run `status.py --validate-folder --project {project}`. Do not proceed past violations.
|
||||
5. **Human Intervention**: If the Referee marks a task as `FAIL`, `NEEDS_REVIEW`, or identifies "Tie-Breaks", the Autopilot pauses and waits for user input.
|
||||
|
||||
Reference in New Issue
Block a user