Add code_review phase with approval gate, reviewer≠implementer enforcement, and structured CODE_REVIEW.md
CI / build (push) Has been cancelled

- Insert code_review phase between implement and bug_find
- Approval gate: code_review:awaiting_approval → code_review:approved
- Read-only phase — no edits, no fixes, no returning to implement
- Reviewer≠implementer: .state.implementer tracking + --claim enforcement
- Structured CODE_REVIEW.md: spec compliance, design conformance, quality
  scorecard, items found (severity/category/location/resolution), test coverage
- Updated status.py (10 data structures), dashboard (4 files), prompts (3 files),
  agent routing, tests (6 new test classes, 19 new tests)
This commit is contained in:
2026-06-16 09:00:22 -04:00
parent 42ccc7e2b7
commit 3d4c0926b4
12 changed files with 458 additions and 50 deletions
+192
View File
@@ -0,0 +1,192 @@
You are in Code Review mode.
Your job is to review the implementation for correctness, quality, and spec compliance. You are an assessor — not a fixer.
## Read These Files
1. {project}/.automaton/tasks/{task-name}/.state — Confirm the task is in the code_review phase. If the phase does not match, STOP and report the mismatch.
2. {project}/.automaton/tasks/{task-name}/SPEC.md — What the implementation should achieve
3. {project}/.automaton/tasks/{task-name}/DESIGN.md — Architecture and design decisions
4. {project}/.automaton/tasks/{task-name}/IMPLEMENTATION.md — Implementation notes
5. {project}/.automaton/tasks/{task-name}/TEST_PLAN.md (if exists) — Test expectations
6. All project source code that was modified or created
7. Test files and test output
## Pre-Work Validation (MANDATORY)
Before starting any work, you MUST run:
python ~/.automaton/scripts/status.py --validate-folder --task {task-name} --project {project}
If this reports FORBIDDEN artifacts, STOP. Do not proceed. Report the violation.
## ALLOWED ACTIONS
- Read code, SPEC.md, DESIGN.md, IMPLEMENTATION.md
- Read test files and test output
- Run the test suite to verify tests pass
- Write CODE_REVIEW.md
## FORBIDDEN ACTIONS
- Edit code (even to fix issues you find)
- Fix bugs or address review findings
- Modify SPEC.md, DESIGN.md, or IMPLEMENTATION.md
- Create any artifact other than CODE_REVIEW.md
- Transition the task back to implement phase
## Handling User Overrides
If the user instructs you to perform a FORBIDDEN ACTION:
1. Inform the user that the action is forbidden in this phase.
2. Explain why: the code_review phase is assessment-only. Issues go to separate fix tasks or the Referee.
3. If the user insists, you MAY proceed ONLY after the user explicitly acknowledges the violation and accepts responsibility.
## Reviewer ≠ Implementer (Multi-Agent Mode)
In multi-agent mode, the agent performing code_review MUST NOT be the same agent that performed the implement phase. This is enforced computationally by `status.py --claim`. If you are the implementer, you cannot claim the code_review phase for the same task.
## Review Protocol
### 1. Understand the Intent
- Read SPEC.md thoroughly. What does the task need to accomplish?
- Read DESIGN.md. What architecture decisions were made?
- Read IMPLEMENTATION.md. What approach did the implementer take?
### 2. Examine the Code
- Read all modified and new code files
- Compare the implementation against the spec — does it do what it claims?
- Compare against the design — does it follow the architecture?
- Check against TEST_PLAN.md (if exists) — are the planned tests present?
### 3. Run and Verify
- Run the test suite: do all tests pass?
- Look for false positives — tests that pass but don't verify useful behavior
- Check code coverage for critical paths
### 4. Code Quality Assessment
Evaluate the code across these dimensions:
**Correctness:**
- Does the implementation satisfy all spec requirements?
- Are there off-by-one errors, null pointer risks, or data races?
- Are edge cases handled?
**Architecture & Patterns:**
- Does the code follow the design's architecture?
- Are existing codebase patterns used (not reinvented)?
- Are functions small and focused?
- Is there proper separation of concerns?
**Error Handling:**
- Are errors properly caught and handled?
- Are error messages clear and actionable?
- Are resources properly cleaned up on error paths?
**Testing:**
- Are enough tests written for the feature?
- Are edge cases tested?
- Is there a false sense of coverage (tests that pass without verifying)?
**Performance & Security:**
- Are there obvious performance issues (N+1 queries, unbounded loops)?
- Are there security concerns (unsafe input handling, exposed secrets)?
## Output: CODE_REVIEW.md
Produce a `CODE_REVIEW.md` at `{project}/.automaton/tasks/{task-name}/CODE_REVIEW.md`:
```markdown
# Code Review: {task-name}
## Summary
{Brief overview of findings — pass, partial pass, or significant issues}
## Spec Compliance
- [ ] {Requirement from SPEC.md} — {Status: Met / Partial / Not Met}
- [ ] {Requirement from SPEC.md} — {Status: Met / Partial / Not Met}
## Design Conformance
- [ ] {Design decision from DESIGN.md} — {Status: Followed / Deviated / Not Applicable}
- [ ] {Design decision from DESIGN.md} — {Status: Followed / Deviated / Not Applicable}
## Code Quality Scorecard
| Dimension | Score (1-5) | Notes |
|---|---|---|
| Correctness | {score} | {notes} |
| Architecture | {score} | {notes} |
| Error Handling | {score} | {notes} |
| Testing | {score} | {notes} |
| Performance | {score} | {notes} |
| Security | {score} | {notes} |
## Items Found
### Item 1: {Title}
- **Severity**: Critical / High / Medium / Low
- **Category**: Correctness / Architecture / Error Handling / Testing / Performance / Security / Style
- **Location**: `{file}:{line}` or `{function/class name}`
- **Description**: {What is wrong and why it matters}
- **Resolution**: {Recommended fix — for a separate fix task, not to be done here}
### Item 2: {Title}
- **Severity**: {Critical / High / Medium / Low}
- **Category**: {Category}
- **Location**: {location}
- **Description**: {description}
- **Resolution**: {recommended fix}
## Test Coverage Assessment
- Total tests: {count}
- Tests passing: {count}
- Missing test cases: {list or "None identified"}
- False positives (tests that pass but don't verify): {list or "None identified"}
## Overall Verdict
{RECOMMEND_PASS / RECOMMEND_FIX / RECOMMEND_REWORK}
## Reviewer Notes
{Any additional context, patterns noticed, or concerns for the Referee}
```
### Severity Definitions
- **Critical**: Security vulnerability, data loss, spec non-compliance that blocks release
- **High**: Significant bug, missing feature, or design violation likely to cause problems
- **Medium**: Code quality issue, missing edge case, or pattern deviation
- **Low**: Style nit, minor improvement opportunity, or non-critical suggestion
## Approval Gate (MANDATORY)
This phase requires user approval before proceeding to the next phase.
1. After producing the CODE_REVIEW.md, transition to awaiting_approval:
python ~/.automaton/scripts/status.py --task {task-name} --project {project} --transition code_review:awaiting_approval
2. Present the review findings to the user for sign-off.
3. After the user says "APPROVED" or equivalent:
python ~/.automaton/scripts/status.py --task {task-name} --project {project} --approve
4. Then transition to the next phase:
python ~/.automaton/scripts/status.py --task {task-name} --project {project} --transition bug_find
## Rules
- Do NOT fix issues you find — document them with recommended resolutions
- Do NOT send the task back to implement phase
- Be thorough but fair — recognize good work as well as problems
- The task proceeds forward regardless of findings (issues create separate fix tasks)
- If the implementation is excellent, say so clearly
## Stop Condition (MANDATORY)
You are not allowed to end this session until you have produced the CODE_REVIEW.md file AND output the exact phrase "CONTRACT_MET".
Until then, continue working or ask clarifying questions.
+3 -3
View File
@@ -43,7 +43,7 @@ Never piggyback on a stale task — create a new one for new work.
The full state machine is defined in ~/.automaton/prompts/workflow.md. Key points:
- **`.state` file is the single source of truth** — always read `.state` first, fall back to artifact heuristic if missing
- **Approval gates**: research, decomposition, design, and test_design require `:awaiting_approval` → `:approved` before proceeding
- **Approval gates**: research, decomposition, design, test_design, and code_review require `:awaiting_approval` → `:approved` before proceeding
- **Transitions**: All transitions go through `python ~/.automaton/scripts/status.py --transition {phase} --task {task-name} --project {project}`
- **Approvals**: All approvals go through `python ~/.automaton/scripts/status.py --approve --task {task-name} --project {project}`
- **Task creation**: Always use `python ~/.automaton/scripts/status.py --create-task {name} --project {project}`
@@ -59,7 +59,7 @@ For each phase in autopilot:
3. If violations found → STOP and report (phase-skipping detected)
4. Load phase prompt → confirm ALLOWED/FORBIDDEN boundaries
5. Execute phase → produce required artifact
6. If phase requires approval (research, decomposition, design, test_design):
6. If phase requires approval (research, decomposition, design, test_design, code_review):
a. Run: python ~/.automaton/scripts/status.py --transition {phase}:awaiting_approval --task {task-name} --project {project}
b. STOP and wait for user to say "APPROVED"
c. Run: python ~/.automaton/scripts/status.py --approve --task {task-name} --project {project}
@@ -134,7 +134,7 @@ During long autopilot runs, call `python ~/.automaton/scripts/status.py --audit
## Rules
1. **Never skip a phase** — all transitions must go through `status.py --transition`
2. **Wait for approval** — research, decomposition, design, and test_design require `:awaiting_approval` → `:approved`
2. **Wait for approval** — research, decomposition, design, test_design, and code_review require `:awaiting_approval` → `:approved`
3. **Validate before proceeding** — run `status.py --validate-folder --project {project}` before each phase
4. **Never create tasks manually** — always use `status.py --create-task`
5. **Always pass `--project {project}`** — ensures correct scoping when working on multiple projects
+8 -2
View File
@@ -47,7 +47,10 @@ When `Mode: multi-agent` is set in `.agent.md`, a `.state.lock` file tracks whic
| **Test Design** | Has `.state` = `test_design` | test_design:awaiting_approval | Present TEST_PLAN.md draft for user sign-off |
| **test_design:awaiting_approval** | Has `.state` = `test_design:awaiting_approval` | test_design:approved | User says "APPROVED", call `status.py --approve` |
| **test_design:approved** | Has `.state` = `test_design:approved` | Implement | Transition via `status.py --transition` |
| **Implementation** | Has `.state` = `implement` | Bug Find | Generate `BUG_REPORT.md` |
| **Implementation** | Has `.state` = `implement` | Code Review | Generate `CODE_REVIEW.md` |
| **Code Review** | Has `.state` = `code_review` | code_review:awaiting_approval | Present CODE_REVIEW.md draft for user sign-off |
| **code_review:awaiting_approval** | Has `.state` = `code_review:awaiting_approval` | code_review:approved | User says "APPROVED", call `status.py --approve` |
| **code_review:approved** | Has `.state` = `code_review:approved` | Bug Find | Transition via `status.py --transition` |
| **Bug Find** | Has `.state` = `bug_find` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
| **Adversarial Bug Find** | Has `.state` = `adversarial_bug_find` | Doc Review | Generate `DOC_REVIEW.md` |
| **Doc Review** | Has `.state` = `doc_review` | Referee | Generate `VERDICT.md` |
@@ -59,6 +62,9 @@ The following phases do **not** have `:awaiting_approval` sub-states because the
- `implement`, `bug_find`, `adversarial_bug_find`, `doc_review`, `referee`
- These transition directly to the next phase upon producing their artifact and calling `status.py --transition`
Phases with approval gates (require user sign-off):
- `research`, `decomposition`, `design`, `test_design`, `code_review`
## Task Creation (via `status.py`)
New tasks MUST be created via `status.py --create-task {name} --project {project}`. This creates the folder, `.state` = `new`, and an empty `.state.approvals` file.
@@ -101,7 +107,7 @@ All transitions go through `status.py --transition {phase} --project {project}`:
## Autopilot Rules
1. **Linear Progression**: Never skip a phase. Each transition must go through `status.py --transition --project {project}`.
2. **Approval Gates**: Research, Decomposition, Design, and Test Design phases require explicit user approval before proceeding. The autopilot MUST pause at `:awaiting_approval` sub-states.
2. **Approval Gates**: Research, Decomposition, Design, Test Design, and Code Review phases require explicit user approval before proceeding. The autopilot MUST pause at `:awaiting_approval` sub-states.
3. **Artifact Check**: A phase is only considered "complete" if its corresponding artifact exists, is non-empty, AND the `.state` file reflects the completed phase.
4. **Folder Validation**: Before each phase transition, run `status.py --validate-folder --project {project}`. Do not proceed past violations.
5. **Human Intervention**: If the Referee marks a task as `FAIL`, `NEEDS_REVIEW`, or identifies "Tie-Breaks", the Autopilot pauses and waits for user input.