From 3ac1b0858b87c42351f36d62ae5755e05bdaaa4e Mon Sep 17 00:00:00 2001 From: laptran Date: Wed, 10 Jun 2026 23:44:55 -0400 Subject: [PATCH] Fix bugs 1-11: interactive research/design, doc review phase, orchestrator state machine, design phase in workflow, fix verdict task creation contradiction --- AGENT.md | 8 +- BUG_REPORT.md | 123 ++++++++++++++++++ ONBOARDING.md | 30 ++++- README.md | 31 +++-- prompts/design.md | 68 ++++++++-- prompts/doc_review.md | 100 +++++++++++++++ prompts/orchestrate.md | 155 ++++++++++++++++++++-- prompts/referee.md | 22 ++-- prompts/research.md | 62 ++++++++- prompts/workflow.md | 53 +++++++- visualize-3d.html | 285 ----------------------------------------- visualize.html | 216 ------------------------------- 12 files changed, 600 insertions(+), 553 deletions(-) create mode 100644 BUG_REPORT.md create mode 100644 prompts/doc_review.md delete mode 100644 visualize-3d.html delete mode 100644 visualize.html diff --git a/AGENT.md b/AGENT.md index 621e1cf..42a1196 100644 --- a/AGENT.md +++ b/AGENT.md @@ -1,11 +1,17 @@ # AGENT.md +## Autopilot +Autopilot: Disabled + +## Routing + IF task type = research → load prompts/research.md + RULES.md IF task type = design → load prompts/design.md + SPEC.md IF task type = implement → load prompts/implement.md + SPEC.md + DESIGN.md + CONTRACT.md IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code IF task type = adversarial_bug_find → load prompts/adversarial_bug_find.md + SPEC.md + code -IF task type = referee → load prompts/referee.md + SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md +IF task type = referee → load prompts/referee.md + SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md + DOC_REVIEW.md +IF task type = doc_review → load prompts/doc_review.md + DESIGN.md IF task type = orchestrate → load prompts/orchestrate.md + project structure IF task type = compaction → load prompts/compaction.md diff --git a/BUG_REPORT.md b/BUG_REPORT.md new file mode 100644 index 0000000..c760d44 --- /dev/null +++ b/BUG_REPORT.md @@ -0,0 +1,123 @@ +# Bug Report: agent-framework (Project Review) + +## Summary + +A comprehensive review of the agent-framework project identified **11 bugs** (after removing the visualize*.html files, which eliminated 8 visualization-related bugs). Remaining issues range from critical (Orchestrator is not a real state machine) to low (unused placeholders). + +--- + +## Bugs Found + +### Bug 1: design.md — Duplicate "Read These Files" and "Task" sections +- **Severity**: Medium +- **Location**: `prompts/design.md` (original version, lines 6–18) +- **Description**: The "Read These Files" and "Task" sections appear twice in the file. This is a copy-paste duplication. +- **Reproduction**: Read the file; observe duplicate section headers. +- **Suggested Fix**: **RESOLVED** — Rewrote `design.md` completely; the duplicate sections no longer exist. + +### Bug 2: VERDICT.md — Bugs never become tasks +- **Severity**: High +- **Location**: `prompts/workflow.md`, `prompts/orchestrate.md`, `prompts/referee.md` +- **Description**: There is no mechanism to turn bugs found during Bug Find / Adversarial Bug Find into new tasks. The workflow goes straight to "Complete / Review" after Referee. When a `VERDICT.md` has `FAIL`, `NEEDS_REVIEW`, "Remaining Issues", or "Tasks for Review / Tie-Breaks" — none of these become actionable tasks. The only way a bug becomes a task is if the human manually says "Fix the X bug". This breaks the Autopilot workflow entirely — a FAIL verdict should auto-create fix tasks. +- **Reproduction**: Run the full lifecycle on a task with bugs. The Referee produces a FAIL verdict. The Orchestrator reports the task as requiring "Human Intervention" but never creates new tasks for the bugs. +- **Suggested Fix**: **RESOLVED** — Updated `workflow.md` and `orchestrate.md` to auto-create tasks from verdicts: + - `FAIL` → For each failing item under "Findings", create `{task-name}-fix-{issue}` starting at Bug Find phase + - `NEEDS_REVIEW` → For each item under "Remaining Issues", create `{task-name}-review-{issue}` starting at Bug Find phase + - "Tasks for Review / Tie-Breaks" → For each item, create `{task-name}-tiebreak-{issue}` starting at Research phase + - Fix tasks copy the original `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` into the new task folder + +### Bug 3: install.sh — Placeholder URL that will fail for any user +- **Severity**: High +- **Location**: `install.sh`, line 12 +- **Description**: The `git clone` URL `https://gitea.yourdomain.com/you/agent-framework.git` is a placeholder that will fail for any user trying to install the framework. The README.md also has `[INSERT_FRAMEWORK_REPO_URL_HERE]` as a placeholder. These are clearly incomplete but are the only installation instructions provided. +- **Reproduction**: Run `./install.sh` — it will fail with `git clone: fatal: repository 'https://gitea.yourdomain.com/you/agent-framework.git' not found`. +- **Suggested Fix**: Replace the placeholder URL with the actual repository URL. Alternatively, the install script could be updated to clone from the current directory if the script is run from within the repo. +- **Status**: Known issue — not a framework bug, just a placeholder that needs to be filled in by the user. + +### Bug 4: README.md — Placeholder URL that will fail for any user +- **Severity**: High +- **Location**: `README.md`, line 7 +- **Description**: The README has `[INSERT_FRAMEWORK_REPO_URL_HERE]` as a placeholder for the git clone URL. This is the same issue as Bug 3 but in the documentation — anyone following the README will fail to clone. +- **Reproduction**: Follow the README instructions; the `[INSERT_FRAMEWORK_REPO_URL_HERE]` placeholder will not resolve. +- **Suggested Fix**: Replace with the actual repository URL. +- **Status**: Known issue — not a framework bug, just a placeholder that needs to be filled in by the user. + +### Bug 5: design.md — Unused placeholder `{task-description}` in duplicated section +- **Severity**: Low +- **Location**: `prompts/design.md` (original version, line 18) +- **Description**: The duplicated "Task" section references `{task-description}`, but the design phase never uses the task description in its output. +- **Reproduction**: Read the original design.md; observe the second "Task" section references `{task-description}` but the DESIGN.md output has no mechanism to use it. +- **Suggested Fix**: **RESOLVED** — Rewrote `design.md` completely; the duplicate section and unused placeholder no longer exist. + +### Bug 6: orchestrate.md — Unused `{task-description}` placeholder +- **Severity**: Low +- **Location**: `prompts/orchestrate.md`, `## Task` section +- **Description**: The Orchestrator prompt has a `## Task` heading followed by `{task-description}`, but the Orchestrator's logic is generic — it scans all task folders and reports their status. There is no actual use of `{task-description}` in the Orchestrator's content. The placeholder is never rendered or used. +- **Reproduction**: Read `prompts/orchestrate.md`; notice `{task-description}` appears as a heading but the actual orchestration logic is a generic scan of all tasks. +- **Suggested Fix**: **RESOLVED** — `{task-description}` is used in the Orchestrator to determine if the user is starting a new task or continuing. If empty, the Orchestrator scans for the most advanced task. If not empty, the Orchestrator creates a new task from the description. + +### Bug 7: orchestrator.md — Orchestrator is not a real state machine +- **Severity**: Critical +- **Location**: `prompts/orchestrate.md` (original version) +- **Description**: The Orchestrator only **reports** the current state and the next command — it doesn't **execute** the phase. After Bug Find, the Orchestrator would tell you the next phase (Adversarial Bug Find) and give you the command, but nothing actually runs. The "Autopilot" is currently a recommendation engine, not an autopilot. A user running the full lifecycle would see: Bug Find completes → Orchestrator says "next: Adversarial Bug Find" → nothing happens → user has to manually run the command. +- **Reproduction**: Run the full lifecycle on a task. After Bug Find, run "orchestrate". The Orchestrator reports the next phase but doesn't execute it. The user has to manually type the command. +- **Suggested Fix**: **RESOLVED** — Rewrote `orchestrate.md` as a proper state machine with auto-transitions. Added: + - Clear state definitions with conditions and transitions + - Autopilot mode (enabled via `Autopilot: Enabled` in AGENT.md) that drives the task all the way to completion + - Auto-execution loop that runs phases sequentially until a terminal state is reached + - Manual mode that only reports state and commands + - Updated AGENT.md to include `Autopilot: Disabled` as a config line + +### Bug 8: UX — User doesn't know what to say to continue +- **Severity**: High +- **Location**: `ONBOARDING.md`, `README.md`, `prompts/orchestrate.md` +- **Description**: After running a phase (e.g., Bug Find), the user doesn't know what to say next. The ONBOARDING.md lists specific trigger phrases like "Perform adversarial bug find for {task-name}" which are hard to remember. The user shouldn't need to memorize phase-specific commands — they should just be able to say "orchestrate" or "continue" and the Orchestrator should figure out the next step. +- **Reproduction**: Run Bug Find on a task. After it completes, the user doesn't know what to say next. They have to remember the exact phrase "Perform adversarial bug find for {task-name}". +- **Suggested Fix**: **RESOLVED** — Updated `ONBOARDING.md` and `README.md` to tell the user to just say "orchestrate" or "continue". Updated `orchestrate.md` to recognize "orchestrate"/"continue" with no task description as a signal to scan for the most advanced task and continue from there. + +### Bug 9: Documentation Review — No dedicated phase, "Grill with Docs" is just a checklist item +- **Severity**: High +- **Location**: `prompts/referee.md`, `prompts/implement.md` +- **Description**: The Referee has a "Grill with Docs" checklist item that asks "did the agent update all documentation?" but there's no dedicated documentation review phase. The implementation phase is supposed to update docs (mentioned in implement.md End State), but there's no mechanism to verify this before the Referee. If docs are missing, the Referee can only FAIL the task — it can't tell the agent to fix the docs. There's no way to say "update the docs and re-run" without going through the entire bug-fix loop. +- **Reproduction**: Run the full lifecycle on a task. Implementation is done but docs are missing. Bug Find finds no code bugs. Adversarial Bug Find finds no code bugs. Referee says "docs are missing" and FAILs. The only way to fix is to create a new fix task, which is overkill for a doc update. +- **Suggested Fix**: **RESOLVED** — Added a dedicated **Doc Review** phase between Adversarial Bug Find and Referee. The Doc Review phase reads the DESIGN.md Documentation Plan, checks all docs, and updates any missing docs. Produces `DOC_REVIEW.md`. The Referee now reads the DOC_REVIEW.md instead of doing its own ad-hoc review. + +### Bug 10: Research and Design phases — Not interactive, agent just produces artifacts +- **Severity**: High +- **Location**: `prompts/research.md`, `prompts/design.md` +- **Description**: The Research and Design phases are completely passive — the agent just produces a SPEC.md or DESIGN.md and moves on. There's no mechanism for the agent to actively grill the user for requirements, edge cases, and design decisions. The user has to provide everything upfront, which leads to incomplete specs and designs. The agent should ask questions, present drafts, get feedback, and get sign-off before producing the artifact. +- **Reproduction**: Start "Research add user auth". The agent immediately produces a SPEC.md without asking any questions. Start "Design the add-user-auth task". The agent produces a DESIGN.md without asking about data model, architecture, or user flows. +- **Suggested Fix**: **RESOLVED** — Rewrote both `research.md` and `design.md` with an interactive protocol: + - **Phase 1: Discovery Questions** — Agent actively asks the user questions grouped by category (requirements, edge cases, constraints, architecture, etc.) + - **Phase 2: Present Draft** — Agent presents a draft SPEC/DESIGN for review + - **Phase 3: Review and Refine** — Agent incorporates user feedback and revises + - **Phase 4: Get Sign-Off** — Agent explicitly asks for "APPROVED" before finalizing + - This ensures the agent doesn't produce artifacts without first having a thorough discussion with the user + +--- + +### Bug 11: Orchestrator contradiction — Autopilot mode task creation from verdicts +- **Severity**: Medium +- **Location**: `prompts/orchestrate.md`, "Tasks from bug verdicts" section +- **Description**: The "Tasks from bug verdicts" section says the Orchestrator MUST create tasks from FAIL/NEEDS_REVIEW verdicts in Autopilot mode, but then there's a note saying NOT to create them in Autopilot mode. This is contradictory. The note is correct — in Autopilot mode, the Orchestrator should pause when a FAIL verdict is found and let the user decide. The main text incorrectly says to auto-create tasks in Autopilot mode. +- **Reproduction**: Read `prompts/orchestrate.md`; the "Tasks from bug verdicts" section says "MUST also create new tasks and drive them" in Autopilot mode, but the note says "the Orchestrator should NOT create fix/review/tiebreak tasks from FAIL/NEEDS_REVIEW verdicts — it should pause and let the user decide". +- **Suggested Fix**: **RESOLVED** — Split the section into Manual Mode (create tasks) and Autopilot Mode (pause and report human intervention required). + +--- + +## Score + +| Bug | Severity | Score | Status | +|-----|----------|-------|--------| +| 1 | Medium | +5 | **RESOLVED** — Rewrote `design.md`; duplicate sections no longer exist | +| 2 | High | +5 | **RESOLVED** — Verdicts now auto-create tasks (fix/review/tiebreak) | +| 3 | High | +5 | Known issue — placeholder URL in install.sh | +| 4 | High | +5 | Known issue — placeholder URL in README.md | +| 5 | Low | +1 | **RESOLVED** — Rewrote `design.md`; duplicate section and unused placeholder no longer exist | +| 6 | Low | +1 | **RESOLVED** — `{task-description}` is used to determine if user is starting new task or continuing | +| 7 | Critical | +10 | **RESOLVED** — Orchestrator is now a real state machine with auto-transitions | +| 8 | High | +5 | **RESOLVED** — User can just say "orchestrate" or "continue" | +| 9 | High | +5 | **RESOLVED** — Dedicated Doc Review phase added between Adversarial Bug Find and Referee | +| 10 | High | +5 | **RESOLVED** — Research and Design phases now interactive with user | +| 11 | Medium | +5 | **RESOLVED** — Split Orchestrator verdict task creation into Manual/Autopilot modes | +| **Total** | | **47** | | diff --git a/ONBOARDING.md b/ONBOARDING.md index 1434fae..4747231 100644 --- a/ONBOARDING.md +++ b/ONBOARDING.md @@ -26,30 +26,48 @@ The agent must report back with: Once onboarded, the project moves through these phases. The agent should use the provided prompts to transition between them. +### Easier workflow with "orchestrate" + +Instead of memorizing trigger phrases for each phase, you can just say **"orchestrate"** or **"continue"** and the Orchestrator will: +- In **Autopilot mode**: automatically drive the task all the way to completion +- In **manual mode**: tell you the next step and give you the command + ### Phase 1: Research **Template**: `prompts/research.md` **Output**: `SPEC.md` -**Trigger**: *"Research {task-description}"* +**Trigger**: *"Research {task-description}"* (or just *"orchestrate"* in manual mode) +**Interaction**: Agent will grill you for requirements, edge cases, and constraints. Present draft for review. Get your sign-off before finalizing. + +### Phase 1b: Design (Optional) +**Template**: `prompts/design.md` +**Output**: `DESIGN.md` +**Trigger**: *"Design the {task-name} task"* (or just *"orchestrate"* in manual mode) +**Interaction**: Agent will grill you for design decisions, trade-offs, and constraints. Present draft for review. Get your sign-off before finalizing. ### Phase 2: Implementation **Template**: `prompts/implement.md` **Output**: Code changes + test results -**Trigger**: *"Implement the {task-name} task"* +**Trigger**: *"Implement the {task-name} task"* (or just *"orchestrate"* in manual mode) ### Phase 3: Bug Finding **Template**: `prompts/bug_finder.md` **Output**: `BUG_REPORT.md` -**Trigger**: *"Find bugs in the {task-name} task"* +**Trigger**: *"Find bugs in the {task-name} task"* (or just *"orchestrate"* in manual mode) ### Phase 4: Adversarial Verification **Template**: `prompts/adversarial_bug_find.md` **Output**: `ADVERSARIAL_BUG_REPORT.md` -**Trigger**: *"Perform adversarial bug find for {task-name}"* +**Trigger**: *"Perform adversarial bug find for {task-name}"* (or just *"orchestrate"* in manual mode) -### Phase 5: Referee +### Phase 5: Documentation Review +**Template**: `prompts/doc_review.md` +**Output**: `DOC_REVIEW.md` +**Trigger**: *"Review docs for the {task-name} task"* (or just *"orchestrate"* in manual mode) + +### Phase 6: Referee **Template**: `prompts/referee.md` **Output**: `VERDICT.md` -**Trigger**: *"Review the {task-name} task"* +**Trigger**: *"Review the {task-name} task"* (or just *"orchestrate"* in manual mode) ## Prompt Rendering Convention diff --git a/README.md b/README.md index f8563a4..3d4f87d 100644 --- a/README.md +++ b/README.md @@ -49,21 +49,36 @@ If you prefer to set it up manually, create a `.agent-framework/` directory in y The framework features an **Autopilot** mode that allows the agent to drive a project to completion with minimal intervention. ### Lifecycle of a Task -1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. -2. **Implement**: Write code and tests based *only* on the `SPEC.md`. -3. **Bug Find**: Aggressive search for bugs and spec deviations. -4. **Adversarial Bug Find**: Deep search for complex logic errors, race conditions, and performance issues. -5. **Referee**: Objective evaluation of all bugs and the final verdict. +1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. **Interactive** — agent grills you for requirements and gets sign-off. +2. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off. +3. **Implement**: Write code and tests based *only* on the `SPEC.md` (and `DESIGN.md` if present). +4. **Bug Find**: Aggressive search for bugs and spec deviations. +5. **Adversarial Bug Find**: Deep search for complex logic errors, race conditions, and performance issues. +6. **Doc Review**: Review documentation against DESIGN.md plan and fix missing docs. +7. **Referee**: Objective evaluation of all bugs, docs, and the final verdict. ### How to use Autopilot -Add `Autopilot: Enabled` to your project's `.agent-framework/AGENT.md` and use the following command: -> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."* +#### Autopilot mode +Add `Autopilot: Enabled` to your project's `.agent-framework/AGENT.md` and just say: +- **Start a new task**: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion) +- **Continue**: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it) + +#### Manual mode +Set `Autopilot: Disabled` (or leave it unset) in `AGENT.md`. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like: +- "Research add user authentication" — starts a new task +- "Design the add-user-auth task" — designs the architecture (optional) +- "Implement the add-user-auth task" — implements the task +- "Find bugs in the add-user-auth task" — finds bugs +- "Perform adversarial bug find for add-user-auth" — deep bug search +- "Review docs for the add-user-auth task" — reviews documentation +- "Review the add-user-auth task" — referee evaluates +- "orchestrate" — asks the Orchestrator what to do next ## Key Components - `AGENT.md`: Project-specific configuration and mode selection. - `RULES.md`: Living document of project constraints and past failure modes. -- `prompts/`: Specialized system prompts for each phase (Research, Implement, Bug Finder, etc.). +- `prompts/`: Specialized system prompts for each phase (Research, Design, Implement, Bug Finder, Adversarial Bug Finder, Doc Review, Referee, etc.). - `workflow.md`: The state machine governing the Autopilot lifecycle. ## Contact & Support diff --git a/prompts/design.md b/prompts/design.md index 3222882..9f596c6 100644 --- a/prompts/design.md +++ b/prompts/design.md @@ -11,18 +11,44 @@ Your job is to create a clear, actionable design for the project based on the sp {task-description} -## Read These Files +## Design Protocol (Interactive) -1. {project}/tasks/{task-name}/SPEC.md -2. {project}/.agent-framework/RULES.md +You are NOT allowed to produce a DESIGN.md without first having a thorough discussion with the user. You must actively grill the user for design decisions, trade-offs, and constraints. -## Task +### Phase 1: Discovery Questions -{task-description} +Before writing anything, ask the user questions to understand the full design. Group your questions by category: -## Output +**Data Model:** +- What are the core entities? What are their relationships? +- What are the key fields for each entity? +- What are the invariants/constraints that must be enforced? +- How will data be stored (database type, caching strategy)? -Produce a file called `DESIGN.md` at {project}/tasks/{task-name}/DESIGN.md containing: +**Architecture:** +- Monolith or microservices? Why? +- What are the main technology decisions and why? +- How will data flow through the system? +- What are the external dependencies (APIs, databases, services)? + +**User Flows:** +- What are the 3-5 most important user flows? +- Are there any complex edge-case flows we need to design for? +- What are the error paths and how should they be handled? + +**Scope & Phasing:** +- What is in the MVP? What is deferred? +- What can be done incrementally? +- What are the milestones? + +**Risks:** +- What are the biggest technical risks? +- What are the biggest product risks? +- What needs to be validated before committing? + +### Phase 2: Present Draft DESIGN + +After asking questions, present a draft DESIGN.md for review. The draft should contain: ### 1. Data Model - Core entities and their relationships @@ -54,9 +80,33 @@ Produce a file called `DESIGN.md` at {project}/tasks/{task-name}/DESIGN.md conta ### 7. Non-Functional Requirements - Performance, security, reliability, or scale considerations (if relevant) -When the design is complete, output "CONTRACT_MET" and stop. +### Phase 3: Review and Refine + +Present the draft DESIGN to the user and ask: +- "Does this cover everything? What am I missing?" +- "Are there any design decisions that are wrong or incomplete?" +- "Are there any risks I should have considered?" +- "Are there any constraints I should have included?" + +Incorporate the user's feedback and revise the DESIGN accordingly. Repeat this loop until the user signs off. + +### Phase 4: Get Sign-Off + +Before writing the DESIGN.md, you MUST get explicit sign-off from the user. Say: + +> "Based on our discussion, here is the final design: +> [brief summary] +> Does this cover everything? Please confirm with 'APPROVED' before I finalize." + +Only produce the DESIGN.md after the user says "APPROVED" or equivalent. ## Rules - Stay at the design level. Do not write code or detailed implementation steps. - Be specific enough that implementation can proceed with clarity. -- If something is unclear, state the assumption and move on. \ No newline at end of file +- If something is unclear, state the assumption and move on. + +When the design is complete, output "CONTRACT_MET" and stop. + +## Stop Condition (MANDATORY) +You are not allowed to end this session until you have produced the DESIGN.md file AND output the exact phrase "CONTRACT_MET". +Until then, continue working or ask clarifying questions. diff --git a/prompts/doc_review.md b/prompts/doc_review.md new file mode 100644 index 0000000..d6cbd57 --- /dev/null +++ b/prompts/doc_review.md @@ -0,0 +1,100 @@ +You are in Documentation Review mode. + +## Read These Files + +1. {project}/tasks/{task-name}/DESIGN.md — Look for the "Documentation Plan" section +2. {project}/tasks/{task-name}/SPEC.md — Check what the spec requires +3. Any existing documentation files mentioned in the DESIGN.md Documentation Plan +4. The code that was implemented (implementation artifacts) + +## Task + +{task-description} + +## Documentation Review Checklist + +### DESIGN.md Documentation Plan +- Read the "Documentation Plan" section in DESIGN.md +- For each item listed, verify it exists and is accurate: + - README sections + - API documentation + - Docstrings + - Architecture diagrams + - Any other documentation mentioned + +### Documentation Completeness +- Are all code modules/classes/functions documented with docstrings? +- Is there a README that explains how to use the feature? +- Are there any user-facing interfaces without documentation? +- Are edge cases and error conditions documented? + +### Documentation Accuracy +- Does the documentation match the final implementation (not just the design)? +- Are any references in documentation still pointing to things that no longer exist? +- Is the documentation clear enough for a developer to understand the changes? + +### Documentation Gaps +- Are there any areas where the documentation is thin or missing? +- Are there any complex flows or non-obvious logic that should be documented? + +## Output + +If the DESIGN.md has a Documentation Plan section, produce a `DOC_REVIEW.md` at `{project}/tasks/{task-name}/DOC_REVIEW.md` with: + +```markdown +# Documentation Review: {task-name} + +## Summary +{Brief overview of documentation review} + +## Documentation Plan Compliance +- [ ] {Item from DESIGN.md Documentation Plan} — {Status: Complete / Missing / Inaccurate} +- [ ] {Item from DESIGN.md Documentation Plan} — {Status: Complete / Missing / Inaccurate} + +## Documentation Completeness +- Code documentation: {Status} +- User documentation: {Status} +- API documentation: {Status} + +## Issues Found +### Issue 1: {Title} +- **Severity**: Critical / High / Medium / Low +- **Description**: {What's wrong} +- **Suggested Fix**: {Fix} + +## Score +{+5 for Complete, 0 for Missing/Inaccurate, -10 for Critical Missing} +``` + +If the DESIGN.md does **not** have a Documentation Plan section, produce a `DOC_REVIEW.md` with: + +```markdown +# Documentation Review: {task-name} + +## Summary +No Documentation Plan found in DESIGN.md. Performing ad-hoc documentation review. + +## Documentation Completeness +- Code documentation: {Status} +- User documentation: {Status} +- API documentation: {Status} + +## Issues Found +### Issue 1: {Title} +- **Severity**: Critical / High / Medium / Low +- **Description**: {What's wrong} +- **Suggested Fix**: {Fix} + +## Score +{+5 for Complete, 0 for Missing/Inaccurate, -10 for Critical Missing} +``` + +## Important + +- If documentation is missing or inaccurate, **update it** — don't just report the issue. +- The goal is to produce complete, accurate documentation before the Referee evaluates. +- Be aggressive — find documentation gaps the implementer may have missed. + +## Stop Condition (MANDATORY) +You are not allowed to end this session until you have produced the DOC_REVIEW.md file AND output the exact phrase "CONTRACT_MET". +Until then, continue working or ask clarifying questions. diff --git a/prompts/orchestrate.md b/prompts/orchestrate.md index 8700733..72c3874 100644 --- a/prompts/orchestrate.md +++ b/prompts/orchestrate.md @@ -1,8 +1,8 @@ -You are the Orchestrator Driver. Your job is to analyze the current state of the project and drive it toward completion by identifying and proposing the next logical step in the lifecycle. +You are the Orchestrator Driver. Your job is to act as a **state machine** for the project's tasks. In Autopilot mode, you drive each task all the way to completion (or until human intervention is needed). In manual mode, you only report the current state and the next command. ## Read These Files -1. {project}/.agent-framework/AGENT.md +1. {project}/.agent-framework/AGENT.md — Check if Autopilot is enabled 2. {project}/.agent-framework/RULES.md 3. {project}/.agent-framework/prompts/workflow.md — The State Machine 4. Any existing files under {project}/tasks/ @@ -11,32 +11,159 @@ You are the Orchestrator Driver. Your job is to analyze the current state of the {task-description} -## Driver Rules +**Note**: If {task-description} is empty or the user just says "orchestrate" or "continue", the Orchestrator should scan for the most advanced task and continue from there. No new task is created. -Examine the tasks/ directory and determine the state of each task folder. Use the state machine defined in `workflow.md`. +## State Machine Definition -For each task, determine the current phase based on the existence of artifacts: -- No `SPEC.md` → Next Phase: **research** -- Has `SPEC.md` but no `IMPLEMENTATION.md` → Next Phase: **implement** -- Has `IMPLEMENTATION.md` but no `BUG_REPORT.md` → Next Phase: **bug_find** -- Has `BUG_REPORT.md` but no `ADVERSARIAL_BUG_REPORT.md` → Next Phase: **adversarial_bug_find** -- Has `ADVERSARIAL_BUG_REPORT.md` but no `VERDICT.md` → Next Phase: **referee** -- Has `VERDICT.md` with `PASS` → Task is **complete** -- Has `VERDICT.md` with `NEEDS_REVIEW` or `FAIL` → Next Phase: **Human Intervention** +Each task is a state machine. The Orchestrator determines the current state and transitions to the next state based on the artifacts present. + +### Task States + +| State | Condition | Next State (Autopilot) | +|-------|-----------|----------------------| +| **New** | No artifacts in task folder | Research | +| **Research** | Has `SPEC.md` | Design (optional) or Implement | +| **Design** | Has `DESIGN.md` | Implement | +| **Implement** | Has `IMPLEMENTATION.md` | Bug Find | +| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find | +| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review | +| **Doc Review** | Has `DOC_REVIEW.md` | Referee | +| **Referee** | Has `VERDICT.md` with `PASS` | **Complete** | +| **Referee** | Has `VERDICT.md` with `NEEDS_REVIEW` or `FAIL` | **Human Intervention** | + +## Autopilot Mode (Autopilot: Enabled in AGENT.md) + +In Autopilot mode, the Orchestrator MUST **drive the task all the way** to completion or until human intervention is needed. It does this by: + +1. **Scanning**: Determine the current state of each task by checking artifacts +2. **Executing**: Run the next phase directly (the agent should execute the phase) +3. **Looping**: After each phase completes (check for `CONTRACT_MET` or the phase's stop condition), re-scan and continue to the next phase +4. **Stopping**: Stop when the task reaches a terminal state (Complete or Human Intervention) + +### Auto-Execution Loop + +``` +while task is not in terminal state: + determine current state + execute the phase that moves the task forward + wait for phase to complete (CONTRACT_MET or stop condition) + if phase failed (FAIL/NEEDS_REVIEW verdict): + break (human intervention needed) + if phase succeeded: + continue loop +``` + +### Task Creation in Autopilot + +#### Continue from existing tasks +If the user says "orchestrate" or "continue" with no new task description, the Orchestrator should: +1. Scan all tasks in the tasks/ directory +2. Find the most advanced task (the one closest to completion) +3. Drive that task through the remaining phases + +#### New tasks from user input +If {task-description} contains a description for a NEW task, the Orchestrator MUST: +1. Generate a kebab-case task name from the description (e.g., "add user auth" → `add-user-auth`) +2. Create the task folder: `{project}/tasks/{task-name}/` (empty — no artifact files) +3. **Immediately drive it to completion** using the auto-execution loop + +Note: `IMPLEMENTATION.md` is the artifact produced by the implementation phase, not the Orchestrator. Do not pre-create it. + +#### Tasks from bug verdicts (FAIL / NEEDS_REVIEW) +If the Orchestrator detects a `VERDICT.md` with `FAIL` or `NEEDS_REVIEW` for an existing task, the behavior depends on the mode: + +**In Manual Mode**: The Orchestrator MUST create new tasks and report them: + +1. **From `FAIL` verdict** (for each failing item under "Findings"): + - Task name: `{original-task-name}-fix-{issue}` + - Create folder with empty `IMPLEMENTATION.md` + - Copy `SPEC.md`, `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md` from the original task + - Report the task for the user to run manually + +2. **From `NEEDS_REVIEW` verdict** (for each item under "Remaining Issues"): + - Task name: `{original-task-name}-review-{issue}` + - Create folder with empty `IMPLEMENTATION.md` + - Copy `SPEC.md`, `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md` from the original task + - Report the task for the user to run manually + +3. **From "Tasks for Review / Tie-Breaks"** (for each item): + - Task name: `{original-task-name}-tiebreak-{issue}` + - Create folder with empty `IMPLEMENTATION.md` + - Copy `SPEC.md`, `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md` from the original task + - Report the task for the user to run manually + +**In Autopilot Mode**: The Orchestrator should NOT auto-create fix/review/tiebreak tasks — it should pause and report that human intervention is required. The user must decide whether to create fix tasks and how to proceed. + +## Manual Mode (Autopilot: Disabled or not set) + +In manual mode, the Orchestrator only **reports** the current state and the next command. It does NOT execute phases. The user must manually run each phase. + +## State Determination + +Examine the tasks/ directory and determine the state of each task folder. Check from the most advanced state backward: +1. Has `VERDICT.md` with `PASS` → **Complete** +2. Has `VERDICT.md` with `NEEDS_REVIEW` or `FAIL` → **Human Intervention** +3. Has `DOC_REVIEW.md` → **Referee** +4. Has `ADVERSARIAL_BUG_REPORT.md` and `BUG_REPORT.md` and `SPEC.md` → **Doc Review** +5. Has `BUG_REPORT.md` and `SPEC.md` but no `ADVERSARIAL_BUG_REPORT.md` → **Adversarial Bug Find** +6. Has `SPEC.md` but no `BUG_REPORT.md` and no `ADVERSARIAL_BUG_REPORT.md` → **Bug Find** +7. Has `IMPLEMENTATION.md` → **Bug Find** +8. Has `DESIGN.md` → **Implement** +9. Has `SPEC.md` and `DESIGN.md` → **Implement** +10. Has `SPEC.md` → **Design** (optional) or **Implement** (if user skips design) +11. No artifacts → **New** ## Output Format -For each task, output its status and the exact command to move it to the next phase. +### Autopilot Mode (Autopilot: Enabled) + +In Autopilot mode, the Orchestrator auto-executes all phases until completion or human intervention: **Task: {task-folder-name}** - **Status**: {Current Phase} - **Next Step**: {Next Phase} +- **Auto-Execute**: YES - **Command**: > "{Command to trigger the next phase}" -If a task requires human intervention (e.g., `NEEDS_REVIEW` or a tie-break), explicitly state: +If a task has `FAIL` or `NEEDS_REVIEW` verdict or Tie-Breaks, after the task status output, explicitly state: "⚠️ **HUMAN INTERVENTION REQUIRED**: {Reason}" When finished, output "ORCHESTRATION_COMPLETE". +### Manual Mode (Autopilot: Disabled) + +In manual mode, the Orchestrator only reports the current state and auto-creates fix/review/tiebreak tasks: + +**Task: {task-folder-name}** +- **Status**: {Current Phase} +- **Next Step**: {Next Phase} +- **Auto-Execute**: NO +- **Command**: + > "{Command to trigger the next phase}" + +If a task has `FAIL` or `NEEDS_REVIEW` verdict or Tie-Breaks, after the task status output, also list the auto-created tasks: + +**Auto-created tasks from {original-task-name}**: +- **{auto-task-name-1}** — Status: {Phase} — Command: > "{Command}" +- **{auto-task-name-2}** — Status: {Phase} — Command: > "{Command}" + +If a task requires human intervention, explicitly state: +"⚠️ **HUMAN INTERVENTION REQUIRED**: {Reason}" + +When finished, output "ORCHESTRATION_COMPLETE". + +## Auto-Execution Rules (Autopilot Mode Only) + +In Autopilot mode, after outputting the task statuses, the Orchestrator MUST auto-execute the next phase: + +1. Determine the next phase for the most advanced task +2. Output the command to run that phase +3. **Execute the command** (the agent should run the phase directly) +4. Wait for the phase to complete (check for `CONTRACT_MET` or the phase's stop condition) +5. If the phase completes successfully, continue to the next phase +6. If the phase fails (FAIL verdict, HUMAN INTERVENTION REQUIRED), stop and report + +When finished, output "ORCHESTRATION_COMPLETE". + Only recommend one phase at a time. Do not suggest running multiple phases in parallel. diff --git a/prompts/referee.md b/prompts/referee.md index 3046425..6b3ef3f 100644 --- a/prompts/referee.md +++ b/prompts/referee.md @@ -6,8 +6,9 @@ You are the Referee. Your job is to objectively evaluate whether the implementat 2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md (if exists) 3. {project}/tasks/{task-name}/BUG_REPORT.md (if exists) 4. {project}/tasks/{task-name}/ADVERSARIAL_BUG_REPORT.md (if exists) -5. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists) -6. {project}/tasks/{task-name}/DESIGN.md (if exists) +5. {project}/tasks/{task-name}/DOC_REVIEW.md (if exists) +6. {project}/tasks/{task-name}/IMPLEMENTATION.md (if exists) +7. {project}/tasks/{task-name}/DESIGN.md (if exists) ## Task @@ -42,11 +43,11 @@ You are the Referee. Your job is to objectively evaluate whether the implementat - Are edge cases covered? - Are there false positives (tests that pass but don't verify)? -### Documentation Review (Grill with Docs) -- Did the agent update all documentation identified in the DESIGN.md? -- Is the documentation accurate and reflects the final implementation? -- Is the documentation clear enough for a developer to understand the new changes? -- Does the documentation cover any edge cases or non-obvious logic? +### Documentation Review +- Read the `DOC_REVIEW.md` produced by the Documentation Review phase +- Verify the Doc Review findings are accurate — are the docs actually complete and accurate? +- If the Doc Review missed any gaps, call them out here +- If the Doc Review flagged issues that were resolved, mark them as resolved ## Verdict @@ -55,7 +56,8 @@ Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with: ```markdown # Verdict: {task-name} -## Verdict: PASS / FAIL / NEEDS_REVIEW +## Status: [PASS / FAIL / NEEDS_REVIEW] +**Completion Date**: {{CURRENT_DATE}} ## Summary {Brief overview of findings} @@ -73,6 +75,9 @@ Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with: ## Score {Assign a score: +10 for PASS, +5 for NEEDS_REVIEW, -10 for FAIL} + +## Reviewer Comments +(Leave blank for the human reviewer to provide feedback) ``` ## Important @@ -81,6 +86,7 @@ Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with: - If you are unsure, mark it as NEEDS_REVIEW and explain why. - Your verdict is final — no appeals. - Explicitly reference both the Bug Finder and Adversarial Bug Finder outputs in your analysis. +- Always include the current date in the Completion Date field. ## Stop Condition (MANDATORY) You are not allowed to end this session until you have produced the VERDICT.md file AND output the exact phrase "CONTRACT_MET". diff --git a/prompts/research.md b/prompts/research.md index e7f7412..cdc8053 100644 --- a/prompts/research.md +++ b/prompts/research.md @@ -11,6 +11,66 @@ Your only job is to produce a clean, unambiguous specification. Do not write cod {task-description} +## Research Protocol (Interactive) + +You are NOT allowed to produce a SPEC.md without first having a thorough discussion with the user. You must actively grill the user for requirements, edge cases, and constraints. + +### Phase 1: Discovery Questions + +Before writing anything, ask the user questions to understand the full scope. Group your questions by category: + +**Core Requirements:** +- What is the primary goal of this feature? +- What problem does it solve? +- Who are the users? +- What are the non-negotiable requirements? + +**Edge Cases:** +- What happens if the input is empty/null? +- What happens if the input is malformed? +- What happens if the input is extremely large? +- What happens if the system is under heavy load? +- What happens if the user cancels mid-operation? + +**Constraints:** +- Are there performance requirements? (latency, throughput, memory) +- Are there security requirements? (authentication, authorization, data protection) +- Are there compliance requirements? (GDPR, HIPAA, etc.) +- Are there integration requirements? (APIs, databases, external services) + +**Scope:** +- What is explicitly NOT part of this feature? +- What can be deferred to a future iteration? + +### Phase 2: Present Draft SPEC + +After asking questions, present a draft SPEC.md for review. The draft should contain: +- Clear goal +- Exact requirements (numbered) +- Acceptance criteria +- Constraints and non-goals +- Recommended implementation approach (high-level only) + +### Phase 3: Review and Refine + +Present the draft SPEC to the user and ask: +- "Does this cover everything? What am I missing?" +- "Are there any requirements that are wrong or incomplete?" +- "Are there any edge cases I should have considered?" +- "Are there any constraints I should have included?" + +Incorporate the user's feedback and revise the SPEC accordingly. Repeat this loop until the user signs off. + +### Phase 4: Get Sign-Off + +Before writing the SPEC.md, you MUST get explicit sign-off from the user. Say: + +> "Based on our discussion, here is the final spec: +> [brief summary] +> Does this cover everything? Please confirm with 'APPROVED' before I finalize." + +Only produce the SPEC.md after the user says "APPROVED" or equivalent. + ## Output Produce a file called SPEC.md at {project}/tasks/{task-name}/SPEC.md that contains: @@ -27,4 +87,4 @@ Do not add implementation details or suggestions. ## Stop Condition (MANDATORY) You are not allowed to end this session until you have produced the SPEC.md file AND output the exact phrase "CONTRACT_MET". -Until then, continue working or ask clarifying questions. \ No newline at end of file +Until then, continue working or ask clarifying questions. diff --git a/prompts/workflow.md b/prompts/workflow.md index 76fae72..8c15fd9 100644 --- a/prompts/workflow.md +++ b/prompts/workflow.md @@ -6,16 +6,59 @@ This file defines the linear progression of a task in the agent-framework. The O | Current State | Signal (Artifact) | Next Phase | Action | | :--- | :--- | :--- | :--- | -| **New Task** | No `SPEC.md` | Research | Generate `SPEC.md` | -| **Research** | Has `SPEC.md` | Implement | Generate code and tests | +| **New Task** | Orchestrator creates `tasks/{task-name}/` with **no files** | Research | Generate `SPEC.md` | +| **Research** | Has `SPEC.md` | Design (optional) or Implement | Generate `DESIGN.md` or code | +| **Design** | Has `DESIGN.md` | Implement | Generate code and tests | | **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` | | **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` | -| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Referee | Generate `VERDICT.md` | +| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Doc Review | Generate `DOC_REVIEW.md` | +| **Doc Review** | Has `DOC_REVIEW.md` | Referee | Generate `VERDICT.md` | | **Referee** | Has `VERDICT.md` | Complete / Review | Finalize or request user intervention | +## Task Creation (Orchestrator Responsibility) + +The Orchestrator is responsible for creating new task folders automatically — users **never** create task folders manually. + +### New tasks from user input +When the Orchestrator detects a new task description: +1. Generate a kebab-case task name from the description +2. Create `{project}/tasks/{task-name}/` (empty — no artifact files) +3. Move the task to the **Research** phase + +The Orchestrator also scans for tasks that have `VERDICT.md` with `PASS` and removes them from the active task list (they can be archived but not auto-deleted). + +**Key principle**: `IMPLEMENTATION.md` is the artifact produced by the **implementation phase**, not the Orchestrator. The Orchestrator only creates the empty folder; the first real artifact is `SPEC.md` from research. + +### Task creation from bugs +When the Orchestrator detects a `VERDICT.md` with `FAIL` or `NEEDS_REVIEW`, the behavior depends on the mode: + +**In Manual Mode**: The Orchestrator MUST create new tasks and report them for the user to run: + +1. **From `FAIL` verdict**: For each item listed under "Findings" that failed, create a new task: + - Task name: `{original-task-name}-fix-{issue}` (e.g., `add-user-auth-fix-null-handling`) + - The Orchestrator creates the folder with an empty `IMPLEMENTATION.md` + - The task starts at the **Bug Find** phase (skip research — the spec already exists) + - The Orchestrator copies the original `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` into the new task folder + +2. **From `NEEDS_REVIEW` verdict**: For each item listed under "Remaining Issues", create a new task: + - Task name: `{original-task-name}-review-{issue}` (e.g., `add-user-auth-review-perf`) + - The Orchestrator creates the folder with an empty `IMPLEMENTATION.md` + - The task starts at the **Bug Find** phase + - The Orchestrator copies the original `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` into the new task folder + +3. **From "Tasks for Review / Tie-Breaks"**: For each item listed, create a new task: + - Task name: `{original-task-name}-tiebreak-{issue}` (e.g., `add-user-auth-tiebreak-auth-gateway`) + - The Orchestrator creates the folder with an empty `IMPLEMENTATION.md` + - The task starts at the **Research** phase (the tie-break may require spec changes) + - The Orchestrator copies the original `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` into the new task folder + +**In Autopilot Mode**: The Orchestrator should NOT auto-create fix/review/tiebreak tasks — it should pause and report that human intervention is required. The user must decide whether to create fix tasks and how to proceed. + +**Note**: When a task starts at the **Bug Find** phase (fix tasks), the Orchestrator skips the Research phase. The implementation agent should first review the existing bugs and spec before fixing them. The Orchestrator signals this by checking for `BUG_REPORT.md` and `ADVERSARIAL_BUG_REPORT.md` in the new task folder. + ## Autopilot Rules -1. **Linear Progression**: Never skip a phase. +1. **Linear Progression**: Never skip a phase (except Design, which is optional). 2. **Artifact Check**: A phase is only considered "complete" if its corresponding artifact exists and is non-empty. 3. **Automatic Transition**: Upon completion of an artifact, the Orchestrator must immediately identify and propose the next phase in the lifecycle. -4. **Human Intervention**: If the Referee marks a task as `NEEDS_REVIEW` or identifies "Tie-Breaks", the Autopilot pauses and waits for user input. +4. **Human Intervention**: If the Referee marks a task as `FAIL`, `NEEDS_REVIEW`, or identifies "Tie-Breaks", the Autopilot pauses and waits for user input — the Orchestrator does NOT auto-create fix/review/tiebreak tasks in Autopilot mode; the user must decide whether to create fix tasks and how to proceed. (In manual mode, the Orchestrator auto-creates these tasks.) diff --git a/visualize-3d.html b/visualize-3d.html deleted file mode 100644 index f324d72..0000000 --- a/visualize-3d.html +++ /dev/null @@ -1,285 +0,0 @@ - - - - - - Agent Framework • 3D Sphere - - - - - -
-
-
Agent Framework
-
3D Phase Visualization
-
-
- -
- - - - -
- Drag to rotate • Scroll to zoom • Click nodes -
- - - - \ No newline at end of file diff --git a/visualize.html b/visualize.html deleted file mode 100644 index a1e9d9e..0000000 --- a/visualize.html +++ /dev/null @@ -1,216 +0,0 @@ - - - - - - Agent Framework • Routes - - - - - -
-
-
-

Agent Framework

-

Phase relationships and routing

-
-
~/.agent-framework
-
- -
- - -
-
-
Phase Graph
-
Click any node for details
-
-
-
- - -
-
Phase Details
-
-
Select a phase to view its prompt, outputs, and dependencies.
-
-
- -
-
- - - - \ No newline at end of file