Files
agent-framework/BUG_REPORT.md
T

124 lines
13 KiB
Markdown
Raw Normal View History

# Bug Report: agent-framework (Project Review)
## Summary
A comprehensive review of the agent-framework project identified **11 bugs** (after removing the visualize*.html files, which eliminated 8 visualization-related bugs). Remaining issues range from critical (Orchestrator is not a real state machine) to low (unused placeholders).
---
## Bugs Found
### Bug 1: design.md — Duplicate "Read These Files" and "Task" sections
- **Severity**: Medium
- **Location**: `prompts/design.md` (original version, lines 6–18)
- **Description**: The "Read These Files" and "Task" sections appear twice in the file. This is a copy-paste duplication.
- **Reproduction**: Read the file; observe duplicate section headers.
- **Suggested Fix**: **RESOLVED** — Rewrote `design.md` completely; the duplicate sections no longer exist.
### Bug 2: VERDICT.md — Bugs never become tasks
- **Severity**: High
- **Location**: `prompts/workflow.md`, `prompts/orchestrate.md`, `prompts/referee.md`
- **Description**: There is no mechanism to turn bugs found during Bug Find / Adversarial Bug Find into new tasks. The workflow goes straight to "Complete / Review" after Referee. When a `VERDICT.md` has `FAIL`, `NEEDS_REVIEW`, "Remaining Issues", or "Tasks for Review / Tie-Breaks" — none of these become actionable tasks. The only way a bug becomes a task is if the human manually says "Fix the X bug". This breaks the Autopilot workflow entirely — a FAIL verdict should auto-create fix tasks.
- **Reproduction**: Run the full lifecycle on a task with bugs. The Referee produces a FAIL verdict. The Orchestrator reports the task as requiring "Human Intervention" but never creates new tasks for the bugs.
- **Suggested Fix**: **RESOLVED** — Updated `workflow.md` and `orchestrate.md` to auto-create tasks from verdicts:
- `FAIL` → For each failing item under "Findings", create `{task-name}-fix-{issue}` starting at Bug Find phase
- `NEEDS_REVIEW` → For each item under "Remaining Issues", create `{task-name}-review-{issue}` starting at Bug Find phase
- "Tasks for Review / Tie-Breaks" → For each item, create `{task-name}-tiebreak-{issue}` starting at Research phase
- Fix tasks copy the original `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` into the new task folder
### Bug 3: install.sh — Placeholder URL that will fail for any user
- **Severity**: High
- **Location**: `install.sh`, line 12
- **Description**: The `git clone` URL `https://gitea.yourdomain.com/you/agent-framework.git` is a placeholder that will fail for any user trying to install the framework. The README.md also has `[INSERT_FRAMEWORK_REPO_URL_HERE]` as a placeholder. These are clearly incomplete but are the only installation instructions provided.
- **Reproduction**: Run `./install.sh` — it will fail with `git clone: fatal: repository 'https://gitea.yourdomain.com/you/agent-framework.git' not found`.
- **Suggested Fix**: Replace the placeholder URL with the actual repository URL. Alternatively, the install script could be updated to clone from the current directory if the script is run from within the repo.
- **Status**: Known issue — not a framework bug, just a placeholder that needs to be filled in by the user.
### Bug 4: README.md — Placeholder URL that will fail for any user
- **Severity**: High
- **Location**: `README.md`, line 7
- **Description**: The README has `[INSERT_FRAMEWORK_REPO_URL_HERE]` as a placeholder for the git clone URL. This is the same issue as Bug 3 but in the documentation — anyone following the README will fail to clone.
- **Reproduction**: Follow the README instructions; the `[INSERT_FRAMEWORK_REPO_URL_HERE]` placeholder will not resolve.
- **Suggested Fix**: Replace with the actual repository URL.
- **Status**: Known issue — not a framework bug, just a placeholder that needs to be filled in by the user.
### Bug 5: design.md — Unused placeholder `{task-description}` in duplicated section
- **Severity**: Low
- **Location**: `prompts/design.md` (original version, line 18)
- **Description**: The duplicated "Task" section references `{task-description}`, but the design phase never uses the task description in its output.
- **Reproduction**: Read the original design.md; observe the second "Task" section references `{task-description}` but the DESIGN.md output has no mechanism to use it.
- **Suggested Fix**: **RESOLVED** — Rewrote `design.md` completely; the duplicate section and unused placeholder no longer exist.
### Bug 6: orchestrate.md — Unused `{task-description}` placeholder
- **Severity**: Low
- **Location**: `prompts/orchestrate.md`, `## Task` section
- **Description**: The Orchestrator prompt has a `## Task` heading followed by `{task-description}`, but the Orchestrator's logic is generic — it scans all task folders and reports their status. There is no actual use of `{task-description}` in the Orchestrator's content. The placeholder is never rendered or used.
- **Reproduction**: Read `prompts/orchestrate.md`; notice `{task-description}` appears as a heading but the actual orchestration logic is a generic scan of all tasks.
- **Suggested Fix**: **RESOLVED** — `{task-description}` is used in the Orchestrator to determine if the user is starting a new task or continuing. If empty, the Orchestrator scans for the most advanced task. If not empty, the Orchestrator creates a new task from the description.
### Bug 7: orchestrator.md — Orchestrator is not a real state machine
- **Severity**: Critical
- **Location**: `prompts/orchestrate.md` (original version)
- **Description**: The Orchestrator only **reports** the current state and the next command — it doesn't **execute** the phase. After Bug Find, the Orchestrator would tell you the next phase (Adversarial Bug Find) and give you the command, but nothing actually runs. The "Autopilot" is currently a recommendation engine, not an autopilot. A user running the full lifecycle would see: Bug Find completes → Orchestrator says "next: Adversarial Bug Find" → nothing happens → user has to manually run the command.
- **Reproduction**: Run the full lifecycle on a task. After Bug Find, run "orchestrate". The Orchestrator reports the next phase but doesn't execute it. The user has to manually type the command.
- **Suggested Fix**: **RESOLVED** — Rewrote `orchestrate.md` as a proper state machine with auto-transitions. Added:
- Clear state definitions with conditions and transitions
- Autopilot mode (enabled via `Autopilot: Enabled` in AGENT.md) that drives the task all the way to completion
- Auto-execution loop that runs phases sequentially until a terminal state is reached
- Manual mode that only reports state and commands
- Updated AGENT.md to include `Autopilot: Disabled` as a config line
### Bug 8: UX — User doesn't know what to say to continue
- **Severity**: High
- **Location**: `ONBOARDING.md`, `README.md`, `prompts/orchestrate.md`
- **Description**: After running a phase (e.g., Bug Find), the user doesn't know what to say next. The ONBOARDING.md lists specific trigger phrases like "Perform adversarial bug find for {task-name}" which are hard to remember. The user shouldn't need to memorize phase-specific commands — they should just be able to say "orchestrate" or "continue" and the Orchestrator should figure out the next step.
- **Reproduction**: Run Bug Find on a task. After it completes, the user doesn't know what to say next. They have to remember the exact phrase "Perform adversarial bug find for {task-name}".
- **Suggested Fix**: **RESOLVED** — Updated `ONBOARDING.md` and `README.md` to tell the user to just say "orchestrate" or "continue". Updated `orchestrate.md` to recognize "orchestrate"/"continue" with no task description as a signal to scan for the most advanced task and continue from there.
### Bug 9: Documentation Review — No dedicated phase, "Grill with Docs" is just a checklist item
- **Severity**: High
- **Location**: `prompts/referee.md`, `prompts/implement.md`
- **Description**: The Referee has a "Grill with Docs" checklist item that asks "did the agent update all documentation?" but there's no dedicated documentation review phase. The implementation phase is supposed to update docs (mentioned in implement.md End State), but there's no mechanism to verify this before the Referee. If docs are missing, the Referee can only FAIL the task — it can't tell the agent to fix the docs. There's no way to say "update the docs and re-run" without going through the entire bug-fix loop.
- **Reproduction**: Run the full lifecycle on a task. Implementation is done but docs are missing. Bug Find finds no code bugs. Adversarial Bug Find finds no code bugs. Referee says "docs are missing" and FAILs. The only way to fix is to create a new fix task, which is overkill for a doc update.
- **Suggested Fix**: **RESOLVED** — Added a dedicated **Doc Review** phase between Adversarial Bug Find and Referee. The Doc Review phase reads the DESIGN.md Documentation Plan, checks all docs, and updates any missing docs. Produces `DOC_REVIEW.md`. The Referee now reads the DOC_REVIEW.md instead of doing its own ad-hoc review.
### Bug 10: Research and Design phases — Not interactive, agent just produces artifacts
- **Severity**: High
- **Location**: `prompts/research.md`, `prompts/design.md`
- **Description**: The Research and Design phases are completely passive — the agent just produces a SPEC.md or DESIGN.md and moves on. There's no mechanism for the agent to actively grill the user for requirements, edge cases, and design decisions. The user has to provide everything upfront, which leads to incomplete specs and designs. The agent should ask questions, present drafts, get feedback, and get sign-off before producing the artifact.
- **Reproduction**: Start "Research add user auth". The agent immediately produces a SPEC.md without asking any questions. Start "Design the add-user-auth task". The agent produces a DESIGN.md without asking about data model, architecture, or user flows.
- **Suggested Fix**: **RESOLVED** — Rewrote both `research.md` and `design.md` with an interactive protocol:
- **Phase 1: Discovery Questions** — Agent actively asks the user questions grouped by category (requirements, edge cases, constraints, architecture, etc.)
- **Phase 2: Present Draft** — Agent presents a draft SPEC/DESIGN for review
- **Phase 3: Review and Refine** — Agent incorporates user feedback and revises
- **Phase 4: Get Sign-Off** — Agent explicitly asks for "APPROVED" before finalizing
- This ensures the agent doesn't produce artifacts without first having a thorough discussion with the user
---
### Bug 11: Orchestrator contradiction — Autopilot mode task creation from verdicts
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Tasks from bug verdicts" section
- **Description**: The "Tasks from bug verdicts" section says the Orchestrator MUST create tasks from FAIL/NEEDS_REVIEW verdicts in Autopilot mode, but then there's a note saying NOT to create them in Autopilot mode. This is contradictory. The note is correct — in Autopilot mode, the Orchestrator should pause when a FAIL verdict is found and let the user decide. The main text incorrectly says to auto-create tasks in Autopilot mode.
- **Reproduction**: Read `prompts/orchestrate.md`; the "Tasks from bug verdicts" section says "MUST also create new tasks and drive them" in Autopilot mode, but the note says "the Orchestrator should NOT create fix/review/tiebreak tasks from FAIL/NEEDS_REVIEW verdicts — it should pause and let the user decide".
- **Suggested Fix**: **RESOLVED** — Split the section into Manual Mode (create tasks) and Autopilot Mode (pause and report human intervention required).
---
## Score
| Bug | Severity | Score | Status |
|-----|----------|-------|--------|
| 1 | Medium | +5 | **RESOLVED** — Rewrote `design.md`; duplicate sections no longer exist |
| 2 | High | +5 | **RESOLVED** — Verdicts now auto-create tasks (fix/review/tiebreak) |
| 3 | High | +5 | Known issue — placeholder URL in install.sh |
| 4 | High | +5 | Known issue — placeholder URL in README.md |
| 5 | Low | +1 | **RESOLVED** — Rewrote `design.md`; duplicate section and unused placeholder no longer exist |
| 6 | Low | +1 | **RESOLVED** — `{task-description}` is used to determine if user is starting new task or continuing |
| 7 | Critical | +10 | **RESOLVED** — Orchestrator is now a real state machine with auto-transitions |
| 8 | High | +5 | **RESOLVED** — User can just say "orchestrate" or "continue" |
| 9 | High | +5 | **RESOLVED** — Dedicated Doc Review phase added between Adversarial Bug Find and Referee |
| 10 | High | +5 | **RESOLVED** — Research and Design phases now interactive with user |
| 11 | Medium | +5 | **RESOLVED** — Split Orchestrator verdict task creation into Manual/Autopilot modes |
| **Total** | | **47** | |