Implement Autopilot mode: state machine, adversarial bug finding, and proactive orchestration
This commit is contained in:
@@ -4,8 +4,8 @@ IF task type = research → load prompts/research.md + RULES.md
|
||||
IF task type = design → load prompts/design.md + SPEC.md
|
||||
IF task type = implement → load prompts/implement.md + SPEC.md + DESIGN.md + CONTRACT.md
|
||||
IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code
|
||||
IF task type = disprove → load prompts/disprover.md + SPEC.md + BUGS.md
|
||||
IF task type = referee → load prompts/referee.md + SPEC.md + BUGS.md + DISPROVALS.md
|
||||
IF task type = adversarial_bug_find → load prompts/adversarial_bug_find.md + SPEC.md + code
|
||||
IF task type = referee → load prompts/referee.md + SPEC.md + BUG_REPORT.md + ADVERSARIAL_BUG_REPORT.md
|
||||
IF task type = orchestrate → load prompts/orchestrate.md + project structure
|
||||
IF task type = compaction → load prompts/compaction.md
|
||||
|
||||
|
||||
+38
-211
@@ -1,4 +1,4 @@
|
||||
# Onboarding a New Project
|
||||
# Onboarding a Project
|
||||
|
||||
## Quick Checklist
|
||||
|
||||
@@ -7,60 +7,45 @@
|
||||
- [ ] Create `RULES.md` (project level)
|
||||
- [ ] Run exploration ritual with agent (fresh session)
|
||||
- [ ] Agent reads global + project AGENT.md and RULES.md
|
||||
- [ ] Agent reports back
|
||||
- [ ] Create first `tasks/{task-name}/` folder
|
||||
- [ ] Start research phase using `prompts/research.md`
|
||||
|
||||
This is the exact sequence to follow when bringing any new project into the framework.
|
||||
## Autopilot Mode
|
||||
|
||||
## Prompt Rendering Convention
|
||||
The framework now includes an **Autopilot** mode driven by the `orchestrate.md` prompt.
|
||||
|
||||
All prompts are stored as template files in `~/.agent-framework/prompts/`. They use `{placeholder}` syntax.
|
||||
### How Autopilot Works
|
||||
|
||||
### Placeholders
|
||||
1. **State Detection**: The agent scans your `tasks/` directory and detects the current progress of each task based on the artifacts produced (e.g., `SPEC.md`, `IMPLEMENTATION.md`, etc.).
|
||||
2. **Proactive Driving**: Instead of waiting for you to tell it what to do next, the agent will recommend the exact command needed to progress to the next phase.
|
||||
3. **Verification Loop**: Every task follows a strict lifecycle:
|
||||
`Research` $\rightarrow$ `Implement` $\rightarrow$ `Bug Find` $\rightarrow$ `Adversarial Bug Find` $\rightarrow$ `Referee`.
|
||||
|
||||
| Placeholder | Example | Description |
|
||||
|---|---|---|
|
||||
| `{project}` | `/home/laptran/ai-env/projects/invest-copilot` | Absolute path to the project root |
|
||||
| `{task-name}` | `fix-alert-test` | The task folder name (kebab-case) |
|
||||
| `{task-description}` | `Fix the alert test and complete sector rotation` | Brief description of what to do |
|
||||
### Using Autopilot
|
||||
|
||||
### How It Works
|
||||
To enable Autopilot, add the following to your project's `AGENT.md`:
|
||||
|
||||
When you ask me to run a phase, I will:
|
||||
1. Read the template file (e.g., `~/.agent-framework/prompts/implement.md`)
|
||||
2. Replace all `{placeholders}` with actual values
|
||||
3. Execute the rendered prompt
|
||||
|
||||
You never need to copy-paste prompts. Just tell me what to do.
|
||||
|
||||
### Example
|
||||
|
||||
You say: *"Implement the sector rotation task"*
|
||||
|
||||
I do:
|
||||
```
|
||||
Read: ~/.agent-framework/prompts/implement.md
|
||||
Replace: {project} → /home/laptran/ai-env/projects/invest-copilot
|
||||
{task-name} → fix-alert-test-and-complete-sector-rotation
|
||||
{task-description} → Fix alert test and complete sector rotation service
|
||||
Execute: The rendered prompt
|
||||
```markdown
|
||||
## Mode
|
||||
Autopilot: Enabled.
|
||||
- The agent should use the Orchestrator to drive tasks to completion.
|
||||
- When a task is completed, the agent should automatically scan for the next pending task.
|
||||
```
|
||||
|
||||
When you want the agent to drive the project, use the following command:
|
||||
> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."*
|
||||
|
||||
---
|
||||
|
||||
## Scenario A: Existing Project (Drop-In)
|
||||
|
||||
Use this when the project already exists with code, tests, and structure.
|
||||
|
||||
### 1. Create the Project Framework Directory
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/.agent-framework
|
||||
```
|
||||
|
||||
### 2. Create the Two Required Files
|
||||
|
||||
Create these two files inside `.agent-framework/`:
|
||||
|
||||
#### AGENT.md (project level)
|
||||
```markdown
|
||||
# AGENT.md (project-name)
|
||||
@@ -69,23 +54,14 @@ This project uses the global framework at ~/.agent-framework/.
|
||||
|
||||
Additional project rules are in RULES.md.
|
||||
|
||||
Default mode: research → implement → optional verification.
|
||||
## Mode
|
||||
Autopilot: Enabled.
|
||||
```
|
||||
|
||||
#### RULES.md (project level)
|
||||
Start with any hard constraints you already know for this project. Keep it short.
|
||||
|
||||
Example:
|
||||
```markdown
|
||||
# RULES.md (project-name)
|
||||
|
||||
- Always separate research from implementation in fresh sessions.
|
||||
- Never assume existing code or schema — explore first.
|
||||
- [Add any other non-negotiables]
|
||||
```
|
||||
Start with any hard constraints you already know for this project.
|
||||
|
||||
### 3. Run the Initial Exploration Ritual
|
||||
|
||||
Give the agent this prompt in a fresh session:
|
||||
|
||||
```
|
||||
@@ -93,7 +69,6 @@ You have been given a new project at this path:
|
||||
/path/to/project/
|
||||
|
||||
Your first actions must be:
|
||||
|
||||
1. Explore the project root using ls and find.
|
||||
2. Read .agent-framework/AGENT.md
|
||||
3. Read .agent-framework/RULES.md
|
||||
@@ -108,175 +83,27 @@ Report back with:
|
||||
Do not start any task yet.
|
||||
```
|
||||
|
||||
### 4. If You Don't Know What Task to Do Next
|
||||
|
||||
Run a research phase whose goal is to discover the highest-value next task.
|
||||
|
||||
Tell me: *"Research the project and find the next task"*
|
||||
|
||||
I will:
|
||||
1. Read `~/.agent-framework/prompts/research.md`
|
||||
2. Replace `{project}` with the project path
|
||||
3. Set `{task-description}` to "Explore the project and identify the highest-value next task"
|
||||
4. Execute the rendered prompt
|
||||
|
||||
### 5. Pick the First Task
|
||||
|
||||
Once the agent has completed the exploration report (or discovery research), give me the first real task.
|
||||
|
||||
### 6. Create the First Task Folder
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/tasks/first-task-name/
|
||||
```
|
||||
|
||||
I will then produce `SPEC.md` inside that folder during the research phase.
|
||||
---
|
||||
|
||||
## Scenario B: Starting From Scratch
|
||||
|
||||
Use this when you have an idea but no code yet.
|
||||
Follow the same steps as Scenario A, but start by asking the agent to design the initial architecture:
|
||||
|
||||
### 1. Create the Project Framework Directory
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/.agent-framework
|
||||
```
|
||||
|
||||
### 2. Create the Two Required Files
|
||||
|
||||
#### AGENT.md (project level)
|
||||
```markdown
|
||||
# AGENT.md (project-name)
|
||||
|
||||
This project uses the global framework at ~/.agent-framework/.
|
||||
|
||||
Additional project rules are in RULES.md.
|
||||
|
||||
Default mode: research → implement → optional verification.
|
||||
```
|
||||
|
||||
#### RULES.md (project level)
|
||||
```markdown
|
||||
# RULES.md (project-name)
|
||||
|
||||
- Always separate research from implementation in fresh sessions.
|
||||
- Start with a minimal viable structure — no over-engineering.
|
||||
- [Add any other non-negotiables]
|
||||
```
|
||||
|
||||
### 3. Run the Discovery Research
|
||||
|
||||
Tell me: *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"*
|
||||
|
||||
I will:
|
||||
1. Read `~/.agent-framework/prompts/research.md`
|
||||
2. Replace `{project}` with the project path
|
||||
3. Set `{task-description}` to "Help design the initial architecture and first implementation task for a new project"
|
||||
4. Execute the rendered prompt
|
||||
|
||||
### 4. Review and Approve the Spec
|
||||
|
||||
Review the SPEC.md. If it looks good, proceed to implementation. If not, iterate.
|
||||
|
||||
### 5. Create the Task Folder and Implement
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/tasks/01-initial-setup/
|
||||
```
|
||||
|
||||
Then tell me: *"Implement the first task"* and I'll run the implementation phase.
|
||||
|
||||
### 6. Iterate
|
||||
|
||||
After the first task is complete, create the next task folder and repeat:
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/tasks/02-next-feature/
|
||||
```
|
||||
> *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"*
|
||||
|
||||
## Phase Prompts (Template Files)
|
||||
|
||||
All phase prompts are stored as template files. I render them automatically.
|
||||
| Phase | Template | Output | Trigger |
|
||||
|---|---|---|---|
|
||||
| **Research** | `research.md` | `SPEC.md` | *"Research {task-description}"* |
|
||||
| **Implementation** | `implement.md` | Code + Tests | *"Implement the {task-name} task"* |
|
||||
| **Bug Find** | `bug_finder.md` | `BUG_REPORT.md` | *"Find bugs in the {task-name} task"* |
|
||||
| **Adversarial** | `adversarial_bug_find.md` | `ADVERSARIAL_BUG_REPORT.md` | *"Perform adversarial bug find for {task-name}"* |
|
||||
| **Referee** | `referee.md` | `VERDICT.md` | *"Review the {task-name} task"* |
|
||||
|
||||
### Phase 1: Research
|
||||
## Core Principles
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/research.md`
|
||||
**Output**: `SPEC.md`
|
||||
|
||||
**Trigger**: *"Research {task-description}"*
|
||||
|
||||
### Phase 2: Implementation
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/implement.md`
|
||||
**Output**: Code changes + test results
|
||||
|
||||
**Trigger**: *"Implement the {task-name} task"*
|
||||
|
||||
### Phase 3: Bug Finding (Optional)
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/bug_finder.md`
|
||||
**Output**: `BUG_REPORT.md`
|
||||
|
||||
**Trigger**: *"Find bugs in the {task-name} task"*
|
||||
|
||||
### Phase 4: Referee (Optional)
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/referee.md`
|
||||
**Output**: `VERDICT.md`
|
||||
|
||||
**Trigger**: *"Review the {task-name} task"*
|
||||
|
||||
## Core Principles (From the Original Article)
|
||||
|
||||
These principles underpin the entire framework. Internalize them.
|
||||
|
||||
### 1. Context Is Everything
|
||||
- Ruthlessly minimize what the agent sees. Irrelevant history, old notes, or too many skills destroys performance.
|
||||
- Separate research from implementation. Don't make one agent both figure out *what* to build and *how* to implement it in the same session.
|
||||
- Use fresh contexts/sessions per major task or "contract."
|
||||
|
||||
### 2. Handle Sycophancy (The "Desire to Please")
|
||||
Agents are optimized to be helpful and agreeable. This leads to a common failure mode: if you say "find me a bug," it will often find (or invent) one because it wants to deliver.
|
||||
|
||||
**Solutions:**
|
||||
- Use **neutral prompts** ("Search through the database, follow the logic of each component, and report all your findings") instead of leading ones.
|
||||
- Exploit it productively with a **multi-agent validation loop**:
|
||||
- Agent A (Bug Finder): Scored +1 / +5 / +10 based on severity. It becomes hyper-aggressive at finding issues.
|
||||
- Agent B (Adversarial): Gets points for every bug it successfully disproves, but loses double if wrong. It aggressively tries to shoot them down.
|
||||
- Agent C (Referee): Told you have ground truth; scores both previous agents. This yields very high-fidelity results.
|
||||
|
||||
### 3. Define Clear End States ("How to End a Task")
|
||||
Agents know how to start but not when to stop (they'll implement stubs and declare victory).
|
||||
|
||||
**Fixes:**
|
||||
- Heavy use of tests as milestones ("Task is not complete until these X tests pass. You are not allowed to delete or modify the tests.")
|
||||
- Create a **{TASK}_CONTRACT.md** that explicitly lists all acceptance criteria, tests, screenshots, etc. Make this the single source of truth for completion.
|
||||
- Use stop-hooks that prevent the agent from ending the session until the contract is satisfied.
|
||||
|
||||
### 4. Rules + Skills (The Actual Memory System)
|
||||
Treat your CLAUDE.md (or equivalent) as a lightweight router/directory, not a massive dump.
|
||||
|
||||
- **Rules**: Encode preferences and prohibitions ("If coding, read coding-rules.md first"). Make them conditional and nested.
|
||||
- **Skills**: Encode repeatable *recipes* ("This is exactly how we implement authentication" or "This is our research process").
|
||||
- Start minimal. Iteratively add rules/skills as you observe unwanted behavior.
|
||||
- When performance degrades (contradictions or bloat), have the agent consolidate, de-duplicate, and ask you to resolve conflicts.
|
||||
|
||||
### 5. Long-Running Agents
|
||||
24/7 autonomous agents often fail due to context accumulation and drift.
|
||||
|
||||
**Better pattern:**
|
||||
- One focused session per contract/task.
|
||||
- An orchestration layer that spawns new clean sessions.
|
||||
- Avoid throwing everything into one forever-running context.
|
||||
|
||||
### 6. Stay Current Without Chasing
|
||||
Just update your CLI regularly and read the release notes. If Anthropic/OpenAI add or acquire something (skills, memory, planning, etc.), pay attention. Most "new hot harness" hype becomes obsolete quickly.
|
||||
|
||||
## Notes
|
||||
|
||||
- Only create the two files in step 2. Do not copy the entire global framework.
|
||||
- The global `~/.agent-framework/` already contains the prompts and contracts.
|
||||
- Keep project RULES.md short and specific to this project.
|
||||
- Always start with a fresh agent session for each phase.
|
||||
- Phase 3 and 4 are optional but recommended for critical features.
|
||||
- Use `{task-name}_CONTRACT.md` for critical tasks to define explicit acceptance criteria.
|
||||
- **Context Is Everything**: Separate research from implementation. Use fresh sessions per task.
|
||||
- **Sycophancy Management**: Use the multi-agent validation loop (Bug Finder $\rightarrow$ Adversarial $\rightarrow$ Referee) to ensure objective results.
|
||||
- **Clear End States**: Tests are mandatory. A task is not complete until all acceptance criteria in the `SPEC.md` (or `CONTRACT.md`) are met.
|
||||
- **Rules + Skills**: Keep `RULES.md` specific to this project. Use the global framework for general behavior.
|
||||
|
||||
@@ -1,67 +1,36 @@
|
||||
# agent-framework
|
||||
# Agent Framework
|
||||
|
||||
Minimal, prompt-driven agent framework for Claude Code / Codex / any CLI agent.
|
||||
A contract-based operating system for LLM agents, designed to enforce disciplined engineering workflows.
|
||||
|
||||
Radical simplicity. One focused session per task. Context hygiene first.
|
||||
## Core Philosophy
|
||||
The framework prevents agents from "hallucinating" features or jumping into code without a plan. It enforces a strict separation between **Planning (Research)** and **Execution (Implementation)**, mediated by a **Verification** loop.
|
||||
|
||||
## Install (One-time)
|
||||
## The Autopilot Workflow
|
||||
The framework features an **Autopilot** mode that allows the agent to drive a project to completion with minimal intervention.
|
||||
|
||||
```bash
|
||||
git clone https://gitea.yourdomain.com/you/agent-framework.git ~/.agent-framework
|
||||
```
|
||||
### Lifecycle of a Task
|
||||
1. **Research**: Produce a `SPEC.md` (Contract). No code allowed.
|
||||
2. **Implement**: Write code and tests based *only* on the `SPEC.md`.
|
||||
3. **Bug Find**: Aggressive search for bugs and spec deviations.
|
||||
4. **Adversarial Bug Find**: Deep search for complex logic errors, race conditions, and performance issues.
|
||||
5. **Referee**: Objective evaluation of all bugs and the final verdict.
|
||||
|
||||
That's it. The framework now lives at `~/.agent-framework`.
|
||||
### How to use Autopilot
|
||||
Add `Autopilot: Enabled` to your project's `.agent-framework/AGENT.md` and use the following command:
|
||||
|
||||
## Per-Project Setup
|
||||
> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."*
|
||||
|
||||
Every project only needs two files:
|
||||
The agent will then:
|
||||
- Scan the `tasks/` directory.
|
||||
- Identify the current phase of every task based on existing artifacts.
|
||||
- Recommend the next command to progress each task.
|
||||
- Automatically move to the next phase once the current one's artifacts are produced.
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/.agent-framework
|
||||
```
|
||||
## Key Components
|
||||
- `AGENT.md`: Project-specific configuration and mode selection.
|
||||
- `RULES.md`: Living document of project constraints and past failure modes.
|
||||
- `prompts/`: Specialized system prompts for each phase (Research, Implement, Bug Finder, etc.).
|
||||
- `workflow.md`: The state machine governing the Autopilot lifecycle.
|
||||
|
||||
Create:
|
||||
|
||||
- `.agent-framework/AGENT.md` — project router (points to global framework)
|
||||
- `.agent-framework/RULES.md` — project-specific hard constraints
|
||||
|
||||
Then run the onboarding prompt:
|
||||
|
||||
> "Onboard this project to the agent framework"
|
||||
|
||||
## Usage
|
||||
|
||||
All work happens through prompts in `~/.agent-framework/prompts/`:
|
||||
|
||||
### Core Phases
|
||||
- Research → `prompts/research.md` (produces SPEC.md)
|
||||
- Design → `prompts/design.md` (produces DESIGN.md) — recommended for greenfield projects
|
||||
- Implement → `prompts/implement.md`
|
||||
|
||||
### Quality Phases (Optional)
|
||||
- Bug finding → `prompts/bug_finder.md`
|
||||
- Disprove → `prompts/disprover.md`
|
||||
- Referee → `prompts/referee.md`
|
||||
|
||||
### Other
|
||||
- Onboarding → `prompts/onboarding.md`
|
||||
- Compaction → `prompts/compaction.md`
|
||||
- Orchestrate → `prompts/orchestrate.md`
|
||||
|
||||
See `ONBOARDING.md` for detailed drop-in vs from-scratch scenarios and the full philosophy.
|
||||
|
||||
## Principles
|
||||
|
||||
- Less is more
|
||||
- Separate research from implementation (fresh sessions)
|
||||
- AGENT.md is a lightweight if/else router, not a dump
|
||||
- Clear contracts and stop conditions
|
||||
- Multi-agent verification loop when quality matters
|
||||
|
||||
## Updating
|
||||
|
||||
Pull the latest changes anytime:
|
||||
|
||||
```bash
|
||||
cd ~/.agent-framework && git pull
|
||||
```
|
||||
## Installation
|
||||
Clone the repository and set up your project following the instructions in `ONBOARDING.md`.
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
You are the Adversarial Bug Finder.
|
||||
|
||||
Read the SPEC.md and the code.
|
||||
|
||||
Your job is to find bugs that are difficult to spot, such as complex logic errors, race conditions, and performance bottlenecks. Be more aggressive and exhaustive than a standard bug finder.
|
||||
|
||||
Output your findings in ADVERSARIAL_BUG_REPORT.md.
|
||||
|
||||
When finished, output "ADVERSARIAL_BUG_FIND_COMPLETE".
|
||||
@@ -1,9 +0,0 @@
|
||||
You are the Disprover.
|
||||
|
||||
Read the SPEC.md, BUGS.md, and the code.
|
||||
|
||||
For each claimed bug, try to disprove it. Confirm real bugs and explain why false ones are not issues.
|
||||
|
||||
Output your analysis in DISPROVALS.md.
|
||||
|
||||
When finished, output "DISPROVE_COMPLETE".
|
||||
+23
-24
@@ -1,43 +1,42 @@
|
||||
You are in orchestration mode.
|
||||
|
||||
Your job is to analyze the current state of the project and recommend the next logical phase. You do **not** run any phases yourself.
|
||||
You are the Orchestrator Driver. Your job is to analyze the current state of the project and drive it toward completion by identifying and proposing the next logical step in the lifecycle.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/.agent-framework/AGENT.md
|
||||
2. {project}/.agent-framework/RULES.md
|
||||
3. Any existing files under {project}/tasks/
|
||||
3. {project}/.agent-framework/prompts/workflow.md — The State Machine
|
||||
4. Any existing files under {project}/tasks/
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Analysis Rules
|
||||
## Driver Rules
|
||||
|
||||
Examine the tasks/ directory and determine the state of each task folder using these signals:
|
||||
Examine the tasks/ directory and determine the state of each task folder. Use the state machine defined in `workflow.md`.
|
||||
|
||||
- No SPEC.md → needs **research**
|
||||
- Has SPEC.md but no IMPLEMENTATION.md → needs **implement**
|
||||
- Has implementation but no BUG_REPORT.md → can optionally run **bug_find**
|
||||
- Has BUG_REPORT.md but no DISPROVALS.md → can run **disprove**
|
||||
- Has BUG_REPORT.md + DISPROVALS.md but no VERDICT.md → needs **referee**
|
||||
- Has VERDICT.md with PASS → task is complete
|
||||
For each task, determine the current phase based on the existence of artifacts:
|
||||
- No `SPEC.md` → Next Phase: **research**
|
||||
- Has `SPEC.md` but no `IMPLEMENTATION.md` → Next Phase: **implement**
|
||||
- Has `IMPLEMENTATION.md` but no `BUG_REPORT.md` → Next Phase: **bug_find**
|
||||
- Has `BUG_REPORT.md` but no `ADVERSARIAL_BUG_REPORT.md` → Next Phase: **adversarial_bug_find**
|
||||
- Has `ADVERSARIAL_BUG_REPORT.md` but no `VERDICT.md` → Next Phase: **referee**
|
||||
- Has `VERDICT.md` with `PASS` → Task is **complete**
|
||||
- Has `VERDICT.md` with `NEEDS_REVIEW` or `FAIL` → Next Phase: **Human Intervention**
|
||||
|
||||
## Output Format
|
||||
|
||||
For each task, output one of the following recommendations:
|
||||
For each task, output its status and the exact command to move it to the next phase.
|
||||
|
||||
**Recommended next action:**
|
||||
- Phase: research / implement / bug_find / disprove / referee
|
||||
- Exact command to give the agent:
|
||||
> "Research {task-description}"
|
||||
> "Implement the {task-name} task"
|
||||
> "Find bugs in the {task-name} task"
|
||||
> "Disprove the bugs in the {task-name} task"
|
||||
> "Review the {task-name} task"
|
||||
**Task: {task-folder-name}**
|
||||
- **Status**: {Current Phase}
|
||||
- **Next Step**: {Next Phase}
|
||||
- **Command**:
|
||||
> "{Command to trigger the next phase}"
|
||||
|
||||
If everything is complete, say so clearly.
|
||||
If a task requires human intervention (e.g., `NEEDS_REVIEW` or a tie-break), explicitly state:
|
||||
"⚠️ **HUMAN INTERVENTION REQUIRED**: {Reason}"
|
||||
|
||||
When finished, output "ORCHESTRATION_COMPLETE".
|
||||
|
||||
Only recommend one phase at a time. Do not suggest running multiple phases in parallel.
|
||||
|
||||
When finished, output "ORCHESTRATION_COMPLETE".
|
||||
+11
-6
@@ -5,7 +5,7 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
|
||||
3. {project}/tasks/{task-name}/BUG_REPORT.md — bugs found by Bug Finder (if exists)
|
||||
4. {project}/tasks/{task-name}/DISPROVALS.md — analysis from Disprover (if exists)
|
||||
4. {project}/tasks/{task-name}/ADVERSARIAL_BUG_REPORT.md — bugs found by Adversarial Bug Finder (if exists)
|
||||
5. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists)
|
||||
|
||||
## Task
|
||||
@@ -21,10 +21,12 @@ You are the Referee. Your job is to objectively evaluate whether the implementat
|
||||
|
||||
### Bug Resolution
|
||||
- Were all bugs from the BUG_REPORT.md addressed?
|
||||
- Compare Bug Finder claims vs Disprover analysis in DISPROVALS.md:
|
||||
- Which bugs were confirmed real by the Disprover?
|
||||
- Which bugs were successfully disproved (false positives)?
|
||||
- Did the Disprover miss any real issues?
|
||||
- Were all bugs from the ADVERSARIAL_BUG_REPORT.md addressed?
|
||||
- Compare Bug Finder claims vs Adversarial Bug Finder claims:
|
||||
- Which bugs were found by both?
|
||||
- Which bugs were found only by Bug Finder?
|
||||
- Which bugs were found only by Adversarial Bug Finder?
|
||||
- Identify any contradictions or areas of uncertainty.
|
||||
- Are the suggested fixes correct?
|
||||
- Are there new bugs introduced by the fixes?
|
||||
|
||||
@@ -56,6 +58,9 @@ Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with:
|
||||
- {What failed}
|
||||
- {What needs review}
|
||||
|
||||
## Tasks for Review / Tie-Breaks
|
||||
- {List specific tasks for the user to review or tie-break. If there are no contradictions or complex issues, state "None".}
|
||||
|
||||
## Remaining Issues
|
||||
- {List any remaining issues}
|
||||
|
||||
@@ -68,7 +73,7 @@ Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with:
|
||||
- Be objective. Do not let ego or politics influence your verdict.
|
||||
- If you are unsure, mark it as NEEDS_REVIEW and explain why.
|
||||
- Your verdict is final — no appeals.
|
||||
- Explicitly reference both the Bug Finder and Disprover outputs in your analysis.
|
||||
- Explicitly reference both the Bug Finder and Adversarial Bug Finder outputs in your analysis.
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced the VERDICT.md file AND output the exact phrase "CONTRACT_MET".
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# Workflow State Machine
|
||||
|
||||
This file defines the linear progression of a task in the agent-framework. The Orchestrator uses this to determine the next phase.
|
||||
|
||||
## Task Lifecycle
|
||||
|
||||
| Current State | Signal (Artifact) | Next Phase | Action |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| **New Task** | No `SPEC.md` | Research | Generate `SPEC.md` |
|
||||
| **Research** | Has `SPEC.md` | Implement | Generate code and tests |
|
||||
| **Implementation** | Has `IMPLEMENTATION.md` | Bug Find | Generate `BUG_REPORT.md` |
|
||||
| **Bug Find** | Has `BUG_REPORT.md` | Adversarial Bug Find | Generate `ADVERSARIAL_BUG_REPORT.md` |
|
||||
| **Adversarial Bug Find** | Has `ADVERSARIAL_BUG_REPORT.md` | Referee | Generate `VERDICT.md` |
|
||||
| **Referee** | Has `VERDICT.md` | Complete / Review | Finalize or request user intervention |
|
||||
|
||||
## Autopilot Rules
|
||||
|
||||
1. **Linear Progression**: Never skip a phase.
|
||||
2. **Artifact Check**: A phase is only considered "complete" if its corresponding artifact exists and is non-empty.
|
||||
3. **Automatic Transition**: Upon completion of an artifact, the Orchestrator must immediately identify and propose the next phase in the lifecycle.
|
||||
4. **Human Intervention**: If the Referee marks a task as `NEEDS_REVIEW` or identifies "Tie-Breaks", the Autopilot pauses and waits for user input.
|
||||
Reference in New Issue
Block a user