Implement Autopilot mode: state machine, adversarial bug finding, and proactive orchestration

This commit is contained in:
2026-06-09 23:58:48 -04:00
parent d235c12ab0
commit 59b6339765
8 changed files with 131 additions and 310 deletions
+38 -211
View File
@@ -1,4 +1,4 @@
# Onboarding a New Project
# Onboarding a Project
## Quick Checklist
@@ -7,60 +7,45 @@
- [ ] Create `RULES.md` (project level)
- [ ] Run exploration ritual with agent (fresh session)
- [ ] Agent reads global + project AGENT.md and RULES.md
- [ ] Agent reports back
- [ ] Create first `tasks/{task-name}/` folder
- [ ] Start research phase using `prompts/research.md`
This is the exact sequence to follow when bringing any new project into the framework.
## Autopilot Mode
## Prompt Rendering Convention
The framework now includes an **Autopilot** mode driven by the `orchestrate.md` prompt.
All prompts are stored as template files in `~/.agent-framework/prompts/`. They use `{placeholder}` syntax.
### How Autopilot Works
### Placeholders
1. **State Detection**: The agent scans your `tasks/` directory and detects the current progress of each task based on the artifacts produced (e.g., `SPEC.md`, `IMPLEMENTATION.md`, etc.).
2. **Proactive Driving**: Instead of waiting for you to tell it what to do next, the agent will recommend the exact command needed to progress to the next phase.
3. **Verification Loop**: Every task follows a strict lifecycle:
`Research` $\rightarrow$ `Implement` $\rightarrow$ `Bug Find` $\rightarrow$ `Adversarial Bug Find` $\rightarrow$ `Referee`.
| Placeholder | Example | Description |
|---|---|---|
| `{project}` | `/home/laptran/ai-env/projects/invest-copilot` | Absolute path to the project root |
| `{task-name}` | `fix-alert-test` | The task folder name (kebab-case) |
| `{task-description}` | `Fix the alert test and complete sector rotation` | Brief description of what to do |
### Using Autopilot
### How It Works
To enable Autopilot, add the following to your project's `AGENT.md`:
When you ask me to run a phase, I will:
1. Read the template file (e.g., `~/.agent-framework/prompts/implement.md`)
2. Replace all `{placeholders}` with actual values
3. Execute the rendered prompt
You never need to copy-paste prompts. Just tell me what to do.
### Example
You say: *"Implement the sector rotation task"*
I do:
```
Read: ~/.agent-framework/prompts/implement.md
Replace: {project} → /home/laptran/ai-env/projects/invest-copilot
{task-name} → fix-alert-test-and-complete-sector-rotation
{task-description} → Fix alert test and complete sector rotation service
Execute: The rendered prompt
```markdown
## Mode
Autopilot: Enabled.
- The agent should use the Orchestrator to drive tasks to completion.
- When a task is completed, the agent should automatically scan for the next pending task.
```
When you want the agent to drive the project, use the following command:
> *"Initialize Autopilot for this project. Scan the tasks/ directory and report the current status of all tasks and the recommended next actions."*
---
## Scenario A: Existing Project (Drop-In)
Use this when the project already exists with code, tests, and structure.
### 1. Create the Project Framework Directory
```bash
mkdir -p /path/to/project/.agent-framework
```
### 2. Create the Two Required Files
Create these two files inside `.agent-framework/`:
#### AGENT.md (project level)
```markdown
# AGENT.md (project-name)
@@ -69,23 +54,14 @@ This project uses the global framework at ~/.agent-framework/.
Additional project rules are in RULES.md.
Default mode: research → implement → optional verification.
## Mode
Autopilot: Enabled.
```
#### RULES.md (project level)
Start with any hard constraints you already know for this project. Keep it short.
Example:
```markdown
# RULES.md (project-name)
- Always separate research from implementation in fresh sessions.
- Never assume existing code or schema — explore first.
- [Add any other non-negotiables]
```
Start with any hard constraints you already know for this project.
### 3. Run the Initial Exploration Ritual
Give the agent this prompt in a fresh session:
```
@@ -93,7 +69,6 @@ You have been given a new project at this path:
/path/to/project/
Your first actions must be:
1. Explore the project root using ls and find.
2. Read .agent-framework/AGENT.md
3. Read .agent-framework/RULES.md
@@ -108,175 +83,27 @@ Report back with:
Do not start any task yet.
```
### 4. If You Don't Know What Task to Do Next
Run a research phase whose goal is to discover the highest-value next task.
Tell me: *"Research the project and find the next task"*
I will:
1. Read `~/.agent-framework/prompts/research.md`
2. Replace `{project}` with the project path
3. Set `{task-description}` to "Explore the project and identify the highest-value next task"
4. Execute the rendered prompt
### 5. Pick the First Task
Once the agent has completed the exploration report (or discovery research), give me the first real task.
### 6. Create the First Task Folder
```bash
mkdir -p /path/to/project/tasks/first-task-name/
```
I will then produce `SPEC.md` inside that folder during the research phase.
---
## Scenario B: Starting From Scratch
Use this when you have an idea but no code yet.
Follow the same steps as Scenario A, but start by asking the agent to design the initial architecture:
### 1. Create the Project Framework Directory
```bash
mkdir -p /path/to/project/.agent-framework
```
### 2. Create the Two Required Files
#### AGENT.md (project level)
```markdown
# AGENT.md (project-name)
This project uses the global framework at ~/.agent-framework/.
Additional project rules are in RULES.md.
Default mode: research → implement → optional verification.
```
#### RULES.md (project level)
```markdown
# RULES.md (project-name)
- Always separate research from implementation in fresh sessions.
- Start with a minimal viable structure — no over-engineering.
- [Add any other non-negotiables]
```
### 3. Run the Discovery Research
Tell me: *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"*
I will:
1. Read `~/.agent-framework/prompts/research.md`
2. Replace `{project}` with the project path
3. Set `{task-description}` to "Help design the initial architecture and first implementation task for a new project"
4. Execute the rendered prompt
### 4. Review and Approve the Spec
Review the SPEC.md. If it looks good, proceed to implementation. If not, iterate.
### 5. Create the Task Folder and Implement
```bash
mkdir -p /path/to/project/tasks/01-initial-setup/
```
Then tell me: *"Implement the first task"* and I'll run the implementation phase.
### 6. Iterate
After the first task is complete, create the next task folder and repeat:
```bash
mkdir -p /path/to/project/tasks/02-next-feature/
```
> *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"*
## Phase Prompts (Template Files)
All phase prompts are stored as template files. I render them automatically.
| Phase | Template | Output | Trigger |
|---|---|---|---|
| **Research** | `research.md` | `SPEC.md` | *"Research {task-description}"* |
| **Implementation** | `implement.md` | Code + Tests | *"Implement the {task-name} task"* |
| **Bug Find** | `bug_finder.md` | `BUG_REPORT.md` | *"Find bugs in the {task-name} task"* |
| **Adversarial** | `adversarial_bug_find.md` | `ADVERSARIAL_BUG_REPORT.md` | *"Perform adversarial bug find for {task-name}"* |
| **Referee** | `referee.md` | `VERDICT.md` | *"Review the {task-name} task"* |
### Phase 1: Research
## Core Principles
**Template**: `~/.agent-framework/prompts/research.md`
**Output**: `SPEC.md`
**Trigger**: *"Research {task-description}"*
### Phase 2: Implementation
**Template**: `~/.agent-framework/prompts/implement.md`
**Output**: Code changes + test results
**Trigger**: *"Implement the {task-name} task"*
### Phase 3: Bug Finding (Optional)
**Template**: `~/.agent-framework/prompts/bug_finder.md`
**Output**: `BUG_REPORT.md`
**Trigger**: *"Find bugs in the {task-name} task"*
### Phase 4: Referee (Optional)
**Template**: `~/.agent-framework/prompts/referee.md`
**Output**: `VERDICT.md`
**Trigger**: *"Review the {task-name} task"*
## Core Principles (From the Original Article)
These principles underpin the entire framework. Internalize them.
### 1. Context Is Everything
- Ruthlessly minimize what the agent sees. Irrelevant history, old notes, or too many skills destroys performance.
- Separate research from implementation. Don't make one agent both figure out *what* to build and *how* to implement it in the same session.
- Use fresh contexts/sessions per major task or "contract."
### 2. Handle Sycophancy (The "Desire to Please")
Agents are optimized to be helpful and agreeable. This leads to a common failure mode: if you say "find me a bug," it will often find (or invent) one because it wants to deliver.
**Solutions:**
- Use **neutral prompts** ("Search through the database, follow the logic of each component, and report all your findings") instead of leading ones.
- Exploit it productively with a **multi-agent validation loop**:
- Agent A (Bug Finder): Scored +1 / +5 / +10 based on severity. It becomes hyper-aggressive at finding issues.
- Agent B (Adversarial): Gets points for every bug it successfully disproves, but loses double if wrong. It aggressively tries to shoot them down.
- Agent C (Referee): Told you have ground truth; scores both previous agents. This yields very high-fidelity results.
### 3. Define Clear End States ("How to End a Task")
Agents know how to start but not when to stop (they'll implement stubs and declare victory).
**Fixes:**
- Heavy use of tests as milestones ("Task is not complete until these X tests pass. You are not allowed to delete or modify the tests.")
- Create a **{TASK}_CONTRACT.md** that explicitly lists all acceptance criteria, tests, screenshots, etc. Make this the single source of truth for completion.
- Use stop-hooks that prevent the agent from ending the session until the contract is satisfied.
### 4. Rules + Skills (The Actual Memory System)
Treat your CLAUDE.md (or equivalent) as a lightweight router/directory, not a massive dump.
- **Rules**: Encode preferences and prohibitions ("If coding, read coding-rules.md first"). Make them conditional and nested.
- **Skills**: Encode repeatable *recipes* ("This is exactly how we implement authentication" or "This is our research process").
- Start minimal. Iteratively add rules/skills as you observe unwanted behavior.
- When performance degrades (contradictions or bloat), have the agent consolidate, de-duplicate, and ask you to resolve conflicts.
### 5. Long-Running Agents
24/7 autonomous agents often fail due to context accumulation and drift.
**Better pattern:**
- One focused session per contract/task.
- An orchestration layer that spawns new clean sessions.
- Avoid throwing everything into one forever-running context.
### 6. Stay Current Without Chasing
Just update your CLI regularly and read the release notes. If Anthropic/OpenAI add or acquire something (skills, memory, planning, etc.), pay attention. Most "new hot harness" hype becomes obsolete quickly.
## Notes
- Only create the two files in step 2. Do not copy the entire global framework.
- The global `~/.agent-framework/` already contains the prompts and contracts.
- Keep project RULES.md short and specific to this project.
- Always start with a fresh agent session for each phase.
- Phase 3 and 4 are optional but recommended for critical features.
- Use `{task-name}_CONTRACT.md` for critical tasks to define explicit acceptance criteria.
- **Context Is Everything**: Separate research from implementation. Use fresh sessions per task.
- **Sycophancy Management**: Use the multi-agent validation loop (Bug Finder $\rightarrow$ Adversarial $\rightarrow$ Referee) to ensure objective results.
- **Clear End States**: Tests are mandatory. A task is not complete until all acceptance criteria in the `SPEC.md` (or `CONTRACT.md`) are met.
- **Rules + Skills**: Keep `RULES.md` specific to this project. Use the global framework for general behavior.