Initial commit: minimal agent framework
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
# AGENT.md
|
||||
|
||||
IF task type = research → load prompts/research.md + RULES.md
|
||||
IF task type = implement → load prompts/implement.md + SPEC.md + CONTRACT.md
|
||||
IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code
|
||||
IF task type = disprove → load prompts/disprover.md + SPEC.md + BUGS.md
|
||||
IF task type = referee → load prompts/referee.md + SPEC.md + BUGS.md + DISPROVALS.md
|
||||
|
||||
Always start by reading this file to determine mode.
|
||||
+282
@@ -0,0 +1,282 @@
|
||||
# Onboarding a New Project
|
||||
|
||||
## Quick Checklist
|
||||
|
||||
- [ ] Create `.agent-framework/` directory in project root
|
||||
- [ ] Create `AGENT.md` (project level)
|
||||
- [ ] Create `RULES.md` (project level)
|
||||
- [ ] Run exploration ritual with agent (fresh session)
|
||||
- [ ] Agent reads global + project AGENT.md and RULES.md
|
||||
- [ ] Agent reports back
|
||||
- [ ] Create first `tasks/{task-name}/` folder
|
||||
- [ ] Start research phase using `prompts/research.md`
|
||||
|
||||
This is the exact sequence to follow when bringing any new project into the framework.
|
||||
|
||||
## Prompt Rendering Convention
|
||||
|
||||
All prompts are stored as template files in `~/.agent-framework/prompts/`. They use `{placeholder}` syntax.
|
||||
|
||||
### Placeholders
|
||||
|
||||
| Placeholder | Example | Description |
|
||||
|---|---|---|
|
||||
| `{project}` | `/home/laptran/ai-env/projects/invest-copilot` | Absolute path to the project root |
|
||||
| `{task-name}` | `fix-alert-test` | The task folder name (kebab-case) |
|
||||
| `{task-description}` | `Fix the alert test and complete sector rotation` | Brief description of what to do |
|
||||
|
||||
### How It Works
|
||||
|
||||
When you ask me to run a phase, I will:
|
||||
1. Read the template file (e.g., `~/.agent-framework/prompts/implement.md`)
|
||||
2. Replace all `{placeholders}` with actual values
|
||||
3. Execute the rendered prompt
|
||||
|
||||
You never need to copy-paste prompts. Just tell me what to do.
|
||||
|
||||
### Example
|
||||
|
||||
You say: *"Implement the sector rotation task"*
|
||||
|
||||
I do:
|
||||
```
|
||||
Read: ~/.agent-framework/prompts/implement.md
|
||||
Replace: {project} → /home/laptran/ai-env/projects/invest-copilot
|
||||
{task-name} → fix-alert-test-and-complete-sector-rotation
|
||||
{task-description} → Fix alert test and complete sector rotation service
|
||||
Execute: The rendered prompt
|
||||
```
|
||||
|
||||
## Scenario A: Existing Project (Drop-In)
|
||||
|
||||
Use this when the project already exists with code, tests, and structure.
|
||||
|
||||
### 1. Create the Project Framework Directory
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/.agent-framework
|
||||
```
|
||||
|
||||
### 2. Create the Two Required Files
|
||||
|
||||
Create these two files inside `.agent-framework/`:
|
||||
|
||||
#### AGENT.md (project level)
|
||||
```markdown
|
||||
# AGENT.md (project-name)
|
||||
|
||||
This project uses the global framework at ~/.agent-framework/.
|
||||
|
||||
Additional project rules are in RULES.md.
|
||||
|
||||
Default mode: research → implement → optional verification.
|
||||
```
|
||||
|
||||
#### RULES.md (project level)
|
||||
Start with any hard constraints you already know for this project. Keep it short.
|
||||
|
||||
Example:
|
||||
```markdown
|
||||
# RULES.md (project-name)
|
||||
|
||||
- Always separate research from implementation in fresh sessions.
|
||||
- Never assume existing code or schema — explore first.
|
||||
- [Add any other non-negotiables]
|
||||
```
|
||||
|
||||
### 3. Run the Initial Exploration Ritual
|
||||
|
||||
Give the agent this prompt in a fresh session:
|
||||
|
||||
```
|
||||
You have been given a new project at this path:
|
||||
/path/to/project/
|
||||
|
||||
Your first actions must be:
|
||||
|
||||
1. Explore the project root using ls and find.
|
||||
2. Read .agent-framework/AGENT.md
|
||||
3. Read .agent-framework/RULES.md
|
||||
4. Read ~/.agent-framework/AGENT.md
|
||||
|
||||
Report back with:
|
||||
- Confirmation the framework files were found and read
|
||||
- Summary of the project rules
|
||||
- What process this project expects
|
||||
- Key observations from the project structure
|
||||
|
||||
Do not start any task yet.
|
||||
```
|
||||
|
||||
### 4. If You Don't Know What Task to Do Next
|
||||
|
||||
Run a research phase whose goal is to discover the highest-value next task.
|
||||
|
||||
Tell me: *"Research the project and find the next task"*
|
||||
|
||||
I will:
|
||||
1. Read `~/.agent-framework/prompts/research.md`
|
||||
2. Replace `{project}` with the project path
|
||||
3. Set `{task-description}` to "Explore the project and identify the highest-value next task"
|
||||
4. Execute the rendered prompt
|
||||
|
||||
### 5. Pick the First Task
|
||||
|
||||
Once the agent has completed the exploration report (or discovery research), give me the first real task.
|
||||
|
||||
### 6. Create the First Task Folder
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/tasks/first-task-name/
|
||||
```
|
||||
|
||||
I will then produce `SPEC.md` inside that folder during the research phase.
|
||||
|
||||
## Scenario B: Starting From Scratch
|
||||
|
||||
Use this when you have an idea but no code yet.
|
||||
|
||||
### 1. Create the Project Framework Directory
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/.agent-framework
|
||||
```
|
||||
|
||||
### 2. Create the Two Required Files
|
||||
|
||||
#### AGENT.md (project level)
|
||||
```markdown
|
||||
# AGENT.md (project-name)
|
||||
|
||||
This project uses the global framework at ~/.agent-framework/.
|
||||
|
||||
Additional project rules are in RULES.md.
|
||||
|
||||
Default mode: research → implement → optional verification.
|
||||
```
|
||||
|
||||
#### RULES.md (project level)
|
||||
```markdown
|
||||
# RULES.md (project-name)
|
||||
|
||||
- Always separate research from implementation in fresh sessions.
|
||||
- Start with a minimal viable structure — no over-engineering.
|
||||
- [Add any other non-negotiables]
|
||||
```
|
||||
|
||||
### 3. Run the Discovery Research
|
||||
|
||||
Tell me: *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"*
|
||||
|
||||
I will:
|
||||
1. Read `~/.agent-framework/prompts/research.md`
|
||||
2. Replace `{project}` with the project path
|
||||
3. Set `{task-description}` to "Help design the initial architecture and first implementation task for a new project"
|
||||
4. Execute the rendered prompt
|
||||
|
||||
### 4. Review and Approve the Spec
|
||||
|
||||
Review the SPEC.md. If it looks good, proceed to implementation. If not, iterate.
|
||||
|
||||
### 5. Create the Task Folder and Implement
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/tasks/01-initial-setup/
|
||||
```
|
||||
|
||||
Then tell me: *"Implement the first task"* and I'll run the implementation phase.
|
||||
|
||||
### 6. Iterate
|
||||
|
||||
After the first task is complete, create the next task folder and repeat:
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/tasks/02-next-feature/
|
||||
```
|
||||
|
||||
## Phase Prompts (Template Files)
|
||||
|
||||
All phase prompts are stored as template files. I render them automatically.
|
||||
|
||||
### Phase 1: Research
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/research.md`
|
||||
**Output**: `SPEC.md`
|
||||
|
||||
**Trigger**: *"Research {task-description}"*
|
||||
|
||||
### Phase 2: Implementation
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/implement.md`
|
||||
**Output**: Code changes + test results
|
||||
|
||||
**Trigger**: *"Implement the {task-name} task"*
|
||||
|
||||
### Phase 3: Bug Finding (Optional)
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/bug_finder.md`
|
||||
**Output**: `BUG_REPORT.md`
|
||||
|
||||
**Trigger**: *"Find bugs in the {task-name} task"*
|
||||
|
||||
### Phase 4: Referee (Optional)
|
||||
|
||||
**Template**: `~/.agent-framework/prompts/referee.md`
|
||||
**Output**: `VERDICT.md`
|
||||
|
||||
**Trigger**: *"Review the {task-name} task"*
|
||||
|
||||
## Core Principles (From the Original Article)
|
||||
|
||||
These principles underpin the entire framework. Internalize them.
|
||||
|
||||
### 1. Context Is Everything
|
||||
- Ruthlessly minimize what the agent sees. Irrelevant history, old notes, or too many skills destroys performance.
|
||||
- Separate research from implementation. Don't make one agent both figure out *what* to build and *how* to implement it in the same session.
|
||||
- Use fresh contexts/sessions per major task or "contract."
|
||||
|
||||
### 2. Handle Sycophancy (The "Desire to Please")
|
||||
Agents are optimized to be helpful and agreeable. This leads to a common failure mode: if you say "find me a bug," it will often find (or invent) one because it wants to deliver.
|
||||
|
||||
**Solutions:**
|
||||
- Use **neutral prompts** ("Search through the database, follow the logic of each component, and report all your findings") instead of leading ones.
|
||||
- Exploit it productively with a **multi-agent validation loop**:
|
||||
- Agent A (Bug Finder): Scored +1 / +5 / +10 based on severity. It becomes hyper-aggressive at finding issues.
|
||||
- Agent B (Adversarial): Gets points for every bug it successfully disproves, but loses double if wrong. It aggressively tries to shoot them down.
|
||||
- Agent C (Referee): Told you have ground truth; scores both previous agents. This yields very high-fidelity results.
|
||||
|
||||
### 3. Define Clear End States ("How to End a Task")
|
||||
Agents know how to start but not when to stop (they'll implement stubs and declare victory).
|
||||
|
||||
**Fixes:**
|
||||
- Heavy use of tests as milestones ("Task is not complete until these X tests pass. You are not allowed to delete or modify the tests.")
|
||||
- Create a **{TASK}_CONTRACT.md** that explicitly lists all acceptance criteria, tests, screenshots, etc. Make this the single source of truth for completion.
|
||||
- Use stop-hooks that prevent the agent from ending the session until the contract is satisfied.
|
||||
|
||||
### 4. Rules + Skills (The Actual Memory System)
|
||||
Treat your CLAUDE.md (or equivalent) as a lightweight router/directory, not a massive dump.
|
||||
|
||||
- **Rules**: Encode preferences and prohibitions ("If coding, read coding-rules.md first"). Make them conditional and nested.
|
||||
- **Skills**: Encode repeatable *recipes* ("This is exactly how we implement authentication" or "This is our research process").
|
||||
- Start minimal. Iteratively add rules/skills as you observe unwanted behavior.
|
||||
- When performance degrades (contradictions or bloat), have the agent consolidate, de-duplicate, and ask you to resolve conflicts.
|
||||
|
||||
### 5. Long-Running Agents
|
||||
24/7 autonomous agents often fail due to context accumulation and drift.
|
||||
|
||||
**Better pattern:**
|
||||
- One focused session per contract/task.
|
||||
- An orchestration layer that spawns new clean sessions.
|
||||
- Avoid throwing everything into one forever-running context.
|
||||
|
||||
### 6. Stay Current Without Chasing
|
||||
Just update your CLI regularly and read the release notes. If Anthropic/OpenAI add or acquire something (skills, memory, planning, etc.), pay attention. Most "new hot harness" hype becomes obsolete quickly.
|
||||
|
||||
## Notes
|
||||
|
||||
- Only create the two files in step 2. Do not copy the entire global framework.
|
||||
- The global `~/.agent-framework/` already contains the prompts and contracts.
|
||||
- Keep project RULES.md short and specific to this project.
|
||||
- Always start with a fresh agent session for each phase.
|
||||
- Phase 3 and 4 are optional but recommended for critical features.
|
||||
- Use `{task-name}_CONTRACT.md` for critical tasks to define explicit acceptance criteria.
|
||||
@@ -0,0 +1,60 @@
|
||||
# agent-framework
|
||||
|
||||
Minimal, prompt-driven agent framework for Claude Code / Codex / any CLI agent.
|
||||
|
||||
Radical simplicity. One focused session per task. Context hygiene first.
|
||||
|
||||
## Install (One-time)
|
||||
|
||||
```bash
|
||||
git clone https://gitea.yourdomain.com/you/agent-framework.git ~/.agent-framework
|
||||
```
|
||||
|
||||
That's it. The framework now lives at `~/.agent-framework`.
|
||||
|
||||
## Per-Project Setup
|
||||
|
||||
Every project only needs two files:
|
||||
|
||||
```bash
|
||||
mkdir -p /path/to/project/.agent-framework
|
||||
```
|
||||
|
||||
Create:
|
||||
|
||||
- `.agent-framework/AGENT.md` — project router (points to global framework)
|
||||
- `.agent-framework/RULES.md` — project-specific hard constraints
|
||||
|
||||
Then run the onboarding prompt:
|
||||
|
||||
> "Onboard this project to the agent framework"
|
||||
|
||||
## Usage
|
||||
|
||||
All work happens through prompts in `~/.agent-framework/prompts/`:
|
||||
|
||||
- Research → `prompts/research.md`
|
||||
- Implement → `prompts/implement.md`
|
||||
- Bug finding → `prompts/bug_finder.md`
|
||||
- Disprove → `prompts/disprover.md`
|
||||
- Referee → `prompts/referee.md`
|
||||
- Onboarding → `prompts/onboarding.md`
|
||||
- Compaction → `prompts/compaction.md`
|
||||
|
||||
See `ONBOARDING.md` for detailed drop-in vs from-scratch scenarios and the full philosophy.
|
||||
|
||||
## Principles
|
||||
|
||||
- Less is more
|
||||
- Separate research from implementation (fresh sessions)
|
||||
- AGENT.md is a lightweight if/else router, not a dump
|
||||
- Clear contracts and stop conditions
|
||||
- Multi-agent verification loop when quality matters
|
||||
|
||||
## Updating
|
||||
|
||||
Pull the latest changes anytime:
|
||||
|
||||
```bash
|
||||
cd ~/.agent-framework && git pull
|
||||
```
|
||||
@@ -0,0 +1,5 @@
|
||||
# RULES.md
|
||||
|
||||
- Add one rule per observed failure mode with a concrete example.
|
||||
- Consolidate contradictions monthly. Remove stale rules.
|
||||
- No rule without a real example of the problem it prevents.
|
||||
@@ -0,0 +1,8 @@
|
||||
## Acceptance Criteria
|
||||
- [ ] Code matches SPEC.md exactly
|
||||
- [ ] All tests pass (if applicable)
|
||||
- [ ] No extra features added
|
||||
- [ ] CONTRACT_MET is output when complete
|
||||
|
||||
## Stop Condition
|
||||
When all checkboxes are checked, output "CONTRACT_MET" and stop.
|
||||
@@ -0,0 +1,8 @@
|
||||
## Acceptance Criteria
|
||||
- [ ] Goal is clearly stated
|
||||
- [ ] Requirements are numbered and unambiguous
|
||||
- [ ] Acceptance criteria are testable
|
||||
- [ ] Non-goals are listed
|
||||
|
||||
## Stop Condition
|
||||
When all checkboxes are checked, output "CONTRACT_MET" and stop.
|
||||
@@ -0,0 +1,7 @@
|
||||
## Acceptance Criteria
|
||||
- [ ] All VALID bugs are fixed or justified
|
||||
- [ ] No new issues introduced
|
||||
- [ ] Code matches SPEC.md line-by-line
|
||||
|
||||
## Stop Condition
|
||||
When all checkboxes are checked, output "CONTRACT_MET" and stop.
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
#!/bin/bash
|
||||
set -e
|
||||
|
||||
FRAMEWORK_DIR="$HOME/.agent-framework"
|
||||
|
||||
if [ -d "$FRAMEWORK_DIR" ]; then
|
||||
echo "agent-framework already installed at $FRAMEWORK_DIR"
|
||||
echo "Run 'cd $FRAMEWORK_DIR && git pull' to update."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "Cloning agent-framework to $FRAMEWORK_DIR..."
|
||||
git clone https://gitea.yourdomain.com/you/agent-framework.git "$FRAMEWORK_DIR"
|
||||
|
||||
echo "Installation complete."
|
||||
echo "Next step: cd into a project and run the onboarding prompt."
|
||||
@@ -0,0 +1,92 @@
|
||||
You are the Bug Finder. Your job is to adversarially test the implementation and find every bug, deviation from spec, and edge case.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
|
||||
3. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Adversarial Checklist
|
||||
|
||||
### Spec Compliance
|
||||
- Does the implementation match the SPEC.md exactly?
|
||||
- Are there missing features or stubs?
|
||||
- Are there features not in the spec (scope creep)?
|
||||
|
||||
### Edge Cases
|
||||
- Empty inputs, null values, zero-length arrays
|
||||
- Large inputs (performance, memory)
|
||||
- Malformed data, unexpected types
|
||||
- Concurrent access, race conditions
|
||||
- Dependency failures (network, database, API)
|
||||
|
||||
### Security
|
||||
- SQL injection, XSS, CSRF
|
||||
- Authentication and authorization gaps
|
||||
- Data exposure (logs, error messages, API responses)
|
||||
- Rate limiting, input validation
|
||||
- File upload, path traversal
|
||||
|
||||
### Data Flow
|
||||
- Trace data from input to output
|
||||
- Are mutations safe?
|
||||
- Is sensitive data exposed?
|
||||
- Can data be lost or corrupted?
|
||||
|
||||
### Concurrency & Race Conditions
|
||||
- Shared state without synchronization
|
||||
- Async operations without error handling
|
||||
- Deadlocks, livelocks
|
||||
- Transaction isolation issues
|
||||
|
||||
### Error Handling
|
||||
- Are all errors caught and logged?
|
||||
- Are silent failures possible?
|
||||
- Are swallowed exceptions present?
|
||||
- Is there graceful degradation?
|
||||
|
||||
### Performance
|
||||
- O(n^2) or worse algorithms
|
||||
- N+1 query patterns
|
||||
- Memory leaks, unbounded caches
|
||||
- Unbounded loops, infinite recursion
|
||||
|
||||
### Testing
|
||||
- Are all edge cases covered by tests?
|
||||
- Are tests actually testing the right things?
|
||||
- Are there false positives (tests that pass but don't verify)?
|
||||
|
||||
## Output Format
|
||||
|
||||
Produce a BUG_REPORT.md at {project}/tasks/{task-name}/BUG_REPORT.md with:
|
||||
|
||||
```markdown
|
||||
# Bug Report: {task-name}
|
||||
|
||||
## Summary
|
||||
{Brief overview of findings}
|
||||
|
||||
## Bugs Found
|
||||
|
||||
### Bug 1: {Title}
|
||||
- **Severity**: Critical / High / Medium / Low
|
||||
- **Description**: {What's wrong}
|
||||
- **Location**: {File:line}
|
||||
- **Reproduction**: {Steps to reproduce}
|
||||
- **Suggested Fix**: {How to fix}
|
||||
|
||||
### Bug 2: ...
|
||||
|
||||
## Score
|
||||
{Assign a score: +1 for low, +5 for medium, +10 for critical}
|
||||
```
|
||||
|
||||
## Important
|
||||
|
||||
- Be aggressive. Your job is to find bugs, not to be nice.
|
||||
- If you find nothing, say so explicitly — but double-check everything first.
|
||||
- Do NOT invent bugs. Only report real issues.
|
||||
@@ -0,0 +1,24 @@
|
||||
You are performing a compaction pass on the agent's rules and skills.
|
||||
|
||||
## Read These Files
|
||||
1. {project}/.agent-framework/RULES.md
|
||||
2. {project}/.agent-framework/AGENT.md
|
||||
3. Any accumulated notes or previous RULES.md versions in the project
|
||||
|
||||
## Task
|
||||
Consolidate and clean up the rules and routing logic.
|
||||
|
||||
## Compaction Rules
|
||||
- Remove duplicate or contradictory rules
|
||||
- Merge related rules into the smallest number of clear statements
|
||||
- Update AGENT.md routing logic if any new patterns have emerged
|
||||
- Keep every rule that still prevents a real observed failure mode
|
||||
- Delete anything that has not been referenced in the last 5 tasks
|
||||
|
||||
## Output
|
||||
Produce an updated RULES.md and AGENT.md.
|
||||
|
||||
At the end, output:
|
||||
"COMPACTION_COMPLETE — X rules removed, Y rules merged, Z rules added."
|
||||
|
||||
Do not start any new tasks.
|
||||
@@ -0,0 +1,9 @@
|
||||
You are the Disprover.
|
||||
|
||||
Read the SPEC.md, BUGS.md, and the code.
|
||||
|
||||
For each claimed bug, try to disprove it. Confirm real bugs and explain why false ones are not issues.
|
||||
|
||||
Output your analysis in DISPROVALS.md.
|
||||
|
||||
When finished, output "DISPROVE_COMPLETE".
|
||||
@@ -0,0 +1,37 @@
|
||||
You are in implementation mode.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md — what needs to be built
|
||||
2. {project}/.agent-framework/RULES.md — project-specific rules
|
||||
3. {project}/.agent-framework/AGENT.md — project agent config (if exists)
|
||||
4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Implementation Rules
|
||||
|
||||
- Follow the SPEC.md exactly. Do not add features not listed.
|
||||
- Write tests first when a correct seam exists.
|
||||
- Run tests and report full output.
|
||||
- Keep functions small and focused.
|
||||
- Use existing patterns in the codebase.
|
||||
|
||||
### End State
|
||||
- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met
|
||||
- Run the full test suite and report results
|
||||
- Do NOT declare victory until tests pass
|
||||
|
||||
## Deliverables
|
||||
|
||||
Report back with:
|
||||
- What you changed (file + summary)
|
||||
- Test results (full output)
|
||||
- Any decisions you made (and why)
|
||||
- Any blockers or open questions
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced all required deliverables AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
@@ -0,0 +1,52 @@
|
||||
You are in onboarding mode for the agent framework.
|
||||
|
||||
Your only job is to set up the minimal agent framework structure in the target project and perform the initial exploration ritual. Do not start any real tasks.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. ~/.agent-framework/AGENT.md — global framework router
|
||||
2. ~/.agent-framework/ONBOARDING.md — human reference for drop-in vs from-scratch scenarios
|
||||
3. {project}/.agent-framework/AGENT.md (if it exists)
|
||||
4. {project}/.agent-framework/RULES.md (if it exists)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Onboarding Ritual (Strict Sequence)
|
||||
|
||||
1. Check if {project}/.agent-framework/ exists. If not, create it.
|
||||
2. Ensure exactly two files exist inside it:
|
||||
- AGENT.md (project-level router)
|
||||
- RULES.md (project-specific constraints)
|
||||
3. If the files are missing or empty, create minimal versions:
|
||||
- AGENT.md should point to the global framework and list any project-specific additions.
|
||||
- RULES.md should contain only hard, non-negotiable constraints for this project.
|
||||
4. Read the global ~/.agent-framework/AGENT.md and the new project-level AGENT.md + RULES.md.
|
||||
5. Explore the project root at a high level (ls, key directories, README if present).
|
||||
6. Produce a short onboarding report.
|
||||
|
||||
## Output
|
||||
|
||||
Create or update the following inside {project}/.agent-framework/:
|
||||
- AGENT.md
|
||||
- RULES.md
|
||||
|
||||
Then produce a file called ONBOARDING_REPORT.md at {project}/tasks/onboarding/ONBOARDING_REPORT.md containing:
|
||||
|
||||
- Confirmation that the framework files were created/read
|
||||
- Summary of the project rules
|
||||
- What process this project expects
|
||||
- Key observations from the project structure
|
||||
- Any missing pieces the human should provide next
|
||||
|
||||
When the ritual is complete, output "CONTRACT_MET" and stop.
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced the ONBOARDING_REPORT.md and output the exact phrase "CONTRACT_MET".
|
||||
Do not begin any research, implementation, or bug-finding tasks.
|
||||
|
||||
## Important
|
||||
- Keep everything minimal. Only create the two required files.
|
||||
- Never copy the entire global framework into the project.
|
||||
- This is a one-time setup. After this session the normal research → implement flow takes over.
|
||||
@@ -0,0 +1,75 @@
|
||||
You are the Referee. Your job is to objectively evaluate whether the implementation meets the spec and addresses all bugs.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built
|
||||
2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists)
|
||||
3. {project}/tasks/{task-name}/BUG_REPORT.md — bugs found by Bug Finder (if exists)
|
||||
4. {project}/tasks/{task-name}/DISPROVALS.md — analysis from Disprover (if exists)
|
||||
5. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Evaluation Checklist
|
||||
|
||||
### Spec Compliance
|
||||
- Does the implementation match the SPEC.md exactly?
|
||||
- Are all features from the spec present and working?
|
||||
- Are there missing features or stubs?
|
||||
|
||||
### Bug Resolution
|
||||
- Were all bugs from the BUG_REPORT.md addressed?
|
||||
- Compare Bug Finder claims vs Disprover analysis in DISPROVALS.md:
|
||||
- Which bugs were confirmed real by the Disprover?
|
||||
- Which bugs were successfully disproved (false positives)?
|
||||
- Did the Disprover miss any real issues?
|
||||
- Are the suggested fixes correct?
|
||||
- Are there new bugs introduced by the fixes?
|
||||
|
||||
### Code Quality
|
||||
- Does the code follow existing patterns?
|
||||
- Are functions small and focused?
|
||||
- Is there proper error handling?
|
||||
- Are there any obvious performance issues?
|
||||
|
||||
### Testing
|
||||
- Do all tests pass?
|
||||
- Are edge cases covered?
|
||||
- Are there false positives (tests that pass but don't verify)?
|
||||
|
||||
## Verdict
|
||||
|
||||
Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with:
|
||||
|
||||
```markdown
|
||||
# Verdict: {task-name}
|
||||
|
||||
## Verdict: PASS / FAIL / NEEDS_REVIEW
|
||||
|
||||
## Summary
|
||||
{Brief overview of findings}
|
||||
|
||||
## Findings
|
||||
- {What passed}
|
||||
- {What failed}
|
||||
- {What needs review}
|
||||
|
||||
## Remaining Issues
|
||||
- {List any remaining issues}
|
||||
|
||||
## Score
|
||||
{Assign a score: +10 for PASS, +5 for NEEDS_REVIEW, -10 for FAIL}
|
||||
```
|
||||
|
||||
## Important
|
||||
|
||||
- Be objective. Do not let ego or politics influence your verdict.
|
||||
- If you are unsure, mark it as NEEDS_REVIEW and explain why.
|
||||
- Your verdict is final — no appeals.
|
||||
- Explicitly reference both the Bug Finder and Disprover outputs in your analysis.
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced the VERDICT.md file AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
@@ -0,0 +1,30 @@
|
||||
You are in research mode.
|
||||
|
||||
Your only job is to produce a clean, unambiguous specification. Do not write code.
|
||||
|
||||
## Read These Files
|
||||
|
||||
1. {project}/.agent-framework/RULES.md — project-specific rules
|
||||
2. {project}/.agent-framework/AGENT.md — project agent config (if exists)
|
||||
|
||||
## Task
|
||||
|
||||
{task-description}
|
||||
|
||||
## Output
|
||||
|
||||
Produce a file called SPEC.md at {project}/tasks/{task-name}/SPEC.md that contains:
|
||||
|
||||
- Clear goal
|
||||
- Exact requirements (numbered)
|
||||
- Acceptance criteria
|
||||
- Any constraints or non-goals
|
||||
- Recommended implementation approach (high-level only)
|
||||
|
||||
When the spec is complete, output "CONTRACT_MET" and stop.
|
||||
|
||||
Do not add implementation details or suggestions.
|
||||
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until you have produced the SPEC.md file AND output the exact phrase "CONTRACT_MET".
|
||||
Until then, continue working or ask clarifying questions.
|
||||
@@ -0,0 +1,59 @@
|
||||
# Stop-Hook Pattern (How to Enforce Task Completion)
|
||||
|
||||
This pattern prevents the agent from ending a session until the contract is satisfied, matching the article's recommendation.
|
||||
|
||||
## Core Idea
|
||||
Instead of relying on the agent to voluntarily output "CONTRACT_MET", the orchestration layer (you or a thin wrapper) only allows the session to terminate when the stop condition is met.
|
||||
|
||||
## Implementation Options
|
||||
|
||||
### Option 1: Explicit Stop Condition in Every Prompt (Recommended)
|
||||
Add this block to the end of `implement.md`, `research.md`, and any other phase that produces a deliverable:
|
||||
|
||||
```
|
||||
## Stop Condition (MANDATORY)
|
||||
You are not allowed to end this session until one of the following is true:
|
||||
- You have produced the required output file(s) AND output the exact phrase "CONTRACT_MET"
|
||||
- You have explicitly stated that the contract cannot be met and explained why
|
||||
|
||||
Until then, continue working or ask clarifying questions.
|
||||
```
|
||||
|
||||
### Option 2: Contract File as Gate
|
||||
Every task folder must contain a `{task-name}_CONTRACT.md`.
|
||||
|
||||
The agent is instructed:
|
||||
- Read the contract at the start of the session
|
||||
- Only output "CONTRACT_MET" after every item in the contract has been verified (tests pass, files exist, acceptance criteria checked)
|
||||
|
||||
### Option 3: External Orchestrator Hook (Strongest)
|
||||
If using a thin wrapper script around the agent CLI:
|
||||
|
||||
```bash
|
||||
# Example pseudo-wrapper
|
||||
while true; do
|
||||
agent run --prompt "$(render_prompt implement.md)"
|
||||
if grep -q "CONTRACT_MET" last_output.txt; then
|
||||
break
|
||||
fi
|
||||
# Otherwise feed the failure back into the same session or new one
|
||||
done
|
||||
```
|
||||
|
||||
## Recommended Addition to referee.md and implement.md
|
||||
After the verdict or implementation, add:
|
||||
|
||||
```
|
||||
## Final Gate
|
||||
Before finishing, re-read the CONTRACT.md (if present) and confirm every line item is satisfied.
|
||||
Only then output "CONTRACT_MET".
|
||||
```
|
||||
|
||||
## Files to Update
|
||||
- Add the stop condition block to:
|
||||
- prompts/implement.md
|
||||
- prompts/research.md
|
||||
- prompts/bug_finder.md (optional)
|
||||
- Update ONBOARDING.md to document this pattern under "Define Clear End States"
|
||||
|
||||
This turns the voluntary "CONTRACT_MET" into a hard mechanical requirement.
|
||||
@@ -0,0 +1,35 @@
|
||||
# Contract: {task-name}
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] {Criterion 1: e.g., All tests pass}
|
||||
- [ ] {Criterion 2: e.g., Feature X works with empty input}
|
||||
- [ ] {Criterion 3: e.g., API returns correct status codes}
|
||||
- [ ] {Criterion 4: e.g., No new lint errors}
|
||||
- [ ] {Criterion 5: e.g., Documentation updated}
|
||||
|
||||
## Tests
|
||||
|
||||
- {List all tests that must pass}
|
||||
- {List any new tests that must be added}
|
||||
|
||||
## Screenshots (if UI)
|
||||
|
||||
- {Describe what screenshots are needed}
|
||||
|
||||
## Files Modified
|
||||
|
||||
- {List exact files changed}
|
||||
|
||||
## Risks & Mitigations
|
||||
|
||||
- {Risk 1}
|
||||
- Mitigation: {how to handle}
|
||||
|
||||
## Notes
|
||||
|
||||
- {Any additional notes or constraints}
|
||||
|
||||
## Stop Condition
|
||||
|
||||
When all checkboxes are checked and tests pass, output "CONTRACT_MET" and stop.
|
||||
Reference in New Issue
Block a user