commit 72ae07c3103c0deb1aea9ef9d7270517942999b1 Author: laptran Date: Sat May 30 23:27:09 2026 -0400 Initial commit: minimal agent framework diff --git a/AGENT.md b/AGENT.md new file mode 100644 index 0000000..c566e62 --- /dev/null +++ b/AGENT.md @@ -0,0 +1,9 @@ +# AGENT.md + +IF task type = research → load prompts/research.md + RULES.md +IF task type = implement → load prompts/implement.md + SPEC.md + CONTRACT.md +IF task type = bug_find → load prompts/bug_finder.md + SPEC.md + code +IF task type = disprove → load prompts/disprover.md + SPEC.md + BUGS.md +IF task type = referee → load prompts/referee.md + SPEC.md + BUGS.md + DISPROVALS.md + +Always start by reading this file to determine mode. \ No newline at end of file diff --git a/ONBOARDING.md b/ONBOARDING.md new file mode 100644 index 0000000..639454e --- /dev/null +++ b/ONBOARDING.md @@ -0,0 +1,282 @@ +# Onboarding a New Project + +## Quick Checklist + +- [ ] Create `.agent-framework/` directory in project root +- [ ] Create `AGENT.md` (project level) +- [ ] Create `RULES.md` (project level) +- [ ] Run exploration ritual with agent (fresh session) +- [ ] Agent reads global + project AGENT.md and RULES.md +- [ ] Agent reports back +- [ ] Create first `tasks/{task-name}/` folder +- [ ] Start research phase using `prompts/research.md` + +This is the exact sequence to follow when bringing any new project into the framework. + +## Prompt Rendering Convention + +All prompts are stored as template files in `~/.agent-framework/prompts/`. They use `{placeholder}` syntax. + +### Placeholders + +| Placeholder | Example | Description | +|---|---|---| +| `{project}` | `/home/laptran/ai-env/projects/invest-copilot` | Absolute path to the project root | +| `{task-name}` | `fix-alert-test` | The task folder name (kebab-case) | +| `{task-description}` | `Fix the alert test and complete sector rotation` | Brief description of what to do | + +### How It Works + +When you ask me to run a phase, I will: +1. Read the template file (e.g., `~/.agent-framework/prompts/implement.md`) +2. Replace all `{placeholders}` with actual values +3. Execute the rendered prompt + +You never need to copy-paste prompts. Just tell me what to do. + +### Example + +You say: *"Implement the sector rotation task"* + +I do: +``` +Read: ~/.agent-framework/prompts/implement.md +Replace: {project} → /home/laptran/ai-env/projects/invest-copilot + {task-name} → fix-alert-test-and-complete-sector-rotation + {task-description} → Fix alert test and complete sector rotation service +Execute: The rendered prompt +``` + +## Scenario A: Existing Project (Drop-In) + +Use this when the project already exists with code, tests, and structure. + +### 1. Create the Project Framework Directory + +```bash +mkdir -p /path/to/project/.agent-framework +``` + +### 2. Create the Two Required Files + +Create these two files inside `.agent-framework/`: + +#### AGENT.md (project level) +```markdown +# AGENT.md (project-name) + +This project uses the global framework at ~/.agent-framework/. + +Additional project rules are in RULES.md. + +Default mode: research → implement → optional verification. +``` + +#### RULES.md (project level) +Start with any hard constraints you already know for this project. Keep it short. + +Example: +```markdown +# RULES.md (project-name) + +- Always separate research from implementation in fresh sessions. +- Never assume existing code or schema — explore first. +- [Add any other non-negotiables] +``` + +### 3. Run the Initial Exploration Ritual + +Give the agent this prompt in a fresh session: + +``` +You have been given a new project at this path: +/path/to/project/ + +Your first actions must be: + +1. Explore the project root using ls and find. +2. Read .agent-framework/AGENT.md +3. Read .agent-framework/RULES.md +4. Read ~/.agent-framework/AGENT.md + +Report back with: +- Confirmation the framework files were found and read +- Summary of the project rules +- What process this project expects +- Key observations from the project structure + +Do not start any task yet. +``` + +### 4. If You Don't Know What Task to Do Next + +Run a research phase whose goal is to discover the highest-value next task. + +Tell me: *"Research the project and find the next task"* + +I will: +1. Read `~/.agent-framework/prompts/research.md` +2. Replace `{project}` with the project path +3. Set `{task-description}` to "Explore the project and identify the highest-value next task" +4. Execute the rendered prompt + +### 5. Pick the First Task + +Once the agent has completed the exploration report (or discovery research), give me the first real task. + +### 6. Create the First Task Folder + +```bash +mkdir -p /path/to/project/tasks/first-task-name/ +``` + +I will then produce `SPEC.md` inside that folder during the research phase. + +## Scenario B: Starting From Scratch + +Use this when you have an idea but no code yet. + +### 1. Create the Project Framework Directory + +```bash +mkdir -p /path/to/project/.agent-framework +``` + +### 2. Create the Two Required Files + +#### AGENT.md (project level) +```markdown +# AGENT.md (project-name) + +This project uses the global framework at ~/.agent-framework/. + +Additional project rules are in RULES.md. + +Default mode: research → implement → optional verification. +``` + +#### RULES.md (project level) +```markdown +# RULES.md (project-name) + +- Always separate research from implementation in fresh sessions. +- Start with a minimal viable structure — no over-engineering. +- [Add any other non-negotiables] +``` + +### 3. Run the Discovery Research + +Tell me: *"Help me design the initial architecture for a new project. Here's what I have in mind: {your idea}"* + +I will: +1. Read `~/.agent-framework/prompts/research.md` +2. Replace `{project}` with the project path +3. Set `{task-description}` to "Help design the initial architecture and first implementation task for a new project" +4. Execute the rendered prompt + +### 4. Review and Approve the Spec + +Review the SPEC.md. If it looks good, proceed to implementation. If not, iterate. + +### 5. Create the Task Folder and Implement + +```bash +mkdir -p /path/to/project/tasks/01-initial-setup/ +``` + +Then tell me: *"Implement the first task"* and I'll run the implementation phase. + +### 6. Iterate + +After the first task is complete, create the next task folder and repeat: + +```bash +mkdir -p /path/to/project/tasks/02-next-feature/ +``` + +## Phase Prompts (Template Files) + +All phase prompts are stored as template files. I render them automatically. + +### Phase 1: Research + +**Template**: `~/.agent-framework/prompts/research.md` +**Output**: `SPEC.md` + +**Trigger**: *"Research {task-description}"* + +### Phase 2: Implementation + +**Template**: `~/.agent-framework/prompts/implement.md` +**Output**: Code changes + test results + +**Trigger**: *"Implement the {task-name} task"* + +### Phase 3: Bug Finding (Optional) + +**Template**: `~/.agent-framework/prompts/bug_finder.md` +**Output**: `BUG_REPORT.md` + +**Trigger**: *"Find bugs in the {task-name} task"* + +### Phase 4: Referee (Optional) + +**Template**: `~/.agent-framework/prompts/referee.md` +**Output**: `VERDICT.md` + +**Trigger**: *"Review the {task-name} task"* + +## Core Principles (From the Original Article) + +These principles underpin the entire framework. Internalize them. + +### 1. Context Is Everything +- Ruthlessly minimize what the agent sees. Irrelevant history, old notes, or too many skills destroys performance. +- Separate research from implementation. Don't make one agent both figure out *what* to build and *how* to implement it in the same session. +- Use fresh contexts/sessions per major task or "contract." + +### 2. Handle Sycophancy (The "Desire to Please") +Agents are optimized to be helpful and agreeable. This leads to a common failure mode: if you say "find me a bug," it will often find (or invent) one because it wants to deliver. + +**Solutions:** +- Use **neutral prompts** ("Search through the database, follow the logic of each component, and report all your findings") instead of leading ones. +- Exploit it productively with a **multi-agent validation loop**: + - Agent A (Bug Finder): Scored +1 / +5 / +10 based on severity. It becomes hyper-aggressive at finding issues. + - Agent B (Adversarial): Gets points for every bug it successfully disproves, but loses double if wrong. It aggressively tries to shoot them down. + - Agent C (Referee): Told you have ground truth; scores both previous agents. This yields very high-fidelity results. + +### 3. Define Clear End States ("How to End a Task") +Agents know how to start but not when to stop (they'll implement stubs and declare victory). + +**Fixes:** +- Heavy use of tests as milestones ("Task is not complete until these X tests pass. You are not allowed to delete or modify the tests.") +- Create a **{TASK}_CONTRACT.md** that explicitly lists all acceptance criteria, tests, screenshots, etc. Make this the single source of truth for completion. +- Use stop-hooks that prevent the agent from ending the session until the contract is satisfied. + +### 4. Rules + Skills (The Actual Memory System) +Treat your CLAUDE.md (or equivalent) as a lightweight router/directory, not a massive dump. + +- **Rules**: Encode preferences and prohibitions ("If coding, read coding-rules.md first"). Make them conditional and nested. +- **Skills**: Encode repeatable *recipes* ("This is exactly how we implement authentication" or "This is our research process"). +- Start minimal. Iteratively add rules/skills as you observe unwanted behavior. +- When performance degrades (contradictions or bloat), have the agent consolidate, de-duplicate, and ask you to resolve conflicts. + +### 5. Long-Running Agents +24/7 autonomous agents often fail due to context accumulation and drift. + +**Better pattern:** +- One focused session per contract/task. +- An orchestration layer that spawns new clean sessions. +- Avoid throwing everything into one forever-running context. + +### 6. Stay Current Without Chasing +Just update your CLI regularly and read the release notes. If Anthropic/OpenAI add or acquire something (skills, memory, planning, etc.), pay attention. Most "new hot harness" hype becomes obsolete quickly. + +## Notes + +- Only create the two files in step 2. Do not copy the entire global framework. +- The global `~/.agent-framework/` already contains the prompts and contracts. +- Keep project RULES.md short and specific to this project. +- Always start with a fresh agent session for each phase. +- Phase 3 and 4 are optional but recommended for critical features. +- Use `{task-name}_CONTRACT.md` for critical tasks to define explicit acceptance criteria. diff --git a/README.md b/README.md new file mode 100644 index 0000000..8d963bd --- /dev/null +++ b/README.md @@ -0,0 +1,60 @@ +# agent-framework + +Minimal, prompt-driven agent framework for Claude Code / Codex / any CLI agent. + +Radical simplicity. One focused session per task. Context hygiene first. + +## Install (One-time) + +```bash +git clone https://gitea.yourdomain.com/you/agent-framework.git ~/.agent-framework +``` + +That's it. The framework now lives at `~/.agent-framework`. + +## Per-Project Setup + +Every project only needs two files: + +```bash +mkdir -p /path/to/project/.agent-framework +``` + +Create: + +- `.agent-framework/AGENT.md` — project router (points to global framework) +- `.agent-framework/RULES.md` — project-specific hard constraints + +Then run the onboarding prompt: + +> "Onboard this project to the agent framework" + +## Usage + +All work happens through prompts in `~/.agent-framework/prompts/`: + +- Research → `prompts/research.md` +- Implement → `prompts/implement.md` +- Bug finding → `prompts/bug_finder.md` +- Disprove → `prompts/disprover.md` +- Referee → `prompts/referee.md` +- Onboarding → `prompts/onboarding.md` +- Compaction → `prompts/compaction.md` + +See `ONBOARDING.md` for detailed drop-in vs from-scratch scenarios and the full philosophy. + +## Principles + +- Less is more +- Separate research from implementation (fresh sessions) +- AGENT.md is a lightweight if/else router, not a dump +- Clear contracts and stop conditions +- Multi-agent verification loop when quality matters + +## Updating + +Pull the latest changes anytime: + +```bash +cd ~/.agent-framework && git pull +``` \ No newline at end of file diff --git a/RULES.md b/RULES.md new file mode 100644 index 0000000..cd1953d --- /dev/null +++ b/RULES.md @@ -0,0 +1,5 @@ +# RULES.md + +- Add one rule per observed failure mode with a concrete example. +- Consolidate contradictions monthly. Remove stale rules. +- No rule without a real example of the problem it prevents. \ No newline at end of file diff --git a/contracts/impl.md b/contracts/impl.md new file mode 100644 index 0000000..15c36be --- /dev/null +++ b/contracts/impl.md @@ -0,0 +1,8 @@ +## Acceptance Criteria +- [ ] Code matches SPEC.md exactly +- [ ] All tests pass (if applicable) +- [ ] No extra features added +- [ ] CONTRACT_MET is output when complete + +## Stop Condition +When all checkboxes are checked, output "CONTRACT_MET" and stop. \ No newline at end of file diff --git a/contracts/spec.md b/contracts/spec.md new file mode 100644 index 0000000..fb5803d --- /dev/null +++ b/contracts/spec.md @@ -0,0 +1,8 @@ +## Acceptance Criteria +- [ ] Goal is clearly stated +- [ ] Requirements are numbered and unambiguous +- [ ] Acceptance criteria are testable +- [ ] Non-goals are listed + +## Stop Condition +When all checkboxes are checked, output "CONTRACT_MET" and stop. \ No newline at end of file diff --git a/contracts/verification.md b/contracts/verification.md new file mode 100644 index 0000000..c687aa9 --- /dev/null +++ b/contracts/verification.md @@ -0,0 +1,7 @@ +## Acceptance Criteria +- [ ] All VALID bugs are fixed or justified +- [ ] No new issues introduced +- [ ] Code matches SPEC.md line-by-line + +## Stop Condition +When all checkboxes are checked, output "CONTRACT_MET" and stop. \ No newline at end of file diff --git a/install.sh b/install.sh new file mode 100644 index 0000000..a8e8d16 --- /dev/null +++ b/install.sh @@ -0,0 +1,16 @@ +#!/bin/bash +set -e + +FRAMEWORK_DIR="$HOME/.agent-framework" + +if [ -d "$FRAMEWORK_DIR" ]; then + echo "agent-framework already installed at $FRAMEWORK_DIR" + echo "Run 'cd $FRAMEWORK_DIR && git pull' to update." + exit 0 +fi + +echo "Cloning agent-framework to $FRAMEWORK_DIR..." +git clone https://gitea.yourdomain.com/you/agent-framework.git "$FRAMEWORK_DIR" + +echo "Installation complete." +echo "Next step: cd into a project and run the onboarding prompt." \ No newline at end of file diff --git a/prompts/bug_finder.md b/prompts/bug_finder.md new file mode 100644 index 0000000..a725537 --- /dev/null +++ b/prompts/bug_finder.md @@ -0,0 +1,92 @@ +You are the Bug Finder. Your job is to adversarially test the implementation and find every bug, deviation from spec, and edge case. + +## Read These Files + +1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built +2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists) +3. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists) + +## Task + +{task-description} + +## Adversarial Checklist + +### Spec Compliance +- Does the implementation match the SPEC.md exactly? +- Are there missing features or stubs? +- Are there features not in the spec (scope creep)? + +### Edge Cases +- Empty inputs, null values, zero-length arrays +- Large inputs (performance, memory) +- Malformed data, unexpected types +- Concurrent access, race conditions +- Dependency failures (network, database, API) + +### Security +- SQL injection, XSS, CSRF +- Authentication and authorization gaps +- Data exposure (logs, error messages, API responses) +- Rate limiting, input validation +- File upload, path traversal + +### Data Flow +- Trace data from input to output +- Are mutations safe? +- Is sensitive data exposed? +- Can data be lost or corrupted? + +### Concurrency & Race Conditions +- Shared state without synchronization +- Async operations without error handling +- Deadlocks, livelocks +- Transaction isolation issues + +### Error Handling +- Are all errors caught and logged? +- Are silent failures possible? +- Are swallowed exceptions present? +- Is there graceful degradation? + +### Performance +- O(n^2) or worse algorithms +- N+1 query patterns +- Memory leaks, unbounded caches +- Unbounded loops, infinite recursion + +### Testing +- Are all edge cases covered by tests? +- Are tests actually testing the right things? +- Are there false positives (tests that pass but don't verify)? + +## Output Format + +Produce a BUG_REPORT.md at {project}/tasks/{task-name}/BUG_REPORT.md with: + +```markdown +# Bug Report: {task-name} + +## Summary +{Brief overview of findings} + +## Bugs Found + +### Bug 1: {Title} +- **Severity**: Critical / High / Medium / Low +- **Description**: {What's wrong} +- **Location**: {File:line} +- **Reproduction**: {Steps to reproduce} +- **Suggested Fix**: {How to fix} + +### Bug 2: ... + +## Score +{Assign a score: +1 for low, +5 for medium, +10 for critical} +``` + +## Important + +- Be aggressive. Your job is to find bugs, not to be nice. +- If you find nothing, say so explicitly — but double-check everything first. +- Do NOT invent bugs. Only report real issues. diff --git a/prompts/compaction.md b/prompts/compaction.md new file mode 100644 index 0000000..35c1f49 --- /dev/null +++ b/prompts/compaction.md @@ -0,0 +1,24 @@ +You are performing a compaction pass on the agent's rules and skills. + +## Read These Files +1. {project}/.agent-framework/RULES.md +2. {project}/.agent-framework/AGENT.md +3. Any accumulated notes or previous RULES.md versions in the project + +## Task +Consolidate and clean up the rules and routing logic. + +## Compaction Rules +- Remove duplicate or contradictory rules +- Merge related rules into the smallest number of clear statements +- Update AGENT.md routing logic if any new patterns have emerged +- Keep every rule that still prevents a real observed failure mode +- Delete anything that has not been referenced in the last 5 tasks + +## Output +Produce an updated RULES.md and AGENT.md. + +At the end, output: +"COMPACTION_COMPLETE — X rules removed, Y rules merged, Z rules added." + +Do not start any new tasks. \ No newline at end of file diff --git a/prompts/disprover.md b/prompts/disprover.md new file mode 100644 index 0000000..f329079 --- /dev/null +++ b/prompts/disprover.md @@ -0,0 +1,9 @@ +You are the Disprover. + +Read the SPEC.md, BUGS.md, and the code. + +For each claimed bug, try to disprove it. Confirm real bugs and explain why false ones are not issues. + +Output your analysis in DISPROVALS.md. + +When finished, output "DISPROVE_COMPLETE". \ No newline at end of file diff --git a/prompts/implement.md b/prompts/implement.md new file mode 100644 index 0000000..9803985 --- /dev/null +++ b/prompts/implement.md @@ -0,0 +1,37 @@ +You are in implementation mode. + +## Read These Files + +1. {project}/tasks/{task-name}/SPEC.md — what needs to be built +2. {project}/.agent-framework/RULES.md — project-specific rules +3. {project}/.agent-framework/AGENT.md — project agent config (if exists) +4. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists) + +## Task + +{task-description} + +## Implementation Rules + +- Follow the SPEC.md exactly. Do not add features not listed. +- Write tests first when a correct seam exists. +- Run tests and report full output. +- Keep functions small and focused. +- Use existing patterns in the codebase. + +### End State +- You are NOT done until ALL acceptance criteria in the SPEC.md (and CONTRACT.md if it exists) are met +- Run the full test suite and report results +- Do NOT declare victory until tests pass + +## Deliverables + +Report back with: +- What you changed (file + summary) +- Test results (full output) +- Any decisions you made (and why) +- Any blockers or open questions + +## Stop Condition (MANDATORY) +You are not allowed to end this session until you have produced all required deliverables AND output the exact phrase "CONTRACT_MET". +Until then, continue working or ask clarifying questions. \ No newline at end of file diff --git a/prompts/onboarding.md b/prompts/onboarding.md new file mode 100644 index 0000000..7b904cd --- /dev/null +++ b/prompts/onboarding.md @@ -0,0 +1,52 @@ +You are in onboarding mode for the agent framework. + +Your only job is to set up the minimal agent framework structure in the target project and perform the initial exploration ritual. Do not start any real tasks. + +## Read These Files + +1. ~/.agent-framework/AGENT.md — global framework router +2. ~/.agent-framework/ONBOARDING.md — human reference for drop-in vs from-scratch scenarios +3. {project}/.agent-framework/AGENT.md (if it exists) +4. {project}/.agent-framework/RULES.md (if it exists) + +## Task + +{task-description} + +## Onboarding Ritual (Strict Sequence) + +1. Check if {project}/.agent-framework/ exists. If not, create it. +2. Ensure exactly two files exist inside it: + - AGENT.md (project-level router) + - RULES.md (project-specific constraints) +3. If the files are missing or empty, create minimal versions: + - AGENT.md should point to the global framework and list any project-specific additions. + - RULES.md should contain only hard, non-negotiable constraints for this project. +4. Read the global ~/.agent-framework/AGENT.md and the new project-level AGENT.md + RULES.md. +5. Explore the project root at a high level (ls, key directories, README if present). +6. Produce a short onboarding report. + +## Output + +Create or update the following inside {project}/.agent-framework/: +- AGENT.md +- RULES.md + +Then produce a file called ONBOARDING_REPORT.md at {project}/tasks/onboarding/ONBOARDING_REPORT.md containing: + +- Confirmation that the framework files were created/read +- Summary of the project rules +- What process this project expects +- Key observations from the project structure +- Any missing pieces the human should provide next + +When the ritual is complete, output "CONTRACT_MET" and stop. + +## Stop Condition (MANDATORY) +You are not allowed to end this session until you have produced the ONBOARDING_REPORT.md and output the exact phrase "CONTRACT_MET". +Do not begin any research, implementation, or bug-finding tasks. + +## Important +- Keep everything minimal. Only create the two required files. +- Never copy the entire global framework into the project. +- This is a one-time setup. After this session the normal research → implement flow takes over. \ No newline at end of file diff --git a/prompts/referee.md b/prompts/referee.md new file mode 100644 index 0000000..1374fd4 --- /dev/null +++ b/prompts/referee.md @@ -0,0 +1,75 @@ +You are the Referee. Your job is to objectively evaluate whether the implementation meets the spec and addresses all bugs. + +## Read These Files + +1. {project}/tasks/{task-name}/SPEC.md — what was supposed to be built +2. {project}/tasks/{task-name}/{task-name}_CONTRACT.md — acceptance criteria (if exists) +3. {project}/tasks/{task-name}/BUG_REPORT.md — bugs found by Bug Finder (if exists) +4. {project}/tasks/{task-name}/DISPROVALS.md — analysis from Disprover (if exists) +5. {project}/tasks/{task-name}/IMPLEMENTATION.md — what was actually implemented (if exists) + +## Task + +{task-description} + +## Evaluation Checklist + +### Spec Compliance +- Does the implementation match the SPEC.md exactly? +- Are all features from the spec present and working? +- Are there missing features or stubs? + +### Bug Resolution +- Were all bugs from the BUG_REPORT.md addressed? +- Compare Bug Finder claims vs Disprover analysis in DISPROVALS.md: + - Which bugs were confirmed real by the Disprover? + - Which bugs were successfully disproved (false positives)? + - Did the Disprover miss any real issues? +- Are the suggested fixes correct? +- Are there new bugs introduced by the fixes? + +### Code Quality +- Does the code follow existing patterns? +- Are functions small and focused? +- Is there proper error handling? +- Are there any obvious performance issues? + +### Testing +- Do all tests pass? +- Are edge cases covered? +- Are there false positives (tests that pass but don't verify)? + +## Verdict + +Produce a VERDICT.md at {project}/tasks/{task-name}/VERDICT.md with: + +```markdown +# Verdict: {task-name} + +## Verdict: PASS / FAIL / NEEDS_REVIEW + +## Summary +{Brief overview of findings} + +## Findings +- {What passed} +- {What failed} +- {What needs review} + +## Remaining Issues +- {List any remaining issues} + +## Score +{Assign a score: +10 for PASS, +5 for NEEDS_REVIEW, -10 for FAIL} +``` + +## Important + +- Be objective. Do not let ego or politics influence your verdict. +- If you are unsure, mark it as NEEDS_REVIEW and explain why. +- Your verdict is final — no appeals. +- Explicitly reference both the Bug Finder and Disprover outputs in your analysis. + +## Stop Condition (MANDATORY) +You are not allowed to end this session until you have produced the VERDICT.md file AND output the exact phrase "CONTRACT_MET". +Until then, continue working or ask clarifying questions. \ No newline at end of file diff --git a/prompts/research.md b/prompts/research.md new file mode 100644 index 0000000..e7f7412 --- /dev/null +++ b/prompts/research.md @@ -0,0 +1,30 @@ +You are in research mode. + +Your only job is to produce a clean, unambiguous specification. Do not write code. + +## Read These Files + +1. {project}/.agent-framework/RULES.md — project-specific rules +2. {project}/.agent-framework/AGENT.md — project agent config (if exists) + +## Task + +{task-description} + +## Output + +Produce a file called SPEC.md at {project}/tasks/{task-name}/SPEC.md that contains: + +- Clear goal +- Exact requirements (numbered) +- Acceptance criteria +- Any constraints or non-goals +- Recommended implementation approach (high-level only) + +When the spec is complete, output "CONTRACT_MET" and stop. + +Do not add implementation details or suggestions. + +## Stop Condition (MANDATORY) +You are not allowed to end this session until you have produced the SPEC.md file AND output the exact phrase "CONTRACT_MET". +Until then, continue working or ask clarifying questions. \ No newline at end of file diff --git a/references/stop-hook-pattern.md b/references/stop-hook-pattern.md new file mode 100644 index 0000000..4189c74 --- /dev/null +++ b/references/stop-hook-pattern.md @@ -0,0 +1,59 @@ +# Stop-Hook Pattern (How to Enforce Task Completion) + +This pattern prevents the agent from ending a session until the contract is satisfied, matching the article's recommendation. + +## Core Idea +Instead of relying on the agent to voluntarily output "CONTRACT_MET", the orchestration layer (you or a thin wrapper) only allows the session to terminate when the stop condition is met. + +## Implementation Options + +### Option 1: Explicit Stop Condition in Every Prompt (Recommended) +Add this block to the end of `implement.md`, `research.md`, and any other phase that produces a deliverable: + +``` +## Stop Condition (MANDATORY) +You are not allowed to end this session until one of the following is true: +- You have produced the required output file(s) AND output the exact phrase "CONTRACT_MET" +- You have explicitly stated that the contract cannot be met and explained why + +Until then, continue working or ask clarifying questions. +``` + +### Option 2: Contract File as Gate +Every task folder must contain a `{task-name}_CONTRACT.md`. + +The agent is instructed: +- Read the contract at the start of the session +- Only output "CONTRACT_MET" after every item in the contract has been verified (tests pass, files exist, acceptance criteria checked) + +### Option 3: External Orchestrator Hook (Strongest) +If using a thin wrapper script around the agent CLI: + +```bash +# Example pseudo-wrapper +while true; do + agent run --prompt "$(render_prompt implement.md)" + if grep -q "CONTRACT_MET" last_output.txt; then + break + fi + # Otherwise feed the failure back into the same session or new one +done +``` + +## Recommended Addition to referee.md and implement.md +After the verdict or implementation, add: + +``` +## Final Gate +Before finishing, re-read the CONTRACT.md (if present) and confirm every line item is satisfied. +Only then output "CONTRACT_MET". +``` + +## Files to Update +- Add the stop condition block to: + - prompts/implement.md + - prompts/research.md + - prompts/bug_finder.md (optional) +- Update ONBOARDING.md to document this pattern under "Define Clear End States" + +This turns the voluntary "CONTRACT_MET" into a hard mechanical requirement. \ No newline at end of file diff --git a/templates/contract-template.md b/templates/contract-template.md new file mode 100644 index 0000000..b059f2e --- /dev/null +++ b/templates/contract-template.md @@ -0,0 +1,35 @@ +# Contract: {task-name} + +## Acceptance Criteria + +- [ ] {Criterion 1: e.g., All tests pass} +- [ ] {Criterion 2: e.g., Feature X works with empty input} +- [ ] {Criterion 3: e.g., API returns correct status codes} +- [ ] {Criterion 4: e.g., No new lint errors} +- [ ] {Criterion 5: e.g., Documentation updated} + +## Tests + +- {List all tests that must pass} +- {List any new tests that must be added} + +## Screenshots (if UI) + +- {Describe what screenshots are needed} + +## Files Modified + +- {List exact files changed} + +## Risks & Mitigations + +- {Risk 1} + - Mitigation: {how to handle} + +## Notes + +- {Any additional notes or constraints} + +## Stop Condition + +When all checkboxes are checked and tests pass, output "CONTRACT_MET" and stop. \ No newline at end of file