Files
automaton/README.md
T

482 lines
22 KiB
Markdown
Raw Normal View History

# Automaton
2026-05-30 23:27:09 -04:00
A contract-based operating system for LLM agents, designed to enforce disciplined engineering workflows.
2026-05-30 23:27:09 -04:00
## 1. Global Framework Installation (One-time setup)
Before you can use the framework in any project, you must install the core logic into your local environment.
**Run these commands in your terminal:**
```bash
# Clone the framework into the global config directory
git clone <your-git-url> ~/.automaton
# Enter the directory
cd ~/.automaton
# Make the installation script executable and run it
# You must provide the git URL as the first argument
chmod +x install.sh
./install.sh <your-git-url>
```
The git URL is required because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling.
*Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory. A self-improvement loop is created and scheduled by default (see Loop Engineering below).*
### Updating the Framework
When the framework is updated, you can update your global installation:
```bash
cd ~/.automaton
./update.sh
```
This will:
- Fetch the latest changes from the repository
- Check for uncommitted changes and warn you
- Pull the latest updates
**Upgrading existing projects:** The framework reads prompts, contracts, and scripts from `~/.automaton/` at runtime. Projects only override `.agent.md` and `.rules.md`. This means updating the global framework (`git pull`) automatically applies to all projects. No per-project upgrade is needed.
If a project was set up under the old model (with copies of framework files), it needs migration first. Tell the agent: "Upgrade automaton for this project" to run the migration.
---
## 2. Project Setup (Per project)
Once the framework is installed globally, you must "onboard" every individual project you work on.
### Option A: The Agent-Driven Way (Recommended)
If you want the agent to handle the configuration for you, navigate to your project root and run:
> *"Onboard this project into automaton."
The agent will automatically:
1. Detect your project type (New, Existing, or Upgrade).
2. Create the `./.automaton/` directory.
3. Generate your `.agent.md` and `.rules.md` files.
4. Initiate the "Exploration Ritual" to understand your codebase.
### Option B: The Manual Way
If you prefer to set it up manually, create a `.automaton/` directory in your project root and add:
- `.agent.md`: Project-specific configuration (Mode, rules, etc.).
- `.rules.md`: Project-specific constraints and past failure modes.
Then install the git pre-commit hook:
```bash
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit
```
This hook blocks commits when no task is in `implement` or `doc_review` phase, preventing accidental code changes without a tracked task.
### The Self-Improvement Loop is framework-scoped
The self-improvement loop installed by default targets the **framework itself** (`work_source.project: ~/.automaton/`), not your project. It ticks hourly against `status.py --audit` on `~/.automaton/` to keep the framework healthy. This is by design — it improves the framework you depend on while you work on your project.
- **Leave it running** if you want the framework maintained in the background (recommended).
- **Disable it** if you want zero background activity: `python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/`
- **Want a loop on your project too?** Create a separate one targeted at the project root:
```bash
python3 ~/.automaton/scripts/status.py --create-loop my-project-loop \
--from-template self-improvement --project /path/to/project
```
---
2026-05-30 23:27:09 -04:00
## The Autopilot Workflow
The framework features an **Autopilot** mode that allows the agent to drive a project to completion with minimal intervention.
2026-05-30 23:27:09 -04:00
### Lifecycle of a Task
1. **Research**: Produce a `SPEC.md` (Contract). No code allowed. **Interactive** — agent grills you for requirements and gets sign-off.
2. **Decomposition** (optional): Break the task into sub-tasks sized for your VRAM. **Interactive** — agent analyzes the spec and proposes sub-tasks with token budget estimates, gets sign-off. See [VRAM Configuration](#vram-configuration) below.
3. **Design** (optional): Produce a `DESIGN.md` (Architecture). No code allowed. **Interactive** — agent grills you for design decisions and gets sign-off.
3. **Test Design** (optional): Produce a `TEST_PLAN.md` (Test specification). No code allowed. **Interactive** — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off.
4. **Implement**: Write code and tests based *only* on the `SPEC.md`, `DESIGN.md` (if present), and `TEST_PLAN.md` (if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows.
5. **Bug Find**: Aggressive search for bugs and spec deviations.
6. **Adversarial Bug Find**: Deep search for complex logic errors, race conditions, and performance issues.
7. **Doc Review**: Review documentation against DESIGN.md plan and fix missing docs.
8. **Referee**: Objective evaluation of all bugs, docs, and the final verdict.
2026-05-30 23:27:09 -04:00
### How to use Autopilot
2026-05-30 23:27:09 -04:00
#### Autopilot mode (default)
The default mode is **Autopilot: Enabled**. The Orchestrator automatically drives tasks through all phases. Just say:
- **Start a new task**: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion)
- **Decompose a task**: "Decompose the add-user-auth task" (the Orchestrator breaks it into sub-tasks sized for your VRAM)
- **Continue**: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it)
#### VRAM Configuration
For low-VRAM systems (8GB, 16GB), set your VRAM limits in `~/.automaton/config.md`:
```markdown
## VRAM Configuration
- **Auto-detect**: Yes # Let the agent detect your VRAM automatically
- **Target VRAM context**: 16k tokens # Override auto-detect if needed
- **Headroom**: 25%
- **Max peak context per sub-task**: 12k tokens
```
**Auto-detection**: When `Auto-detect: Yes`, the Orchestrator runs `scripts/vram_detect.py` to probe:
- GPU VRAM (via `nvidia-smi`)
- System RAM (via `free`)
- Model context window (from config.md or API config files)
- Framework overhead (by counting token load in loaded prompts)
**Manual override**: When `Auto-detect: No`, use the manually specified values:
```markdown
## VRAM Configuration
- **Auto-detect**: No
- **Target VRAM context**: 8k
- **Headroom**: 30%
- **Max peak context per sub-task**: 5.6k
```
#### Model Configuration
When using a local LLM or a specific API model, set the model in `~/.automaton/config.md`:
```markdown
## Model Configuration
- **Model**: auto # Use auto-detection from API config files
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
```
**Auto-detection**: When `Model: auto`, the framework detects the model name from API config files (`.env`, `config.yaml`, etc.) and looks up its context window.
**Manual override**: When you know your model name, specify it:
```markdown
## Model Configuration
- **Model**: gpt-4o
- **Override context window**: 128k
```
When you run "Decompose the X task", the Orchestrator will:
1. Analyze the task's SPEC.md
2. Detect VRAM limits (auto or manual)
3. Break it into sub-tasks, each sized to fit within your VRAM limit
4. Estimate the token budget for each sub-task
5. Create sub-task folders under `tasks/{parent-task}/subtasks/{sub-task}/`
6. Propagate VRAM config to each sub-task
Sub-tasks run independently through the full lifecycle. The parent task is complete only when ALL sub-tasks pass.
#### Manual mode (opt-in)
Set `Autopilot: Disabled` in your project's `.automaton/.agent.md` if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like:
- "Research add user authentication" — starts a new task
- "Decompose the add-user-auth task" — breaks into VRAM-sized sub-tasks (optional)
- "Design the add-user-auth task" — designs the architecture (optional)
- "Design tests for the add-user-auth task" — designs test cases (optional)
- "Implement the add-user-auth task" — implements the task with tests
- "Find bugs in the add-user-auth task" — finds bugs
- "Perform adversarial bug find for add-user-auth" — deep bug search
- "Review docs for the add-user-auth task" — reviews documentation
- "Review the add-user-auth task" — referee evaluates
- "orchestrate" — asks the Orchestrator what to do next
2026-05-30 23:27:09 -04:00
## Sub-Task Management
When a task is decomposed, the Orchestrator creates sub-tasks under `tasks/{parent-task}/subtasks/`. Each sub-task:
- Has its own full lifecycle (Research → Decomposition → Design → ... → Complete)
- Starts at the Research phase with an empty folder
- Receives a `PARENT_SPEC.md` with the parent task's context
- Receives a `VRAM_CONFIG.md` with the VRAM constraints
- Runs independently — sub-tasks in the same wave can run in parallel
- The parent task is NOT complete until ALL sub-tasks pass
## Key Components
- `.agent.md`: Project-specific agent behavior (Autopilot mode, routing rules, optional Agent Configuration for multi-agent).
- `config.md`: Global framework settings (VRAM, model, system requirements, version).
- `.rules.md`: Living document of project constraints and past failure modes.
- `prompts/`: Specialized system prompts for each phase with ALLOWED/FORBIDDEN sections and approval gates.
- `workflow.md`: The state machine governing the task lifecycle, with `.state` file as canonical phase indicator.
- `scripts/status.py`: Enforcement script — task status, phase transitions, approval gates, folder validation, audits, pre-edit hook (`--can-edit`), multi-agent claiming.
- `scripts/vram_detect.py`: Auto-detects GPU VRAM, RAM, model context window, and framework overhead.
- `scripts/git-hooks/pre-commit`: Blocks commits when no task is in an edit-allowed phase.
- `contracts/harness-integration.md`: Integration contract for agent harnesses (opencode, aider, etc.).
- `plugins/automaton-guard/`: opencode plugin that intercepts `edit`/`write` calls and checks `--can-edit` before allowing them.
2026-05-30 23:27:09 -04:00
## Loop Engineering (beta, v1)
Automaton can run unattended workflow loops: each loop has an OS-level schedule (launchd / cron / schtasks) and is constrained by 6 brake gates enforced in `status.py`. The single source of truth for loop runtime state is `.state.loop` per loop at `{project}/.automaton/loops/<name>/.state.loop`.
```bash
# Create a loop from a template (only way to bootstrap)
python3 ~/.automaton/scripts/status.py --create-loop my-ci-triage --from-template ci-triage --project /path/to/project
# Install the native OS schedule unit (launchd/cron/schtasks)
python3 ~/.automaton/scripts/status.py --install-schedule my-ci-triage --interval 3600 --project /path/to/project
# Pre-tick gate check (6 brakes; first failure halts the loop)
python3 ~/.automaton/scripts/status.py --check-gate my-ci-triage --project /path/to/project
# Clear a halt (only way; no auto-approve in v1)
python3 ~/.automaton/scripts/status.py --approve --loop my-ci-triage --project /path/to/project
# List all loops and their status
python3 ~/.automaton/scripts/status.py --loop-list --project /path/to/project
```
Loop states: `running`, `halted`, `paused`, `complete`. Halt reasons: `iterations_exhausted`, `budget_exhausted`, `verifier_failed`, `drift_detected`, `human_intervention`. The per-tick engine is `scripts/loop-runner.py --mode tick`; it gates first, spawns Implement / Verify / Orchestrate role sessions, and writes `.state.loop` atomically. Concurrent ticks on the same loop and concurrent `--pause-loop` / `--approve --loop` writes are serialized via a cross-process file lock on `<loop_path>/.state.lock` (POSIX `fcntl.flock`, Windows `msvcrt.locking`); see `design/loops/technical.md` §7 "Lock serialization". See `design/loops/technical.md` §7 for the full 11-step flow.
### Quick Start
```bash
# 1. Create a loop from a template
python3 ~/.automaton/scripts/status.py --create-loop my-ci-triage \
--from-template ci-triage --project /path/to/project
# 2. Install the OS schedule (launchd on macOS, cron on Linux, schtasks on Windows)
python3 ~/.automaton/scripts/status.py --install-schedule my-ci-triage \
--interval 3600 --project /path/to/project
# 3. Monitor
python3 ~/.automaton/scripts/status.py --loop-list --project /path/to/project
python3 ~/.automaton/scripts/status.py --audit --project /path/to/project
```
### Tick Cycle
Each tick runs this 11-step flow (see `design/loops/technical.md` §7 for details):
```
gate check -> find work -> ensure worktree -> spawn Implement -> spawn Verify
-> parse verdict -> spawn Orchestrate -> atomic state write -> log
```
The runner resolves prompt files from `loop.json` `roles.*.prompt` (e.g. `loop-implement.md`), substitutes content-level tokens (`{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{artifact_content}`, etc.), writes the resolved prompt to `outputs/tickN-<role>-prompt.md`, and passes it to the harness.
### Configuration (`loop.json`)
| Field | Description |
|-------|-------------|
| `name` | Loop name (kebab-case) |
| `schedule.interval_seconds` | Tick interval for daemon mode |
| `brakes.max_iterations` | Max ticks before halt |
| `brakes.max_budget_usd` | Optional USD budget cap (null = unlimited) |
| `brakes.score_plateau_window` | Score plateau detection window |
| `blast_radius.file_scope` | List of paths the loop may edit |
| `blast_radius.use_worktree` | If true, tick runs in a per-loop git worktree |
| `work_source.kind` | `single`, `audit`, or `backlog` |
| `work_source.area` | Design area for `backlog` kind (default `"loops"`; `"framework"` reads `design/framework/BACKLOG.md`) |
| `roles.implement.prompt` | Prompt file for Implement role |
| `roles.verify.prompt` | Prompt file for Verify role |
| `roles.orchestrate.prompt` | Prompt file for Orchestrate role |
| `harness.command` | Command template with `{prompt}`, `{prompt_content}`, `{cwd}` tokens (default invokes `opencode run --dir <cwd> <prompt>`; override for other harnesses -- Pi Dev, aider, etc.) |
| `acceptance_criteria` | List of criteria for the verifier to check |
### Monitoring
- `--loop-list`: show all loops and their status
- `--audit`: check for violations across all tasks and loops
- `.state.log`: per-loop tick log (ISO-timestamped entries)
- `outputs/`: per-tick artifacts and resolved prompts
### Halt and Resume
```bash
# Pause a loop (disables the OS schedule unit)
python3 ~/.automaton/scripts/status.py --pause-loop my-ci-triage --project /path/to/project
# Resume a paused loop
python3 ~/.automaton/scripts/status.py --resume-loop my-ci-triage --project /path/to/project
# Clear a halt (the only way; no auto-approve in v1)
python3 ~/.automaton/scripts/status.py --approve --loop my-ci-triage --project /path/to/project
```
### Self-Improvement Loop (Default-On)
The framework installs a self-improvement loop by default at install time. This loop ticks against `status.py --audit` on the framework's own repo, picking up audit violations and resolving them unattended. It runs every 3600 seconds (1 hour) with `max_iterations: 10` and a score plateau window of 3.
```bash
# Disable the self-improvement loop
python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/
# Re-enable it
python3 ~/.automaton/scripts/status.py --resume-loop self-improvement --project ~/.automaton/
```
The loop uses a git worktree at `~/.automaton/loops/self-improvement/worktree/` and is scoped to `scripts/`, `prompts/`, `tests/`, and `design/` directories.
## State Enforcement (v2.0)
Automaton v2.0 enforces the state machine computationally, not just via prompts:
- **`.state` file**: Each task has a `.state` file that is the single source of truth for its current phase
- **`status.py --transition`**: All phase transitions must go through this command; illegal transitions are refused
- **Approval gates**: Research, Design, Decomposition, and Test Design phases require explicit user approval before proceeding
- **`status.py --validate-folder`**: Detects out-of-order artifacts (phase skipping)
- **`status.py --audit`**: Comprehensive audit across all tasks for violations
- **`status.py --create-task`**: The only valid way to create task folders
- **FORBIDDEN actions in prompts**: Each phase prompt explicitly lists what agents cannot do
### Quick Reference
```bash
# Create a new task
python3 ~/.automaton/scripts/status.py --create-task add-user-auth --project /path/to/project
# Check task status
python3 ~/.automaton/scripts/status.py --task add-user-auth --project /path/to/project
# List all tasks
python3 ~/.automaton/scripts/status.py --list --project /path/to/project
# Transition to next phase
python3 ~/.automaton/scripts/status.py --transition research --task add-user-auth --project /path/to/project
python3 ~/.automaton/scripts/status.py --transition research:awaiting_approval --task add-user-auth --project /path/to/project
# Approve a phase (after user sign-off)
python3 ~/.automaton/scripts/status.py --approve --task add-user-auth --project /path/to/project
# Validate task folder
python3 ~/.automaton/scripts/status.py --validate-folder --task add-user-auth --project /path/to/project
# Audit all tasks
python3 ~/.automaton/scripts/status.py --audit --project /path/to/project
# Upgrade pre-v2.0 tasks (bootstrap .state files)
python3 ~/.automaton/scripts/status.py --upgrade --project /path/to/project
# Check if code edits are allowed (harness integration)
python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project
python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project --file src/main.py
python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project --task add-user-auth --json
```
**Important**: Always pass `--project` to ensure correct scoping when multiple projects exist. Without it, `status.py` resolves the project from the current directory and errors if not in a project.
### Untracked Tasks
Tasks without `.state` files are UNTRACKED — all commands (`--transition`, `--can-edit`, `--task`, `--approve`) refuse to operate on them. This prevents agents from working on tasks created before v2.0 state enforcement.
To fix untracked tasks:
```bash
# Upgrade a single task
python3 ~/.automaton/scripts/status.py --upgrade --task my-old-task --project /path/to/project
# Upgrade all tasks at once
python3 ~/.automaton/scripts/status.py --upgrade --project /path/to/project
```
### Enforcement
The framework enforces the state machine computationally. No phase can be skipped, no approval can be bypassed, and no code edits can happen without a task in an edit-allowed phase. This is enforced through three layers:
1. **Harness pre-edit hook** (`--can-edit`) — blocks edits before they happen. Supported by opencode via the `automaton-guard` plugin.
2. **Git pre-commit hook** — blocks commits when no task is in `implement` or `doc_review` phase. Works for ALL harnesses.
3. **Prompt-based rules** (ALLOWED/FORBIDDEN sections) — advisory only, relies on agent discipline.
See `contracts/harness-integration.md` for integration details.
### Multi-Agent (Optional)
Add an `Agent Configuration` section to `.agent.md` to enable multi-agent mode:
```markdown
## Agent Configuration
Mode: multi-agent
Agents:
- id: researcher
phases: [research, decomposition, design, test_design]
- id: implementer
phases: [implement]
- id: bug-hunter
phases: [bug_find, adversarial_bug_find]
- id: referee
phases: [referee]
- id: orchestrator
phases: [new, complete, human_intervention]
role: coordinator
Lock timeout: 30m
```
In multi-agent mode, agents claim tasks and discover work via `status.py --claim` and `--next-available`. In single-agent mode (the default), these commands are no-ops.
## Layered File System
The framework uses a **layered approach** to file management, with a clear precedence:
1. **Project overrides** (highest precedence): `{project}/.automaton/` — contains project-specific customizations
2. **Global framework** (default): `~/.automaton/` — contains the base framework files
**Precedence rule**: If a file exists in the project's `.automaton/` directory, the Orchestrator reads it from there. If it doesn't exist, the Orchestrator reads it from the global `~/.automaton/` directory.
### What files belong in each layer?
- **Project's `.automaton/`**: .agent.md (project-specific settings like Autopilot mode, rules override), .rules.md (project-specific constraints), extensions/ (optional additive overrides)
- **Global `~/.automaton/`**: All prompt files, contracts, scripts, config.md, workflow.md
### Upgrading
#### Updating the framework
The framework reads all base files from `~/.automaton/` at runtime. To update the framework:
```bash
cd ~/.automaton && git pull
```
This automatically applies changes to all projects — no per-project file update needed.
#### Upgrading an existing project to v2.0
If a project was created before v2.0 state enforcement (`.state` files), it needs an upgrade to bootstrap `.state` files and install the pre-commit hook:
```bash
# From the project root:
bash ~/.automaton/scripts/upgrade.sh /path/to/project
```
This will:
1. Bootstrap `.state` files for all existing tasks (inferring phase from artifacts)
2. Add a version marker to `~/.automaton/config.md`
3. Install the git pre-commit hook (blocks commits without a task in implement/doc_review)
You can also upgrade tasks individually:
```bash
python3 ~/.automaton/scripts/status.py --upgrade --task my-task --project /path/to/project
```
#### Installing the pre-commit hook manually
If you skipped the upgrade script or are setting up a new project:
```bash
# From the project root:
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit
```
To verify the hook is working:
```bash
python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project
# Should return exit code 1 (DENIED) if no tasks are in implement/doc_review
```
If a project has stale framework file copies (from the old model), tell the agent:
> "Upgrade automaton for this project."
The agent will run `migrate-project.sh` to clean up stale files and move customizations to `extensions/`.
## Dashboard
The dashboard provides a web-based Kanban board, statistics, and timeline views for monitoring task progress.
```bash
# Start from any project root or ~/.automaton/
python3 -m automaton.dashboard
# Or use the convenience wrapper
bash ~/.automaton/scripts/dashboard.sh
```
See `automaton/dashboard/README.md` for full documentation on views, keyboard shortcuts, configuration, and scope detection.
## Contact & Support
[Insert Contact Info]