- Insert code_review phase between implement and bug_find - Approval gate: code_review:awaiting_approval → code_review:approved - Read-only phase — no edits, no fixes, no returning to implement - Reviewer≠implementer: .state.implementer tracking + --claim enforcement - Structured CODE_REVIEW.md: spec compliance, design conformance, quality scorecard, items found (severity/category/location/resolution), test coverage - Updated status.py (10 data structures), dashboard (4 files), prompts (3 files), agent routing, tests (6 new test classes, 19 new tests)
Automaton
A contract-based operating system for LLM agents, designed to enforce disciplined engineering workflows.
1. Global Framework Installation (One-time setup)
Before you can use the framework in any project, you must install the core logic into your local environment.
Run these commands in your terminal:
# Clone the framework into the global config directory
git clone http://10.37.0.86:3003/hermes/automaton ~/.automaton
# Enter the directory
cd ~/.automaton
# Make the installation script executable and run it
chmod +x install.sh
./install.sh
Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory.
Updating the Framework
When the framework is updated, you can update your global installation:
cd ~/.automaton
./update.sh
This will:
- Fetch the latest changes from the repository
- Check for uncommitted changes and warn you
- Pull the latest updates
Upgrading existing projects: The framework reads prompts, contracts, and scripts from ~/.automaton/ at runtime. Projects only override .agent.md and .rules.md. This means updating the global framework (git pull) automatically applies to all projects. No per-project upgrade is needed.
If a project was set up under the old model (with copies of framework files), it needs migration first. Tell the agent: "Upgrade automaton for this project" to run the migration.
2. Project Setup (Per project)
Once the framework is installed globally, you must "onboard" every individual project you work on.
Option A: The Agent-Driven Way (Recommended)
If you want the agent to handle the configuration for you, navigate to your project root and run:
*"Onboard this project into automaton."
The agent will automatically:
- Detect your project type (New, Existing, or Upgrade).
- Create the
./.automaton/directory. - Generate your
.agent.mdand.rules.mdfiles. - Initiate the "Exploration Ritual" to understand your codebase.
Option B: The Manual Way
If you prefer to set it up manually, create a .automaton/ directory in your project root and add:
.agent.md: Project-specific configuration (Mode, rules, etc.)..rules.md: Project-specific constraints and past failure modes.
Then install the git pre-commit hook:
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit
This hook blocks commits when no task is in implement or doc_review phase, preventing accidental code changes without a tracked task.
The Autopilot Workflow
The framework features an Autopilot mode that allows the agent to drive a project to completion with minimal intervention.
Lifecycle of a Task
- Research: Produce a
SPEC.md(Contract). No code allowed. Interactive — agent grills you for requirements and gets sign-off. - Decomposition (optional): Break the task into sub-tasks sized for your VRAM. Interactive — agent analyzes the spec and proposes sub-tasks with token budget estimates, gets sign-off. See VRAM Configuration below.
- Design (optional): Produce a
DESIGN.md(Architecture). No code allowed. Interactive — agent grills you for design decisions and gets sign-off. - Test Design (optional): Produce a
TEST_PLAN.md(Test specification). No code allowed. Interactive — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off. - Implement: Write code and tests based only on the
SPEC.md,DESIGN.md(if present), andTEST_PLAN.md(if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows. - Bug Find: Aggressive search for bugs and spec deviations.
- Adversarial Bug Find: Deep search for complex logic errors, race conditions, and performance issues.
- Doc Review: Review documentation against DESIGN.md plan and fix missing docs.
- Referee: Objective evaluation of all bugs, docs, and the final verdict.
How to use Autopilot
Autopilot mode (default)
The default mode is Autopilot: Enabled. The Orchestrator automatically drives tasks through all phases. Just say:
- Start a new task: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion)
- Decompose a task: "Decompose the add-user-auth task" (the Orchestrator breaks it into sub-tasks sized for your VRAM)
- Continue: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it)
VRAM Configuration
For low-VRAM systems (8GB, 16GB), set your VRAM limits in ~/.automaton/config.md:
## VRAM Configuration
- **Auto-detect**: Yes # Let the agent detect your VRAM automatically
- **Target VRAM context**: 16k tokens # Override auto-detect if needed
- **Headroom**: 25%
- **Max peak context per sub-task**: 12k tokens
Auto-detection: When Auto-detect: Yes, the Orchestrator runs scripts/vram_detect.py to probe:
- GPU VRAM (via
nvidia-smi) - System RAM (via
free) - Model context window (from config.md or API config files)
- Framework overhead (by counting token load in loaded prompts)
Manual override: When Auto-detect: No, use the manually specified values:
## VRAM Configuration
- **Auto-detect**: No
- **Target VRAM context**: 8k
- **Headroom**: 30%
- **Max peak context per sub-task**: 5.6k
Model Configuration
When using a local LLM or a specific API model, set the model in ~/.automaton/config.md:
## Model Configuration
- **Model**: auto # Use auto-detection from API config files
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
Auto-detection: When Model: auto, the framework detects the model name from API config files (.env, config.yaml, etc.) and looks up its context window.
Manual override: When you know your model name, specify it:
## Model Configuration
- **Model**: gpt-4o
- **Override context window**: 128k
When you run "Decompose the X task", the Orchestrator will:
- Analyze the task's SPEC.md
- Detect VRAM limits (auto or manual)
- Break it into sub-tasks, each sized to fit within your VRAM limit
- Estimate the token budget for each sub-task
- Create sub-task folders under
tasks/{parent-task}/subtasks/{sub-task}/ - Propagate VRAM config to each sub-task
Sub-tasks run independently through the full lifecycle. The parent task is complete only when ALL sub-tasks pass.
Manual mode (opt-in)
Set Autopilot: Disabled in your project's .automaton/.agent.md if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like:
- "Research add user authentication" — starts a new task
- "Decompose the add-user-auth task" — breaks into VRAM-sized sub-tasks (optional)
- "Design the add-user-auth task" — designs the architecture (optional)
- "Design tests for the add-user-auth task" — designs test cases (optional)
- "Implement the add-user-auth task" — implements the task with tests
- "Find bugs in the add-user-auth task" — finds bugs
- "Perform adversarial bug find for add-user-auth" — deep bug search
- "Review docs for the add-user-auth task" — reviews documentation
- "Review the add-user-auth task" — referee evaluates
- "orchestrate" — asks the Orchestrator what to do next
Sub-Task Management
When a task is decomposed, the Orchestrator creates sub-tasks under tasks/{parent-task}/subtasks/. Each sub-task:
- Has its own full lifecycle (Research → Decomposition → Design → ... → Complete)
- Starts at the Research phase with an empty folder
- Receives a
PARENT_SPEC.mdwith the parent task's context - Receives a
VRAM_CONFIG.mdwith the VRAM constraints - Runs independently — sub-tasks in the same wave can run in parallel
- The parent task is NOT complete until ALL sub-tasks pass
Key Components
.agent.md: Project-specific agent behavior (Autopilot mode, routing rules, optional Agent Configuration for multi-agent).config.md: Global framework settings (VRAM, model, system requirements, version)..rules.md: Living document of project constraints and past failure modes.prompts/: Specialized system prompts for each phase with ALLOWED/FORBIDDEN sections and approval gates.workflow.md: The state machine governing the task lifecycle, with.statefile as canonical phase indicator.scripts/status.py: Enforcement script — task status, phase transitions, approval gates, folder validation, audits, pre-edit hook (--can-edit), multi-agent claiming.scripts/vram_detect.py: Auto-detects GPU VRAM, RAM, model context window, and framework overhead.scripts/git-hooks/pre-commit: Blocks commits when no task is in an edit-allowed phase.contracts/harness-integration.md: Integration contract for agent harnesses (opencode, aider, etc.).plugins/automaton-guard/: opencode plugin that interceptsedit/writecalls and checks--can-editbefore allowing them.
State Enforcement (v2.0)
Automaton v2.0 enforces the state machine computationally, not just via prompts:
.statefile: Each task has a.statefile that is the single source of truth for its current phasestatus.py --transition: All phase transitions must go through this command; illegal transitions are refused- Approval gates: Research, Design, Decomposition, and Test Design phases require explicit user approval before proceeding
status.py --validate-folder: Detects out-of-order artifacts (phase skipping)status.py --audit: Comprehensive audit across all tasks for violationsstatus.py --create-task: The only valid way to create task folders- FORBIDDEN actions in prompts: Each phase prompt explicitly lists what agents cannot do
Quick Reference
# Create a new task
python ~/.automaton/scripts/status.py --create-task add-user-auth --project /path/to/project
# Check task status
python ~/.automaton/scripts/status.py --task add-user-auth --project /path/to/project
# List all tasks
python ~/.automaton/scripts/status.py --list --project /path/to/project
# Transition to next phase
python ~/.automaton/scripts/status.py --transition research --task add-user-auth --project /path/to/project
python ~/.automaton/scripts/status.py --transition research:awaiting_approval --task add-user-auth --project /path/to/project
# Approve a phase (after user sign-off)
python ~/.automaton/scripts/status.py --approve --task add-user-auth --project /path/to/project
# Validate task folder
python ~/.automaton/scripts/status.py --validate-folder --task add-user-auth --project /path/to/project
# Audit all tasks
python ~/.automaton/scripts/status.py --audit --project /path/to/project
# Upgrade pre-v2.0 tasks (bootstrap .state files)
python ~/.automaton/scripts/status.py --upgrade --project /path/to/project
# Check if code edits are allowed (harness integration)
python ~/.automaton/scripts/status.py --can-edit --project /path/to/project
python ~/.automaton/scripts/status.py --can-edit --project /path/to/project --file src/main.py
python ~/.automaton/scripts/status.py --can-edit --project /path/to/project --task add-user-auth --json
Important: Always pass --project to ensure correct scoping when multiple projects exist. Without it, status.py resolves the project from the current directory and errors if not in a project.
Untracked Tasks
Tasks without .state files are UNTRACKED — all commands (--transition, --can-edit, --task, --approve) refuse to operate on them. This prevents agents from working on tasks created before v2.0 state enforcement.
To fix untracked tasks:
# Upgrade a single task
python ~/.automaton/scripts/status.py --upgrade --task my-old-task --project /path/to/project
# Upgrade all tasks at once
python ~/.automaton/scripts/status.py --upgrade --project /path/to/project
Enforcement
The framework enforces the state machine computationally. No phase can be skipped, no approval can be bypassed, and no code edits can happen without a task in an edit-allowed phase. This is enforced through three layers:
- Harness pre-edit hook (
--can-edit) — blocks edits before they happen. Supported by opencode via theautomaton-guardplugin. - Git pre-commit hook — blocks commits when no task is in
implementordoc_reviewphase. Works for ALL harnesses. - Prompt-based rules (ALLOWED/FORBIDDEN sections) — advisory only, relies on agent discipline.
See contracts/harness-integration.md for integration details.
Multi-Agent (Optional)
Add an Agent Configuration section to .agent.md to enable multi-agent mode:
## Agent Configuration
Mode: multi-agent
Agents:
- id: researcher
phases: [research, decomposition, design, test_design]
- id: implementer
phases: [implement]
- id: bug-hunter
phases: [bug_find, adversarial_bug_find]
- id: referee
phases: [referee]
- id: orchestrator
phases: [new, complete, human_intervention]
role: coordinator
Lock timeout: 30m
In multi-agent mode, agents claim tasks and discover work via status.py --claim and --next-available. In single-agent mode (the default), these commands are no-ops.
Layered File System
The framework uses a layered approach to file management, with a clear precedence:
- Project overrides (highest precedence):
{project}/.automaton/— contains project-specific customizations - Global framework (default):
~/.automaton/— contains the base framework files
Precedence rule: If a file exists in the project's .automaton/ directory, the Orchestrator reads it from there. If it doesn't exist, the Orchestrator reads it from the global ~/.automaton/ directory.
What files belong in each layer?
- Project's
.automaton/: .agent.md (project-specific settings like Autopilot mode, rules override), .rules.md (project-specific constraints), extensions/ (optional additive overrides) - Global
~/.automaton/: All prompt files, contracts, scripts, config.md, workflow.md
Upgrading
Updating the framework
The framework reads all base files from ~/.automaton/ at runtime. To update the framework:
cd ~/.automaton && git pull
This automatically applies changes to all projects — no per-project file update needed.
Upgrading an existing project to v2.0
If a project was created before v2.0 state enforcement (.state files), it needs an upgrade to bootstrap .state files and install the pre-commit hook:
# From the project root:
bash ~/.automaton/scripts/upgrade.sh /path/to/project
This will:
- Bootstrap
.statefiles for all existing tasks (inferring phase from artifacts) - Add a version marker to
~/.automaton/config.md - Install the git pre-commit hook (blocks commits without a task in implement/doc_review)
You can also upgrade tasks individually:
python ~/.automaton/scripts/status.py --upgrade --task my-task --project /path/to/project
Installing the pre-commit hook manually
If you skipped the upgrade script or are setting up a new project:
# From the project root:
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit
To verify the hook is working:
python ~/.automaton/scripts/status.py --can-edit --project /path/to/project
# Should return exit code 1 (DENIED) if no tasks are in implement/doc_review
If a project has stale framework file copies (from the old model), tell the agent:
"Upgrade automaton for this project."
The agent will run migrate-project.sh to clean up stale files and move customizations to extensions/.
Dashboard
The dashboard provides a web-based Kanban board, statistics, and timeline views for monitoring task progress.
# Start from any project root or ~/.automaton/
python -m automaton.dashboard
# Or use the convenience wrapper
bash ~/.automaton/scripts/dashboard.sh
See automaton/dashboard/README.md for full documentation on views, keyboard shortcuts, configuration, and scope detection.
Contact & Support
[Insert Contact Info]