Lap Tran 7336db282d
CI / build (push) Has been cancelled
Bootstrap self-improvement loop, decompose model-divergence-enforcement
Self-improvement loop:
- Created via --create-loop --from-template self-improvement
- Scheduled via launchd (3600s interval)
- State: running

model-divergence-enforcement task:
- Research approved, decomposed into 3 sequential subtasks:
  1. mde-manifest-detection (models.json + detect_models.py)
  2. mde-interactive-enforcement (conflict matrix + --model args + audit)
  3. mde-loop-enforcement (loop.json roles + {model} substitution + check-gate)
- Parent at decomposition:approved (stays active until subtasks complete)
- Each subtask has BRIEF.md with scope, deliverables, acceptance criteria
2026-06-25 07:15:25 -04:00
…

Automaton

A contract-based operating system for LLM agents, designed to enforce disciplined engineering workflows.

1. Global Framework Installation (One-time setup)

Before you can use the framework in any project, you must install the core logic into your local environment.

Run these commands in your terminal:

# Clone the framework into the global config directory
git clone <your-git-url> ~/.automaton

# Enter the directory
cd ~/.automaton

# Make the installation script executable and run it
# You must provide the git URL as the first argument
chmod +x install.sh
./install.sh <your-git-url>

The git URL is required because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling.

Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory. A self-improvement loop is created and scheduled by default (see Loop Engineering below).

Updating the Framework

When the framework is updated, you can update your global installation:

cd ~/.automaton
./update.sh

This will:

  • Fetch the latest changes from the repository
  • Check for uncommitted changes and warn you
  • Pull the latest updates

Upgrading existing projects: The framework reads prompts, contracts, and scripts from ~/.automaton/ at runtime. Projects only override .agent.md and .rules.md. This means updating the global framework (git pull) automatically applies to all projects. No per-project upgrade is needed.

If a project was set up under the old model (with copies of framework files), it needs migration first. Tell the agent: "Upgrade automaton for this project" to run the migration.


2. Project Setup (Per project)

Once the framework is installed globally, you must "onboard" every individual project you work on.

If you want the agent to handle the configuration for you, navigate to your project root and run:

*"Onboard this project into automaton."

The agent will automatically:

  1. Detect your project type (New, Existing, or Upgrade).
  2. Create the ./.automaton/ directory.
  3. Generate your .agent.md and .rules.md files.
  4. Initiate the "Exploration Ritual" to understand your codebase.

Option B: The Manual Way

If you prefer to set it up manually, create a .automaton/ directory in your project root and add:

  • .agent.md: Project-specific configuration (Mode, rules, etc.).
  • .rules.md: Project-specific constraints and past failure modes.

Then install the git pre-commit hook:

ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit

This hook blocks commits when no task is in implement or doc_review phase, preventing accidental code changes without a tracked task.


The Autopilot Workflow

The framework features an Autopilot mode that allows the agent to drive a project to completion with minimal intervention.

Lifecycle of a Task

  1. Research: Produce a SPEC.md (Contract). No code allowed. Interactive — agent grills you for requirements and gets sign-off.
  2. Decomposition (optional): Break the task into sub-tasks sized for your VRAM. Interactive — agent analyzes the spec and proposes sub-tasks with token budget estimates, gets sign-off. See VRAM Configuration below.
  3. Design (optional): Produce a DESIGN.md (Architecture). No code allowed. Interactive — agent grills you for design decisions and gets sign-off.
  4. Test Design (optional): Produce a TEST_PLAN.md (Test specification). No code allowed. Interactive — agent grills you for test coverage and edge cases, then presents draft test cases for review and sign-off.
  5. Implement: Write code and tests based only on the SPEC.md, DESIGN.md (if present), and TEST_PLAN.md (if present). Follow TDD (Red/Green/Refactor). The TEST_PLAN.md (if present) serves as the test specification the implementer follows.
  6. Bug Find: Aggressive search for bugs and spec deviations.
  7. Adversarial Bug Find: Deep search for complex logic errors, race conditions, and performance issues.
  8. Doc Review: Review documentation against DESIGN.md plan and fix missing docs.
  9. Referee: Objective evaluation of all bugs, docs, and the final verdict.

How to use Autopilot

Autopilot mode (default)

The default mode is Autopilot: Enabled. The Orchestrator automatically drives tasks through all phases. Just say:

  • Start a new task: "Research add user authentication" (the Orchestrator creates the task and drives it all the way to completion)
  • Decompose a task: "Decompose the add-user-auth task" (the Orchestrator breaks it into sub-tasks sized for your VRAM)
  • Continue: "orchestrate" or "continue" (the Orchestrator finds the most advanced task and drives it)

VRAM Configuration

For low-VRAM systems (8GB, 16GB), set your VRAM limits in ~/.automaton/config.md:

## VRAM Configuration
- **Auto-detect**: Yes  # Let the agent detect your VRAM automatically
- **Target VRAM context**: 16k tokens  # Override auto-detect if needed
- **Headroom**: 25%
- **Max peak context per sub-task**: 12k tokens

Auto-detection: When Auto-detect: Yes, the Orchestrator runs scripts/vram_detect.py to probe:

  • GPU VRAM (via nvidia-smi)
  • System RAM (via free)
  • Model context window (from config.md or API config files)
  • Framework overhead (by counting token load in loaded prompts)

Manual override: When Auto-detect: No, use the manually specified values:

## VRAM Configuration
- **Auto-detect**: No
- **Target VRAM context**: 8k
- **Headroom**: 30%
- **Max peak context per sub-task**: 5.6k

Model Configuration

When using a local LLM or a specific API model, set the model in ~/.automaton/config.md:

## Model Configuration
- **Model**: auto  # Use auto-detection from API config files
- **Override context window**: auto  # Override auto-detection, or specify (e.g., 128k, 200k)

Auto-detection: When Model: auto, the framework detects the model name from API config files (.env, config.yaml, etc.) and looks up its context window.

Manual override: When you know your model name, specify it:

## Model Configuration
- **Model**: gpt-4o
- **Override context window**: 128k

When you run "Decompose the X task", the Orchestrator will:

  1. Analyze the task's SPEC.md
  2. Detect VRAM limits (auto or manual)
  3. Break it into sub-tasks, each sized to fit within your VRAM limit
  4. Estimate the token budget for each sub-task
  5. Create sub-task folders under tasks/{parent-task}/subtasks/{sub-task}/
  6. Propagate VRAM config to each sub-task

Sub-tasks run independently through the full lifecycle. The parent task is complete only when ALL sub-tasks pass.

Manual mode (opt-in)

Set Autopilot: Disabled in your project's .automaton/.agent.md if you prefer to manually run each phase. The Orchestrator reports the current state and tells you the next command. Then run phases by saying things like:

  • "Research add user authentication" — starts a new task
  • "Decompose the add-user-auth task" — breaks into VRAM-sized sub-tasks (optional)
  • "Design the add-user-auth task" — designs the architecture (optional)
  • "Design tests for the add-user-auth task" — designs test cases (optional)
  • "Implement the add-user-auth task" — implements the task with tests
  • "Find bugs in the add-user-auth task" — finds bugs
  • "Perform adversarial bug find for add-user-auth" — deep bug search
  • "Review docs for the add-user-auth task" — reviews documentation
  • "Review the add-user-auth task" — referee evaluates
  • "orchestrate" — asks the Orchestrator what to do next

Sub-Task Management

When a task is decomposed, the Orchestrator creates sub-tasks under tasks/{parent-task}/subtasks/. Each sub-task:

  • Has its own full lifecycle (Research → Decomposition → Design → ... → Complete)
  • Starts at the Research phase with an empty folder
  • Receives a PARENT_SPEC.md with the parent task's context
  • Receives a VRAM_CONFIG.md with the VRAM constraints
  • Runs independently — sub-tasks in the same wave can run in parallel
  • The parent task is NOT complete until ALL sub-tasks pass

Key Components

  • .agent.md: Project-specific agent behavior (Autopilot mode, routing rules, optional Agent Configuration for multi-agent).
  • config.md: Global framework settings (VRAM, model, system requirements, version).
  • .rules.md: Living document of project constraints and past failure modes.
  • prompts/: Specialized system prompts for each phase with ALLOWED/FORBIDDEN sections and approval gates.
  • workflow.md: The state machine governing the task lifecycle, with .state file as canonical phase indicator.
  • scripts/status.py: Enforcement script — task status, phase transitions, approval gates, folder validation, audits, pre-edit hook (--can-edit), multi-agent claiming.
  • scripts/vram_detect.py: Auto-detects GPU VRAM, RAM, model context window, and framework overhead.
  • scripts/git-hooks/pre-commit: Blocks commits when no task is in an edit-allowed phase.
  • contracts/harness-integration.md: Integration contract for agent harnesses (opencode, aider, etc.).
  • plugins/automaton-guard/: opencode plugin that intercepts edit/write calls and checks --can-edit before allowing them.

Loop Engineering (beta, v1)

Automaton can run unattended workflow loops: each loop has an OS-level schedule (launchd / cron / schtasks) and is constrained by 6 brake gates enforced in status.py. The single source of truth for loop runtime state is .state.loop per loop at {project}/.automaton/loops/<name>/.state.loop.

# Create a loop from a template (only way to bootstrap)
python3 ~/.automaton/scripts/status.py --create-loop my-ci-triage --from-template ci-triage --project /path/to/project

# Install the native OS schedule unit (launchd/cron/schtasks)
python3 ~/.automaton/scripts/status.py --install-schedule my-ci-triage --interval 3600 --project /path/to/project

# Pre-tick gate check (6 brakes; first failure halts the loop)
python3 ~/.automaton/scripts/status.py --check-gate my-ci-triage --project /path/to/project

# Clear a halt (only way; no auto-approve in v1)
python3 ~/.automaton/scripts/status.py --approve --loop my-ci-triage --project /path/to/project

# List all loops and their status
python3 ~/.automaton/scripts/status.py --loop-list --project /path/to/project

Loop states: running, halted, paused, complete. Halt reasons: iterations_exhausted, budget_exhausted, verifier_failed, drift_detected, human_intervention. The per-tick engine is scripts/loop-runner.py --mode tick; it gates first, spawns Implement / Verify / Orchestrate role sessions, and writes .state.loop atomically. Concurrent ticks on the same loop and concurrent --pause-loop / --approve --loop writes are serialized via a cross-process file lock on <loop_path>/.state.lock (POSIX fcntl.flock, Windows msvcrt.locking); see design/loops/technical.md §7 "Lock serialization". See design/loops/technical.md §7 for the full 11-step flow.

Quick Start

# 1. Create a loop from a template
python3 ~/.automaton/scripts/status.py --create-loop my-ci-triage \
  --from-template ci-triage --project /path/to/project

# 2. Install the OS schedule (launchd on macOS, cron on Linux, schtasks on Windows)
python3 ~/.automaton/scripts/status.py --install-schedule my-ci-triage \
  --interval 3600 --project /path/to/project

# 3. Monitor
python3 ~/.automaton/scripts/status.py --loop-list --project /path/to/project
python3 ~/.automaton/scripts/status.py --audit --project /path/to/project

Tick Cycle

Each tick runs this 11-step flow (see design/loops/technical.md §7 for details):

gate check -> find work -> ensure worktree -> spawn Implement -> spawn Verify
-> parse verdict -> spawn Orchestrate -> atomic state write -> log

The runner resolves prompt files from loop.json roles.*.prompt (e.g. loop-implement.md), substitutes content-level tokens ({task_brief}, {acceptance_criteria}, {next_hint}, {artifact_content}, etc.), writes the resolved prompt to outputs/tickN-<role>-prompt.md, and passes it to the harness.

Configuration (loop.json)

Field Description
name Loop name (kebab-case)
schedule.interval_seconds Tick interval for daemon mode
brakes.max_iterations Max ticks before halt
brakes.max_budget_usd Optional USD budget cap (null = unlimited)
brakes.score_plateau_window Score plateau detection window
blast_radius.file_scope List of paths the loop may edit
blast_radius.use_worktree If true, tick runs in a per-loop git worktree
work_source.kind single, audit, or backlog
work_source.area Design area for backlog kind (default "loops"; "framework" reads design/framework/BACKLOG.md)
roles.implement.prompt Prompt file for Implement role
roles.verify.prompt Prompt file for Verify role
roles.orchestrate.prompt Prompt file for Orchestrate role
harness.command Command template with {prompt}, {prompt_content}, {cwd} tokens (default invokes opencode run --dir <cwd> <prompt>; override for other harnesses -- Pi Dev, aider, etc.)
acceptance_criteria List of criteria for the verifier to check

Monitoring

  • --loop-list: show all loops and their status
  • --audit: check for violations across all tasks and loops
  • .state.log: per-loop tick log (ISO-timestamped entries)
  • outputs/: per-tick artifacts and resolved prompts

Halt and Resume

# Pause a loop (disables the OS schedule unit)
python3 ~/.automaton/scripts/status.py --pause-loop my-ci-triage --project /path/to/project

# Resume a paused loop
python3 ~/.automaton/scripts/status.py --resume-loop my-ci-triage --project /path/to/project

# Clear a halt (the only way; no auto-approve in v1)
python3 ~/.automaton/scripts/status.py --approve --loop my-ci-triage --project /path/to/project

Self-Improvement Loop (Default-On)

The framework installs a self-improvement loop by default at install time. This loop ticks against status.py --audit on the framework's own repo, picking up audit violations and resolving them unattended. It runs every 3600 seconds (1 hour) with max_iterations: 10 and a score plateau window of 3.

# Disable the self-improvement loop
python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/

# Re-enable it
python3 ~/.automaton/scripts/status.py --resume-loop self-improvement --project ~/.automaton/

The loop uses a git worktree at ~/.automaton/loops/self-improvement/worktree/ and is scoped to scripts/, prompts/, tests/, and design/ directories.

State Enforcement (v2.0)

Automaton v2.0 enforces the state machine computationally, not just via prompts:

  • .state file: Each task has a .state file that is the single source of truth for its current phase
  • status.py --transition: All phase transitions must go through this command; illegal transitions are refused
  • Approval gates: Research, Design, Decomposition, and Test Design phases require explicit user approval before proceeding
  • status.py --validate-folder: Detects out-of-order artifacts (phase skipping)
  • status.py --audit: Comprehensive audit across all tasks for violations
  • status.py --create-task: The only valid way to create task folders
  • FORBIDDEN actions in prompts: Each phase prompt explicitly lists what agents cannot do

Quick Reference

# Create a new task
python3 ~/.automaton/scripts/status.py --create-task add-user-auth --project /path/to/project

# Check task status
python3 ~/.automaton/scripts/status.py --task add-user-auth --project /path/to/project

# List all tasks
python3 ~/.automaton/scripts/status.py --list --project /path/to/project

# Transition to next phase
python3 ~/.automaton/scripts/status.py --transition research --task add-user-auth --project /path/to/project
python3 ~/.automaton/scripts/status.py --transition research:awaiting_approval --task add-user-auth --project /path/to/project

# Approve a phase (after user sign-off)
python3 ~/.automaton/scripts/status.py --approve --task add-user-auth --project /path/to/project

# Validate task folder
python3 ~/.automaton/scripts/status.py --validate-folder --task add-user-auth --project /path/to/project

# Audit all tasks
python3 ~/.automaton/scripts/status.py --audit --project /path/to/project

# Upgrade pre-v2.0 tasks (bootstrap .state files)
python3 ~/.automaton/scripts/status.py --upgrade --project /path/to/project

# Check if code edits are allowed (harness integration)
python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project
python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project --file src/main.py
python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project --task add-user-auth --json

Important: Always pass --project to ensure correct scoping when multiple projects exist. Without it, status.py resolves the project from the current directory and errors if not in a project.

Untracked Tasks

Tasks without .state files are UNTRACKED — all commands (--transition, --can-edit, --task, --approve) refuse to operate on them. This prevents agents from working on tasks created before v2.0 state enforcement.

To fix untracked tasks:

# Upgrade a single task
python3 ~/.automaton/scripts/status.py --upgrade --task my-old-task --project /path/to/project

# Upgrade all tasks at once
python3 ~/.automaton/scripts/status.py --upgrade --project /path/to/project

Enforcement

The framework enforces the state machine computationally. No phase can be skipped, no approval can be bypassed, and no code edits can happen without a task in an edit-allowed phase. This is enforced through three layers:

  1. Harness pre-edit hook (--can-edit) — blocks edits before they happen. Supported by opencode via the automaton-guard plugin.
  2. Git pre-commit hook — blocks commits when no task is in implement or doc_review phase. Works for ALL harnesses.
  3. Prompt-based rules (ALLOWED/FORBIDDEN sections) — advisory only, relies on agent discipline.

See contracts/harness-integration.md for integration details.

Multi-Agent (Optional)

Add an Agent Configuration section to .agent.md to enable multi-agent mode:

## Agent Configuration
Mode: multi-agent
Agents:
  - id: researcher
    phases: [research, decomposition, design, test_design]
  - id: implementer
    phases: [implement]
  - id: bug-hunter
    phases: [bug_find, adversarial_bug_find]
  - id: referee
    phases: [referee]
  - id: orchestrator
    phases: [new, complete, human_intervention]
    role: coordinator
Lock timeout: 30m

In multi-agent mode, agents claim tasks and discover work via status.py --claim and --next-available. In single-agent mode (the default), these commands are no-ops.

Layered File System

The framework uses a layered approach to file management, with a clear precedence:

  1. Project overrides (highest precedence): {project}/.automaton/ — contains project-specific customizations
  2. Global framework (default): ~/.automaton/ — contains the base framework files

Precedence rule: If a file exists in the project's .automaton/ directory, the Orchestrator reads it from there. If it doesn't exist, the Orchestrator reads it from the global ~/.automaton/ directory.

What files belong in each layer?

  • Project's .automaton/: .agent.md (project-specific settings like Autopilot mode, rules override), .rules.md (project-specific constraints), extensions/ (optional additive overrides)
  • Global ~/.automaton/: All prompt files, contracts, scripts, config.md, workflow.md

Upgrading

Updating the framework

The framework reads all base files from ~/.automaton/ at runtime. To update the framework:

cd ~/.automaton && git pull

This automatically applies changes to all projects — no per-project file update needed.

Upgrading an existing project to v2.0

If a project was created before v2.0 state enforcement (.state files), it needs an upgrade to bootstrap .state files and install the pre-commit hook:

# From the project root:
bash ~/.automaton/scripts/upgrade.sh /path/to/project

This will:

  1. Bootstrap .state files for all existing tasks (inferring phase from artifacts)
  2. Add a version marker to ~/.automaton/config.md
  3. Install the git pre-commit hook (blocks commits without a task in implement/doc_review)

You can also upgrade tasks individually:

python3 ~/.automaton/scripts/status.py --upgrade --task my-task --project /path/to/project

Installing the pre-commit hook manually

If you skipped the upgrade script or are setting up a new project:

# From the project root:
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit

To verify the hook is working:

python3 ~/.automaton/scripts/status.py --can-edit --project /path/to/project
# Should return exit code 1 (DENIED) if no tasks are in implement/doc_review

If a project has stale framework file copies (from the old model), tell the agent:

"Upgrade automaton for this project."

The agent will run migrate-project.sh to clean up stale files and move customizations to extensions/.

Dashboard

The dashboard provides a web-based Kanban board, statistics, and timeline views for monitoring task progress.

# Start from any project root or ~/.automaton/
python3 -m automaton.dashboard

# Or use the convenience wrapper
bash ~/.automaton/scripts/dashboard.sh

See automaton/dashboard/README.md for full documentation on views, keyboard shortcuts, configuration, and scope detection.

Contact & Support

[Insert Contact Info]

S
Description
Automaton - AI coding agent framework
Readme
2.3 MiB
Languages
Python 82.4%
JavaScript 5.5%
Shell 4.9%
CSS 3.4%
TypeScript 3.1%
Other 0.7%