- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
16 KiB
SPEC: Multi-Agent Support
Goal
Add opt-in multi-agent coordination to Automaton so that specialized agents (researcher, implementer, bug finder) can claim and work on different tasks or phases simultaneously. Single-agent mode must remain the default with zero configuration changes.
Background
Automaton currently assumes one agent that switches between personas sequentially. The .state file and status.py foundation from the state-file-enforcement and status-script specs provide shared state and validation — but they lack three things needed for multi-agent:
- Claiming — no way to mark a task as "owned" by a specific agent, so two agents can race on the same task
- Role binding — no way to restrict an agent to specific phases (e.g., "Agent A only does research")
- Work discovery — no way for an idle agent to find available work matching its role
These are only needed when multiple agents are running. In single-agent mode (the default), none of this machinery activates.
Design Principle: Single-Agent Is the Zero-Config Default
When no multi-agent configuration exists:
- No
.state.lockfiles are ever created status.pyworks exactly as specified in the status-script spec- The orchestrator drives the full lifecycle in one session (current behavior)
- All claiming/releasing/work-queue commands are no-ops or return "single-agent mode"
- Zero performance overhead, zero behavioral change
Multi-agent activates only when ## Agent Configuration is present in .agent.md.
Requirements
1. Agent Configuration (.agent.md)
Add an optional section to .agent.md:
## Agent Configuration
Mode: multi-agent
Agents:
- id: researcher
phases: [research, decomposition, design, test_design]
- id: implementer
phases: [implement]
- id: bug-hunter
phases: [bug_find, adversarial_bug_find]
- id: doc-reviewer
phases: [doc_review]
- id: referee
phases: [referee]
- id: orchestrator
phases: [new, complete, human_intervention]
role: coordinator
Rules:
- If
## Agent Configurationis absent orMode: single-agent, everything works as today — no locks, no claiming, no work queue - If
Mode: multi-agent, the claiming/work-queue system activates - Agent
idvalues are free-form strings (alphanumeric + hyphens) - Each agent has an explicit list of phases it's allowed to work on
- Only one agent can have
role: coordinator— this is the orchestrator, which claims tasks, transitions state, and delegates - An agent can claim multiple phases
- Every phase must be covered by at least one agent (validated by
status.py) - Phases not listed under any agent are handled by the coordinator
Agent identity resolution (in priority order):
--agentflag onstatus.pycommands (e.g.,--agent implementer)AUTOMATON_AGENT_IDenvironment variableagent.idfield in the project's.agent.md- If none of the above: "default" (single-agent mode)
When --agent is provided, the agent must exist in the Agent Configuration. If not, status.py errors.
2. Task Claiming
python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {agent-id} [--project {project-path}]
Behavior in single-agent mode (no Agent Configuration):
- Output: "Claimed task '{task-name}' (single-agent mode — no lock needed)"
- No
.state.lockfile created - The command succeeds as a no-op
Behavior in multi-agent mode:
- Check if
.state.lockexists for the task - If no lock exists:
- Create
.state.lockatomically (write to.state.lock.tmp, rename to.state.lock) - Lock file format:
agent: {agent-id} phase: {current-phase} claimed: {ISO-8601-timestamp} expires: {ISO-8601-timestamp + lock-timeout} - Output: "Claimed task '{task-name}' for agent '{agent-id}' — phase: {phase}"
- Create
- If lock exists and not expired:
- Output: "ERROR: Task '{task-name}' is claimed by agent '{agent-id}' (expires: {timestamp}). Retry after expiry or release the claim."
- Exit code 1
- If lock exists and expired:
- Overwrite the lock with the new agent's claim
- Output: "WARN: Task '{task-name}' had stale lock from agent '{old-agent-id}' (expired: {timestamp}). Overclaiming for agent '{new-agent-id}'."
- Exit code 0
Phase validation on claim:
- The agent must be configured for the task's current phase
- If
agent-idis not allowed to work onresearchand the task is inresearchphase, refuse the claim - Output: "ERROR: Agent '{agent-id}' is not configured for phase '{phase}'. Allowed phases: {list}"
- Exit code 1
Lock timeout:
- Default: 30 minutes
- Configurable in
.agent.md:## Agent Configuration Mode: multi-agent Lock timeout: 60m - If
Lock timeoutis absent, default to 30 minutes - TTL values:
5m,10m,15m,30m,60m,120m
Lock file location: tasks/{task-name}/.state.lock
Sub-task claiming:
- Sub-tasks have their own
.state.lockin their own folder - Parent task lock is independent of sub-task locks
- Claiming a parent task does NOT claim its sub-tasks
3. Task Releasing
python ~/.automaton/scripts/status.py --release --task {task-name} --agent {agent-id} [--project {project-path}]
Behavior in single-agent mode:
- Output: "Released task '{task-name}' (single-agent mode — no lock to release)"
- No-op
Behavior in multi-agent mode:
- Check if
.state.lockexists for the task - If lock exists and owned by
{agent-id}:- Delete
.state.lock - Output: "Released task '{task-name}' from agent '{agent-id}'"
- Delete
- If lock exists but owned by a different agent:
- Output: "ERROR: Task '{task-name}' is claimed by agent '{other-agent-id}', not '{agent-id}'. Only the claiming agent can release."
- Exit code 1
- If no lock exists:
- Output: "WARN: Task '{task-name}' has no lock. Nothing to release."
- Exit code 0
Automatic release on phase transition:
When status.py --transition succeeds in multi-agent mode:
- The lock is updated to reflect the new phase (if the claiming agent is still valid for the new phase)
- If the claiming agent is NOT valid for the new phase, the lock is released automatically
- This enables handoff: researcher claims, does research, transitions to design, lock auto-releases, designer claims
4. Work Discovery
python ~/.automaton/scripts/status.py --next-available --agent {agent-id} [--project {project-path}]
Behavior in single-agent mode:
- Output: "single-agent mode — use --list to see all tasks"
Behavior in multi-agent mode:
- Scan all tasks in
{project}/.automaton/tasks/(including sub-tasks) - For each task:
- Read
.stateto determine current phase - Check if the task is unclaimed (no
.state.lock) or has an expired lock - Check if
{agent-id}is configured for the task's current phase
- Read
- Return the first available task sorted by priority:
- Task closest to completion first (referee > doc_review > adversarial_bug_find > bug_find > implement > test_design > design > decomposition > research > new)
- Within the same priority level, alphabetical by task name
- Output:
Next available task for agent 'implementer': Task: fix-login-bug Phase: implement Phase priority: 7 (high — close to completion) Status: unclaimed To claim: python ~/.automaton/scripts/status.py --claim --task fix-login-bug --agent implementer - If no tasks available:
- Output: "No tasks available for agent '{agent-id}'. All tasks are claimed by other agents, in phases this agent cannot work on, or completed."
Work queue (list all available):
python ~/.automaton/scripts/status.py --available --agent {agent-id} [--project {project-path}]
Same logic as --next-available but returns ALL matching tasks, not just the first:
Available tasks for agent 'implementer':
1. Task: fix-login-bug | Phase: implement | Status: unclaimed
2. Task: add-payment-api | Phase: implement | Status: claimed by 'other-agent' (expires: 2026-06-14T15:30:00Z)
5. Lock File Details
Format:
agent: {agent-id}
phase: {current-phase-from-state-file}
claimed: 2026-06-14T14:30:00Z
expires: 2026-06-14T15:00:00Z
Atomic writes: Same pattern as .state — write to .state.lock.tmp, then rename to .state.lock.
File classification: .state.lock is metadata (like .state), not a phase deliverable. It is excluded from --validate-folder checks and artifact heuristics.
Git: .state.lock should be in .gitignore (it's ephemeral agent state, not project state).
Lock expiry check: Every status.py command that reads .state.lock must also check expiry. Expired locks are treated as non-existent (available for claiming).
6. Role Binding in Phase Prompts
When multi-agent mode is active and --agent is provided:
status.py --task {task}output includes the agent's allowed phases:Task: add-user-auth Phase: research (from .state) Agent: researcher Agent allowed phases: research, decomposition, design, test_design ALLOWED for this agent: - Read project files, ask questions, write SPEC.md - Transition to decompose, design (if agent is configured for those phases) FORBIDDEN for this agent: - Edit code (implement phase) - Write BUG_REPORT.md (bug_find phase) - Write ADVERSARIAL_BUG_REPORT.md (adversarial_bug_find phase) - Write DOC_REVIEW.md (doc_review phase) - Write VERDICT.md (referee phase)- Phase prompts gain an additional section when the agent is role-bound:
## Agent Role You are agent '{agent-id}'. Your allowed phases are: {phases}. You may NOT perform actions from phases not in your allowed list. If the task requires a phase you cannot perform, output: "TASK_HANDOFF — task '{task-name}' requires phase '{required-phase}' which is outside my allowed phases. Agent '{allowed-agent-list}' can continue."
Single-agent mode: This section is absent. The agent has full access to all phases.
7. Coordinator Role
The orchestrator agent has special privileges:
- Can create new task folders
- Can transition
.statebetween phases (other agents can only request transitions) - Can claim tasks on behalf of other agents (work assignment)
- Can release claims from other agents (override)
- Can force-transition a task (override validation, with
--forceflag)
Coordinator claiming on behalf of another agent:
python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {target-agent-id} --as-coordinator
Coordinator force-release:
python ~/.automaton/scripts/status.py --release --task {task-name} --as-coordinator
Coordinator force-transition:
python ~/.automaton/scripts/status.py --transition {phase} --task {task-name} --force
These are ONLY available when --agent orchestrator (or whatever agent has role: coordinator) is provided.
8. Autopilot Mode in Multi-Agent Configuration
When Mode: multi-agent and Autopilot: Enabled:
- The coordinator agent drives the
drive_all()loop as before - But instead of executing each phase directly, it:
- Claims the task on behalf of the appropriate agent
- Loads the phase prompt for that agent's role
- Transitions
.statewhen the phase produces its artifact - Releases the claim and moves to the next phase
- If the appropriate agent is unavailable, the coordinator can execute the phase itself (coordinator has access to all phases by default)
- This preserves the autopilot behavior while respecting agent roles
When Mode: multi-agent and Autopilot: Disabled:
- Each agent uses
--next-available --agent {my-id}to find work - Each agent claims, works, transitions, and releases independently
- The coordinator monitors progress via
--listor--audit
9. Conflict Resolution
Two agents claim simultaneously:
- Atomic write (
.state.lock.tmp→.state.lock) ensures only one wins - The loser gets "ERROR: Task already claimed by agent '{winner}'"
- This is the same pattern used by
.stateatomic writes
Agent dies mid-phase:
- Lock expires after
Lock timeout(default 30 min) - Any agent can re-claim after expiry
status.py --listshows expired locks with "STALE" statusstatus.py --next-availabletreats expired locks as unclaimed
Phase mismatch after claim:
- Agent claims task in "research" phase
- By the time agent starts, another agent transitioned the task to "design"
- Agent discovers phase mismatch when it reads
.stateor runs--validate-folder - Agent should release the claim and find new work
Task completed while claimed:
- If a task reaches "complete" or "human_intervention" state, all locks on that task are automatically released
status.py --transition completeandstatus.py --transition human_interventiondelete.state.lockas part of the transition
10. Status Output with Multi-Agent Info
status.py --task {task} in multi-agent mode adds claim info:
Task: add-user-auth
Phase: research (from .state)
Agent: researcher (claimed: 2026-06-14T14:30:00Z, expires: 2026-06-14T15:00:00Z)
Agent allowed phases: research, decomposition, design, test_design
Allowed actions:
- Read project files, ask clarifying questions, write SPEC.md
Forbidden actions:
- Edit code (implement phase)
- Write BUG_REPORT.md (bug_find phase)
Next artifact needed: SPEC.md
Next phase: design or implement
status.py --list in multi-agent mode adds a "Claimed By" column:
Task Phase Claimed By Expires
add-user-auth research researcher 15:00 UTC
fix-login-bug implement implementer 15:15 UTC
add-payment-api design — —
Acceptance Criteria
- Single-agent mode (no Agent Configuration) has zero behavioral change from current status-script spec
Agent Configurationsection in.agent.mdactivates multi-agent mode--claimcreates.state.lockatomically; refuses if already claimed; overclaims if expired--releaseremoves.state.lock; refuses if wrong agent- Automatic lock release on phase transition when claiming agent can't work the next phase
--next-availablefinds highest-priority unclaimed task for a given agent role--availablelists all unclaimed tasks for a given agent role- Lock expiry works (default 30 min, configurable)
- Role binding adds agent-specific FORBIDDEN section to
status.py --taskoutput - Phase prompts include
## Agent Rolesection when agent is role-bound TASK_HANDOFFsignal defined for when agent can't perform a required phase- Coordinator agent can claim on behalf of others, force-release, force-transition
- Autopilot mode respects agent roles (claims on behalf, delegates)
- Manual mode uses
--next-availablefor self-organizing agents .state.lockis excluded from--validate-folderand artifact heuristics- Completed / human_intervention tasks auto-release locks
--listshows claim info in multi-agent mode--taskshows agent and claim info in multi-agent mode- Phase validation on claim (agent must be configured for the task's current phase)
- Schema validation for Agent Configuration (every phase covered, only one coordinator)
- Tests in
tests/test_status.pyfor all multi-agent commands - Tests for lock expiry and overclaiming
- Tests for conflict resolution (simultaneous claims, agent death, phase mismatch)
Non-Goals
- This spec does not cover agent-to-agent messaging or notification (agents discover work via
--next-availablepolling or coordinator assignment) - This spec does not cover distributed locking across network filesystems (file-based locks are sufficient for local agent coordination)
- This spec does not cover dashboard integration for multi-agent (future work)
- This spec does not cover the
.statefile format or--validate-folder/--audit(covered by state-file-enforcement and status-script specs) - This spec does not cover prompt restructuring (covered by phase-scoped-prompts)
- This spec does not cover CI/CD integration for multi-agent pipeline orchestration