Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

16 KiB

SPEC: Multi-Agent Support

Goal

Add opt-in multi-agent coordination to Automaton so that specialized agents (researcher, implementer, bug finder) can claim and work on different tasks or phases simultaneously. Single-agent mode must remain the default with zero configuration changes.

Background

Automaton currently assumes one agent that switches between personas sequentially. The .state file and status.py foundation from the state-file-enforcement and status-script specs provide shared state and validation — but they lack three things needed for multi-agent:

  1. Claiming — no way to mark a task as "owned" by a specific agent, so two agents can race on the same task
  2. Role binding — no way to restrict an agent to specific phases (e.g., "Agent A only does research")
  3. Work discovery — no way for an idle agent to find available work matching its role

These are only needed when multiple agents are running. In single-agent mode (the default), none of this machinery activates.

Design Principle: Single-Agent Is the Zero-Config Default

When no multi-agent configuration exists:

  • No .state.lock files are ever created
  • status.py works exactly as specified in the status-script spec
  • The orchestrator drives the full lifecycle in one session (current behavior)
  • All claiming/releasing/work-queue commands are no-ops or return "single-agent mode"
  • Zero performance overhead, zero behavioral change

Multi-agent activates only when ## Agent Configuration is present in .agent.md.

Requirements

1. Agent Configuration (.agent.md)

Add an optional section to .agent.md:

## Agent Configuration

Mode: multi-agent
Agents:
  - id: researcher
    phases: [research, decomposition, design, test_design]
  - id: implementer
    phases: [implement]
  - id: bug-hunter
    phases: [bug_find, adversarial_bug_find]
  - id: doc-reviewer
    phases: [doc_review]
  - id: referee
    phases: [referee]
  - id: orchestrator
    phases: [new, complete, human_intervention]
    role: coordinator

Rules:

  • If ## Agent Configuration is absent or Mode: single-agent, everything works as today — no locks, no claiming, no work queue
  • If Mode: multi-agent, the claiming/work-queue system activates
  • Agent id values are free-form strings (alphanumeric + hyphens)
  • Each agent has an explicit list of phases it's allowed to work on
  • Only one agent can have role: coordinator — this is the orchestrator, which claims tasks, transitions state, and delegates
  • An agent can claim multiple phases
  • Every phase must be covered by at least one agent (validated by status.py)
  • Phases not listed under any agent are handled by the coordinator

Agent identity resolution (in priority order):

  1. --agent flag on status.py commands (e.g., --agent implementer)
  2. AUTOMATON_AGENT_ID environment variable
  3. agent.id field in the project's .agent.md
  4. If none of the above: "default" (single-agent mode)

When --agent is provided, the agent must exist in the Agent Configuration. If not, status.py errors.

2. Task Claiming

python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {agent-id} [--project {project-path}]

Behavior in single-agent mode (no Agent Configuration):

  • Output: "Claimed task '{task-name}' (single-agent mode — no lock needed)"
  • No .state.lock file created
  • The command succeeds as a no-op

Behavior in multi-agent mode:

  1. Check if .state.lock exists for the task
  2. If no lock exists:
    • Create .state.lock atomically (write to .state.lock.tmp, rename to .state.lock)
    • Lock file format:
      agent: {agent-id}
      phase: {current-phase}
      claimed: {ISO-8601-timestamp}
      expires: {ISO-8601-timestamp + lock-timeout}
      
    • Output: "Claimed task '{task-name}' for agent '{agent-id}' — phase: {phase}"
  3. If lock exists and not expired:
    • Output: "ERROR: Task '{task-name}' is claimed by agent '{agent-id}' (expires: {timestamp}). Retry after expiry or release the claim."
    • Exit code 1
  4. If lock exists and expired:
    • Overwrite the lock with the new agent's claim
    • Output: "WARN: Task '{task-name}' had stale lock from agent '{old-agent-id}' (expired: {timestamp}). Overclaiming for agent '{new-agent-id}'."
    • Exit code 0

Phase validation on claim:

  • The agent must be configured for the task's current phase
  • If agent-id is not allowed to work on research and the task is in research phase, refuse the claim
  • Output: "ERROR: Agent '{agent-id}' is not configured for phase '{phase}'. Allowed phases: {list}"
  • Exit code 1

Lock timeout:

  • Default: 30 minutes
  • Configurable in .agent.md:
    ## Agent Configuration
    Mode: multi-agent
    Lock timeout: 60m
    
  • If Lock timeout is absent, default to 30 minutes
  • TTL values: 5m, 10m, 15m, 30m, 60m, 120m

Lock file location: tasks/{task-name}/.state.lock

Sub-task claiming:

  • Sub-tasks have their own .state.lock in their own folder
  • Parent task lock is independent of sub-task locks
  • Claiming a parent task does NOT claim its sub-tasks

3. Task Releasing

python ~/.automaton/scripts/status.py --release --task {task-name} --agent {agent-id} [--project {project-path}]

Behavior in single-agent mode:

  • Output: "Released task '{task-name}' (single-agent mode — no lock to release)"
  • No-op

Behavior in multi-agent mode:

  1. Check if .state.lock exists for the task
  2. If lock exists and owned by {agent-id}:
    • Delete .state.lock
    • Output: "Released task '{task-name}' from agent '{agent-id}'"
  3. If lock exists but owned by a different agent:
    • Output: "ERROR: Task '{task-name}' is claimed by agent '{other-agent-id}', not '{agent-id}'. Only the claiming agent can release."
    • Exit code 1
  4. If no lock exists:
    • Output: "WARN: Task '{task-name}' has no lock. Nothing to release."
    • Exit code 0

Automatic release on phase transition: When status.py --transition succeeds in multi-agent mode:

  • The lock is updated to reflect the new phase (if the claiming agent is still valid for the new phase)
  • If the claiming agent is NOT valid for the new phase, the lock is released automatically
  • This enables handoff: researcher claims, does research, transitions to design, lock auto-releases, designer claims

4. Work Discovery

python ~/.automaton/scripts/status.py --next-available --agent {agent-id} [--project {project-path}]

Behavior in single-agent mode:

  • Output: "single-agent mode — use --list to see all tasks"

Behavior in multi-agent mode:

  1. Scan all tasks in {project}/.automaton/tasks/ (including sub-tasks)
  2. For each task:
    • Read .state to determine current phase
    • Check if the task is unclaimed (no .state.lock) or has an expired lock
    • Check if {agent-id} is configured for the task's current phase
  3. Return the first available task sorted by priority:
    • Task closest to completion first (referee > doc_review > adversarial_bug_find > bug_find > implement > test_design > design > decomposition > research > new)
    • Within the same priority level, alphabetical by task name
  4. Output:
    Next available task for agent 'implementer':
    Task: fix-login-bug
    Phase: implement
    Phase priority: 7 (high — close to completion)
    Status: unclaimed
    
    To claim: python ~/.automaton/scripts/status.py --claim --task fix-login-bug --agent implementer
    
  5. If no tasks available:
    • Output: "No tasks available for agent '{agent-id}'. All tasks are claimed by other agents, in phases this agent cannot work on, or completed."

Work queue (list all available):

python ~/.automaton/scripts/status.py --available --agent {agent-id} [--project {project-path}]

Same logic as --next-available but returns ALL matching tasks, not just the first:

Available tasks for agent 'implementer':
1. Task: fix-login-bug | Phase: implement | Status: unclaimed
2. Task: add-payment-api | Phase: implement | Status: claimed by 'other-agent' (expires: 2026-06-14T15:30:00Z)

5. Lock File Details

Format:

agent: {agent-id}
phase: {current-phase-from-state-file}
claimed: 2026-06-14T14:30:00Z
expires: 2026-06-14T15:00:00Z

Atomic writes: Same pattern as .state — write to .state.lock.tmp, then rename to .state.lock.

File classification: .state.lock is metadata (like .state), not a phase deliverable. It is excluded from --validate-folder checks and artifact heuristics.

Git: .state.lock should be in .gitignore (it's ephemeral agent state, not project state).

Lock expiry check: Every status.py command that reads .state.lock must also check expiry. Expired locks are treated as non-existent (available for claiming).

6. Role Binding in Phase Prompts

When multi-agent mode is active and --agent is provided:

  • status.py --task {task} output includes the agent's allowed phases:
    Task: add-user-auth
    Phase: research (from .state)
    Agent: researcher
    Agent allowed phases: research, decomposition, design, test_design
    
    ALLOWED for this agent:
    - Read project files, ask questions, write SPEC.md
    - Transition to decompose, design (if agent is configured for those phases)
    
    FORBIDDEN for this agent:
    - Edit code (implement phase)
    - Write BUG_REPORT.md (bug_find phase)
    - Write ADVERSARIAL_BUG_REPORT.md (adversarial_bug_find phase)
    - Write DOC_REVIEW.md (doc_review phase)
    - Write VERDICT.md (referee phase)
    
  • Phase prompts gain an additional section when the agent is role-bound:
    ## Agent Role
    You are agent '{agent-id}'. Your allowed phases are: {phases}.
    You may NOT perform actions from phases not in your allowed list.
    If the task requires a phase you cannot perform, output: "TASK_HANDOFF — task '{task-name}' requires phase '{required-phase}' which is outside my allowed phases. Agent '{allowed-agent-list}' can continue."
    

Single-agent mode: This section is absent. The agent has full access to all phases.

7. Coordinator Role

The orchestrator agent has special privileges:

  • Can create new task folders
  • Can transition .state between phases (other agents can only request transitions)
  • Can claim tasks on behalf of other agents (work assignment)
  • Can release claims from other agents (override)
  • Can force-transition a task (override validation, with --force flag)

Coordinator claiming on behalf of another agent:

python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {target-agent-id} --as-coordinator

Coordinator force-release:

python ~/.automaton/scripts/status.py --release --task {task-name} --as-coordinator

Coordinator force-transition:

python ~/.automaton/scripts/status.py --transition {phase} --task {task-name} --force

These are ONLY available when --agent orchestrator (or whatever agent has role: coordinator) is provided.

8. Autopilot Mode in Multi-Agent Configuration

When Mode: multi-agent and Autopilot: Enabled:

  • The coordinator agent drives the drive_all() loop as before
  • But instead of executing each phase directly, it:
    1. Claims the task on behalf of the appropriate agent
    2. Loads the phase prompt for that agent's role
    3. Transitions .state when the phase produces its artifact
    4. Releases the claim and moves to the next phase
  • If the appropriate agent is unavailable, the coordinator can execute the phase itself (coordinator has access to all phases by default)
  • This preserves the autopilot behavior while respecting agent roles

When Mode: multi-agent and Autopilot: Disabled:

  • Each agent uses --next-available --agent {my-id} to find work
  • Each agent claims, works, transitions, and releases independently
  • The coordinator monitors progress via --list or --audit

9. Conflict Resolution

Two agents claim simultaneously:

  • Atomic write (.state.lock.tmp → .state.lock) ensures only one wins
  • The loser gets "ERROR: Task already claimed by agent '{winner}'"
  • This is the same pattern used by .state atomic writes

Agent dies mid-phase:

  • Lock expires after Lock timeout (default 30 min)
  • Any agent can re-claim after expiry
  • status.py --list shows expired locks with "STALE" status
  • status.py --next-available treats expired locks as unclaimed

Phase mismatch after claim:

  • Agent claims task in "research" phase
  • By the time agent starts, another agent transitioned the task to "design"
  • Agent discovers phase mismatch when it reads .state or runs --validate-folder
  • Agent should release the claim and find new work

Task completed while claimed:

  • If a task reaches "complete" or "human_intervention" state, all locks on that task are automatically released
  • status.py --transition complete and status.py --transition human_intervention delete .state.lock as part of the transition

10. Status Output with Multi-Agent Info

status.py --task {task} in multi-agent mode adds claim info:

Task: add-user-auth
Phase: research (from .state)
Agent: researcher (claimed: 2026-06-14T14:30:00Z, expires: 2026-06-14T15:00:00Z)
Agent allowed phases: research, decomposition, design, test_design
Allowed actions:
  - Read project files, ask clarifying questions, write SPEC.md
Forbidden actions:
  - Edit code (implement phase)
  - Write BUG_REPORT.md (bug_find phase)
Next artifact needed: SPEC.md
Next phase: design or implement

status.py --list in multi-agent mode adds a "Claimed By" column:

Task                     Phase               Claimed By          Expires
add-user-auth            research             researcher          15:00 UTC
fix-login-bug            implement            implementer         15:15 UTC
add-payment-api          design               —                   —

Acceptance Criteria

  • Single-agent mode (no Agent Configuration) has zero behavioral change from current status-script spec
  • Agent Configuration section in .agent.md activates multi-agent mode
  • --claim creates .state.lock atomically; refuses if already claimed; overclaims if expired
  • --release removes .state.lock; refuses if wrong agent
  • Automatic lock release on phase transition when claiming agent can't work the next phase
  • --next-available finds highest-priority unclaimed task for a given agent role
  • --available lists all unclaimed tasks for a given agent role
  • Lock expiry works (default 30 min, configurable)
  • Role binding adds agent-specific FORBIDDEN section to status.py --task output
  • Phase prompts include ## Agent Role section when agent is role-bound
  • TASK_HANDOFF signal defined for when agent can't perform a required phase
  • Coordinator agent can claim on behalf of others, force-release, force-transition
  • Autopilot mode respects agent roles (claims on behalf, delegates)
  • Manual mode uses --next-available for self-organizing agents
  • .state.lock is excluded from --validate-folder and artifact heuristics
  • Completed / human_intervention tasks auto-release locks
  • --list shows claim info in multi-agent mode
  • --task shows agent and claim info in multi-agent mode
  • Phase validation on claim (agent must be configured for the task's current phase)
  • Schema validation for Agent Configuration (every phase covered, only one coordinator)
  • Tests in tests/test_status.py for all multi-agent commands
  • Tests for lock expiry and overclaiming
  • Tests for conflict resolution (simultaneous claims, agent death, phase mismatch)

Non-Goals

  • This spec does not cover agent-to-agent messaging or notification (agents discover work via --next-available polling or coordinator assignment)
  • This spec does not cover distributed locking across network filesystems (file-based locks are sufficient for local agent coordination)
  • This spec does not cover dashboard integration for multi-agent (future work)
  • This spec does not cover the .state file format or --validate-folder/--audit (covered by state-file-enforcement and status-script specs)
  • This spec does not cover prompt restructuring (covered by phase-scoped-prompts)
  • This spec does not cover CI/CD integration for multi-agent pipeline orchestration