Files
automaton/tasks/complete/multi-agent-support/SPEC.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

366 lines
16 KiB
Markdown

# SPEC: Multi-Agent Support
## Goal
Add opt-in multi-agent coordination to Automaton so that specialized agents (researcher, implementer, bug finder) can claim and work on different tasks or phases simultaneously. **Single-agent mode must remain the default with zero configuration changes.**
## Background
Automaton currently assumes one agent that switches between personas sequentially. The `.state` file and `status.py` foundation from the state-file-enforcement and status-script specs provide shared state and validation — but they lack three things needed for multi-agent:
1. **Claiming** — no way to mark a task as "owned" by a specific agent, so two agents can race on the same task
2. **Role binding** — no way to restrict an agent to specific phases (e.g., "Agent A only does research")
3. **Work discovery** — no way for an idle agent to find available work matching its role
These are only needed when multiple agents are running. In single-agent mode (the default), none of this machinery activates.
## Design Principle: Single-Agent Is the Zero-Config Default
When no multi-agent configuration exists:
- No `.state.lock` files are ever created
- `status.py` works exactly as specified in the status-script spec
- The orchestrator drives the full lifecycle in one session (current behavior)
- All claiming/releasing/work-queue commands are no-ops or return "single-agent mode"
- Zero performance overhead, zero behavioral change
Multi-agent activates only when `## Agent Configuration` is present in `.agent.md`.
## Requirements
### 1. Agent Configuration (`.agent.md`)
Add an optional section to `.agent.md`:
```markdown
## Agent Configuration
Mode: multi-agent
Agents:
- id: researcher
phases: [research, decomposition, design, test_design]
- id: implementer
phases: [implement]
- id: bug-hunter
phases: [bug_find, adversarial_bug_find]
- id: doc-reviewer
phases: [doc_review]
- id: referee
phases: [referee]
- id: orchestrator
phases: [new, complete, human_intervention]
role: coordinator
```
**Rules:**
- If `## Agent Configuration` is absent or `Mode: single-agent`, everything works as today — no locks, no claiming, no work queue
- If `Mode: multi-agent`, the claiming/work-queue system activates
- Agent `id` values are free-form strings (alphanumeric + hyphens)
- Each agent has an explicit list of phases it's allowed to work on
- Only one agent can have `role: coordinator` — this is the orchestrator, which claims tasks, transitions state, and delegates
- An agent can claim multiple phases
- Every phase must be covered by at least one agent (validated by `status.py`)
- Phases not listed under any agent are handled by the coordinator
**Agent identity resolution (in priority order):**
1. `--agent` flag on `status.py` commands (e.g., `--agent implementer`)
2. `AUTOMATON_AGENT_ID` environment variable
3. `agent.id` field in the project's `.agent.md`
4. If none of the above: "default" (single-agent mode)
When `--agent` is provided, the agent must exist in the Agent Configuration. If not, `status.py` errors.
### 2. Task Claiming
```
python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {agent-id} [--project {project-path}]
```
**Behavior in single-agent mode (no Agent Configuration):**
- Output: "Claimed task '{task-name}' (single-agent mode — no lock needed)"
- No `.state.lock` file created
- The command succeeds as a no-op
**Behavior in multi-agent mode:**
1. Check if `.state.lock` exists for the task
2. If no lock exists:
- Create `.state.lock` atomically (write to `.state.lock.tmp`, rename to `.state.lock`)
- Lock file format:
```
agent: {agent-id}
phase: {current-phase}
claimed: {ISO-8601-timestamp}
expires: {ISO-8601-timestamp + lock-timeout}
```
- Output: "Claimed task '{task-name}' for agent '{agent-id}' — phase: {phase}"
3. If lock exists and not expired:
- Output: "ERROR: Task '{task-name}' is claimed by agent '{agent-id}' (expires: {timestamp}). Retry after expiry or release the claim."
- Exit code 1
4. If lock exists and expired:
- Overwrite the lock with the new agent's claim
- Output: "WARN: Task '{task-name}' had stale lock from agent '{old-agent-id}' (expired: {timestamp}). Overclaiming for agent '{new-agent-id}'."
- Exit code 0
**Phase validation on claim:**
- The agent must be configured for the task's current phase
- If `agent-id` is not allowed to work on `research` and the task is in `research` phase, refuse the claim
- Output: "ERROR: Agent '{agent-id}' is not configured for phase '{phase}'. Allowed phases: {list}"
- Exit code 1
**Lock timeout:**
- Default: 30 minutes
- Configurable in `.agent.md`:
```markdown
## Agent Configuration
Mode: multi-agent
Lock timeout: 60m
```
- If `Lock timeout` is absent, default to 30 minutes
- TTL values: `5m`, `10m`, `15m`, `30m`, `60m`, `120m`
**Lock file location:** `tasks/{task-name}/.state.lock`
**Sub-task claiming:**
- Sub-tasks have their own `.state.lock` in their own folder
- Parent task lock is independent of sub-task locks
- Claiming a parent task does NOT claim its sub-tasks
### 3. Task Releasing
```
python ~/.automaton/scripts/status.py --release --task {task-name} --agent {agent-id} [--project {project-path}]
```
**Behavior in single-agent mode:**
- Output: "Released task '{task-name}' (single-agent mode — no lock to release)"
- No-op
**Behavior in multi-agent mode:**
1. Check if `.state.lock` exists for the task
2. If lock exists and owned by `{agent-id}`:
- Delete `.state.lock`
- Output: "Released task '{task-name}' from agent '{agent-id}'"
3. If lock exists but owned by a different agent:
- Output: "ERROR: Task '{task-name}' is claimed by agent '{other-agent-id}', not '{agent-id}'. Only the claiming agent can release."
- Exit code 1
4. If no lock exists:
- Output: "WARN: Task '{task-name}' has no lock. Nothing to release."
- Exit code 0
**Automatic release on phase transition:**
When `status.py --transition` succeeds in multi-agent mode:
- The lock is updated to reflect the new phase (if the claiming agent is still valid for the new phase)
- If the claiming agent is NOT valid for the new phase, the lock is released automatically
- This enables handoff: researcher claims, does research, transitions to design, lock auto-releases, designer claims
### 4. Work Discovery
```
python ~/.automaton/scripts/status.py --next-available --agent {agent-id} [--project {project-path}]
```
**Behavior in single-agent mode:**
- Output: "single-agent mode — use --list to see all tasks"
**Behavior in multi-agent mode:**
1. Scan all tasks in `{project}/.automaton/tasks/` (including sub-tasks)
2. For each task:
- Read `.state` to determine current phase
- Check if the task is unclaimed (no `.state.lock`) or has an expired lock
- Check if `{agent-id}` is configured for the task's current phase
3. Return the first available task sorted by priority:
- Task closest to completion first (referee > doc_review > adversarial_bug_find > bug_find > implement > test_design > design > decomposition > research > new)
- Within the same priority level, alphabetical by task name
4. Output:
```
Next available task for agent 'implementer':
Task: fix-login-bug
Phase: implement
Phase priority: 7 (high — close to completion)
Status: unclaimed
To claim: python ~/.automaton/scripts/status.py --claim --task fix-login-bug --agent implementer
```
5. If no tasks available:
- Output: "No tasks available for agent '{agent-id}'. All tasks are claimed by other agents, in phases this agent cannot work on, or completed."
**Work queue (list all available):**
```
python ~/.automaton/scripts/status.py --available --agent {agent-id} [--project {project-path}]
```
Same logic as `--next-available` but returns ALL matching tasks, not just the first:
```
Available tasks for agent 'implementer':
1. Task: fix-login-bug | Phase: implement | Status: unclaimed
2. Task: add-payment-api | Phase: implement | Status: claimed by 'other-agent' (expires: 2026-06-14T15:30:00Z)
```
### 5. Lock File Details
**Format:**
```
agent: {agent-id}
phase: {current-phase-from-state-file}
claimed: 2026-06-14T14:30:00Z
expires: 2026-06-14T15:00:00Z
```
**Atomic writes:** Same pattern as `.state` — write to `.state.lock.tmp`, then rename to `.state.lock`.
**File classification:** `.state.lock` is metadata (like `.state`), not a phase deliverable. It is excluded from `--validate-folder` checks and artifact heuristics.
**Git:** `.state.lock` should be in `.gitignore` (it's ephemeral agent state, not project state).
**Lock expiry check:** Every `status.py` command that reads `.state.lock` must also check expiry. Expired locks are treated as non-existent (available for claiming).
### 6. Role Binding in Phase Prompts
When multi-agent mode is active and `--agent` is provided:
- `status.py --task {task}` output includes the agent's allowed phases:
```
Task: add-user-auth
Phase: research (from .state)
Agent: researcher
Agent allowed phases: research, decomposition, design, test_design
ALLOWED for this agent:
- Read project files, ask questions, write SPEC.md
- Transition to decompose, design (if agent is configured for those phases)
FORBIDDEN for this agent:
- Edit code (implement phase)
- Write BUG_REPORT.md (bug_find phase)
- Write ADVERSARIAL_BUG_REPORT.md (adversarial_bug_find phase)
- Write DOC_REVIEW.md (doc_review phase)
- Write VERDICT.md (referee phase)
```
- Phase prompts gain an additional section when the agent is role-bound:
```markdown
## Agent Role
You are agent '{agent-id}'. Your allowed phases are: {phases}.
You may NOT perform actions from phases not in your allowed list.
If the task requires a phase you cannot perform, output: "TASK_HANDOFF — task '{task-name}' requires phase '{required-phase}' which is outside my allowed phases. Agent '{allowed-agent-list}' can continue."
```
**Single-agent mode:** This section is absent. The agent has full access to all phases.
### 7. Coordinator Role
The `orchestrator` agent has special privileges:
- Can create new task folders
- Can transition `.state` between phases (other agents can only request transitions)
- Can claim tasks on behalf of other agents (work assignment)
- Can release claims from other agents (override)
- Can force-transition a task (override validation, with `--force` flag)
**Coordinator claiming on behalf of another agent:**
```
python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {target-agent-id} --as-coordinator
```
**Coordinator force-release:**
```
python ~/.automaton/scripts/status.py --release --task {task-name} --as-coordinator
```
**Coordinator force-transition:**
```
python ~/.automaton/scripts/status.py --transition {phase} --task {task-name} --force
```
These are ONLY available when `--agent orchestrator` (or whatever agent has `role: coordinator`) is provided.
### 8. Autopilot Mode in Multi-Agent Configuration
When `Mode: multi-agent` and `Autopilot: Enabled`:
- The coordinator agent drives the `drive_all()` loop as before
- But instead of executing each phase directly, it:
1. Claims the task on behalf of the appropriate agent
2. Loads the phase prompt for that agent's role
3. Transitions `.state` when the phase produces its artifact
4. Releases the claim and moves to the next phase
- If the appropriate agent is unavailable, the coordinator can execute the phase itself (coordinator has access to all phases by default)
- This preserves the autopilot behavior while respecting agent roles
When `Mode: multi-agent` and `Autopilot: Disabled`:
- Each agent uses `--next-available --agent {my-id}` to find work
- Each agent claims, works, transitions, and releases independently
- The coordinator monitors progress via `--list` or `--audit`
### 9. Conflict Resolution
**Two agents claim simultaneously:**
- Atomic write (`.state.lock.tmp` → `.state.lock`) ensures only one wins
- The loser gets "ERROR: Task already claimed by agent '{winner}'"
- This is the same pattern used by `.state` atomic writes
**Agent dies mid-phase:**
- Lock expires after `Lock timeout` (default 30 min)
- Any agent can re-claim after expiry
- `status.py --list` shows expired locks with "STALE" status
- `status.py --next-available` treats expired locks as unclaimed
**Phase mismatch after claim:**
- Agent claims task in "research" phase
- By the time agent starts, another agent transitioned the task to "design"
- Agent discovers phase mismatch when it reads `.state` or runs `--validate-folder`
- Agent should release the claim and find new work
**Task completed while claimed:**
- If a task reaches "complete" or "human_intervention" state, all locks on that task are automatically released
- `status.py --transition complete` and `status.py --transition human_intervention` delete `.state.lock` as part of the transition
### 10. Status Output with Multi-Agent Info
`status.py --task {task}` in multi-agent mode adds claim info:
```
Task: add-user-auth
Phase: research (from .state)
Agent: researcher (claimed: 2026-06-14T14:30:00Z, expires: 2026-06-14T15:00:00Z)
Agent allowed phases: research, decomposition, design, test_design
Allowed actions:
- Read project files, ask clarifying questions, write SPEC.md
Forbidden actions:
- Edit code (implement phase)
- Write BUG_REPORT.md (bug_find phase)
Next artifact needed: SPEC.md
Next phase: design or implement
```
`status.py --list` in multi-agent mode adds a "Claimed By" column:
```
Task Phase Claimed By Expires
add-user-auth research researcher 15:00 UTC
fix-login-bug implement implementer 15:15 UTC
add-payment-api design — —
```
## Acceptance Criteria
- [ ] Single-agent mode (no Agent Configuration) has zero behavioral change from current status-script spec
- [ ] `Agent Configuration` section in `.agent.md` activates multi-agent mode
- [ ] `--claim` creates `.state.lock` atomically; refuses if already claimed; overclaims if expired
- [ ] `--release` removes `.state.lock`; refuses if wrong agent
- [ ] Automatic lock release on phase transition when claiming agent can't work the next phase
- [ ] `--next-available` finds highest-priority unclaimed task for a given agent role
- [ ] `--available` lists all unclaimed tasks for a given agent role
- [ ] Lock expiry works (default 30 min, configurable)
- [ ] Role binding adds agent-specific FORBIDDEN section to `status.py --task` output
- [ ] Phase prompts include `## Agent Role` section when agent is role-bound
- [ ] `TASK_HANDOFF` signal defined for when agent can't perform a required phase
- [ ] Coordinator agent can claim on behalf of others, force-release, force-transition
- [ ] Autopilot mode respects agent roles (claims on behalf, delegates)
- [ ] Manual mode uses `--next-available` for self-organizing agents
- [ ] `.state.lock` is excluded from `--validate-folder` and artifact heuristics
- [ ] Completed / human_intervention tasks auto-release locks
- [ ] `--list` shows claim info in multi-agent mode
- [ ] `--task` shows agent and claim info in multi-agent mode
- [ ] Phase validation on claim (agent must be configured for the task's current phase)
- [ ] Schema validation for Agent Configuration (every phase covered, only one coordinator)
- [ ] Tests in `tests/test_status.py` for all multi-agent commands
- [ ] Tests for lock expiry and overclaiming
- [ ] Tests for conflict resolution (simultaneous claims, agent death, phase mismatch)
## Non-Goals
- This spec does not cover agent-to-agent messaging or notification (agents discover work via `--next-available` polling or coordinator assignment)
- This spec does not cover distributed locking across network filesystems (file-based locks are sufficient for local agent coordination)
- This spec does not cover dashboard integration for multi-agent (future work)
- This spec does not cover the `.state` file format or `--validate-folder`/`--audit` (covered by state-file-enforcement and status-script specs)
- This spec does not cover prompt restructuring (covered by phase-scoped-prompts)
- This spec does not cover CI/CD integration for multi-agent pipeline orchestration