State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks
Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)
Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)
Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
This commit is contained in:
@@ -0,0 +1,366 @@
|
||||
# SPEC: Multi-Agent Support
|
||||
|
||||
## Goal
|
||||
Add opt-in multi-agent coordination to Automaton so that specialized agents (researcher, implementer, bug finder) can claim and work on different tasks or phases simultaneously. **Single-agent mode must remain the default with zero configuration changes.**
|
||||
|
||||
## Background
|
||||
Automaton currently assumes one agent that switches between personas sequentially. The `.state` file and `status.py` foundation from the state-file-enforcement and status-script specs provide shared state and validation — but they lack three things needed for multi-agent:
|
||||
1. **Claiming** — no way to mark a task as "owned" by a specific agent, so two agents can race on the same task
|
||||
2. **Role binding** — no way to restrict an agent to specific phases (e.g., "Agent A only does research")
|
||||
3. **Work discovery** — no way for an idle agent to find available work matching its role
|
||||
|
||||
These are only needed when multiple agents are running. In single-agent mode (the default), none of this machinery activates.
|
||||
|
||||
## Design Principle: Single-Agent Is the Zero-Config Default
|
||||
|
||||
When no multi-agent configuration exists:
|
||||
- No `.state.lock` files are ever created
|
||||
- `status.py` works exactly as specified in the status-script spec
|
||||
- The orchestrator drives the full lifecycle in one session (current behavior)
|
||||
- All claiming/releasing/work-queue commands are no-ops or return "single-agent mode"
|
||||
- Zero performance overhead, zero behavioral change
|
||||
|
||||
Multi-agent activates only when `## Agent Configuration` is present in `.agent.md`.
|
||||
|
||||
## Requirements
|
||||
|
||||
### 1. Agent Configuration (`.agent.md`)
|
||||
|
||||
Add an optional section to `.agent.md`:
|
||||
|
||||
```markdown
|
||||
## Agent Configuration
|
||||
|
||||
Mode: multi-agent
|
||||
Agents:
|
||||
- id: researcher
|
||||
phases: [research, decomposition, design, test_design]
|
||||
- id: implementer
|
||||
phases: [implement]
|
||||
- id: bug-hunter
|
||||
phases: [bug_find, adversarial_bug_find]
|
||||
- id: doc-reviewer
|
||||
phases: [doc_review]
|
||||
- id: referee
|
||||
phases: [referee]
|
||||
- id: orchestrator
|
||||
phases: [new, complete, human_intervention]
|
||||
role: coordinator
|
||||
```
|
||||
|
||||
**Rules:**
|
||||
- If `## Agent Configuration` is absent or `Mode: single-agent`, everything works as today — no locks, no claiming, no work queue
|
||||
- If `Mode: multi-agent`, the claiming/work-queue system activates
|
||||
- Agent `id` values are free-form strings (alphanumeric + hyphens)
|
||||
- Each agent has an explicit list of phases it's allowed to work on
|
||||
- Only one agent can have `role: coordinator` — this is the orchestrator, which claims tasks, transitions state, and delegates
|
||||
- An agent can claim multiple phases
|
||||
- Every phase must be covered by at least one agent (validated by `status.py`)
|
||||
- Phases not listed under any agent are handled by the coordinator
|
||||
|
||||
**Agent identity resolution (in priority order):**
|
||||
1. `--agent` flag on `status.py` commands (e.g., `--agent implementer`)
|
||||
2. `AUTOMATON_AGENT_ID` environment variable
|
||||
3. `agent.id` field in the project's `.agent.md`
|
||||
4. If none of the above: "default" (single-agent mode)
|
||||
|
||||
When `--agent` is provided, the agent must exist in the Agent Configuration. If not, `status.py` errors.
|
||||
|
||||
### 2. Task Claiming
|
||||
|
||||
```
|
||||
python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {agent-id} [--project {project-path}]
|
||||
```
|
||||
|
||||
**Behavior in single-agent mode (no Agent Configuration):**
|
||||
- Output: "Claimed task '{task-name}' (single-agent mode — no lock needed)"
|
||||
- No `.state.lock` file created
|
||||
- The command succeeds as a no-op
|
||||
|
||||
**Behavior in multi-agent mode:**
|
||||
1. Check if `.state.lock` exists for the task
|
||||
2. If no lock exists:
|
||||
- Create `.state.lock` atomically (write to `.state.lock.tmp`, rename to `.state.lock`)
|
||||
- Lock file format:
|
||||
```
|
||||
agent: {agent-id}
|
||||
phase: {current-phase}
|
||||
claimed: {ISO-8601-timestamp}
|
||||
expires: {ISO-8601-timestamp + lock-timeout}
|
||||
```
|
||||
- Output: "Claimed task '{task-name}' for agent '{agent-id}' — phase: {phase}"
|
||||
3. If lock exists and not expired:
|
||||
- Output: "ERROR: Task '{task-name}' is claimed by agent '{agent-id}' (expires: {timestamp}). Retry after expiry or release the claim."
|
||||
- Exit code 1
|
||||
4. If lock exists and expired:
|
||||
- Overwrite the lock with the new agent's claim
|
||||
- Output: "WARN: Task '{task-name}' had stale lock from agent '{old-agent-id}' (expired: {timestamp}). Overclaiming for agent '{new-agent-id}'."
|
||||
- Exit code 0
|
||||
|
||||
**Phase validation on claim:**
|
||||
- The agent must be configured for the task's current phase
|
||||
- If `agent-id` is not allowed to work on `research` and the task is in `research` phase, refuse the claim
|
||||
- Output: "ERROR: Agent '{agent-id}' is not configured for phase '{phase}'. Allowed phases: {list}"
|
||||
- Exit code 1
|
||||
|
||||
**Lock timeout:**
|
||||
- Default: 30 minutes
|
||||
- Configurable in `.agent.md`:
|
||||
```markdown
|
||||
## Agent Configuration
|
||||
Mode: multi-agent
|
||||
Lock timeout: 60m
|
||||
```
|
||||
- If `Lock timeout` is absent, default to 30 minutes
|
||||
- TTL values: `5m`, `10m`, `15m`, `30m`, `60m`, `120m`
|
||||
|
||||
**Lock file location:** `tasks/{task-name}/.state.lock`
|
||||
|
||||
**Sub-task claiming:**
|
||||
- Sub-tasks have their own `.state.lock` in their own folder
|
||||
- Parent task lock is independent of sub-task locks
|
||||
- Claiming a parent task does NOT claim its sub-tasks
|
||||
|
||||
### 3. Task Releasing
|
||||
|
||||
```
|
||||
python ~/.automaton/scripts/status.py --release --task {task-name} --agent {agent-id} [--project {project-path}]
|
||||
```
|
||||
|
||||
**Behavior in single-agent mode:**
|
||||
- Output: "Released task '{task-name}' (single-agent mode — no lock to release)"
|
||||
- No-op
|
||||
|
||||
**Behavior in multi-agent mode:**
|
||||
1. Check if `.state.lock` exists for the task
|
||||
2. If lock exists and owned by `{agent-id}`:
|
||||
- Delete `.state.lock`
|
||||
- Output: "Released task '{task-name}' from agent '{agent-id}'"
|
||||
3. If lock exists but owned by a different agent:
|
||||
- Output: "ERROR: Task '{task-name}' is claimed by agent '{other-agent-id}', not '{agent-id}'. Only the claiming agent can release."
|
||||
- Exit code 1
|
||||
4. If no lock exists:
|
||||
- Output: "WARN: Task '{task-name}' has no lock. Nothing to release."
|
||||
- Exit code 0
|
||||
|
||||
**Automatic release on phase transition:**
|
||||
When `status.py --transition` succeeds in multi-agent mode:
|
||||
- The lock is updated to reflect the new phase (if the claiming agent is still valid for the new phase)
|
||||
- If the claiming agent is NOT valid for the new phase, the lock is released automatically
|
||||
- This enables handoff: researcher claims, does research, transitions to design, lock auto-releases, designer claims
|
||||
|
||||
### 4. Work Discovery
|
||||
|
||||
```
|
||||
python ~/.automaton/scripts/status.py --next-available --agent {agent-id} [--project {project-path}]
|
||||
```
|
||||
|
||||
**Behavior in single-agent mode:**
|
||||
- Output: "single-agent mode — use --list to see all tasks"
|
||||
|
||||
**Behavior in multi-agent mode:**
|
||||
1. Scan all tasks in `{project}/.automaton/tasks/` (including sub-tasks)
|
||||
2. For each task:
|
||||
- Read `.state` to determine current phase
|
||||
- Check if the task is unclaimed (no `.state.lock`) or has an expired lock
|
||||
- Check if `{agent-id}` is configured for the task's current phase
|
||||
3. Return the first available task sorted by priority:
|
||||
- Task closest to completion first (referee > doc_review > adversarial_bug_find > bug_find > implement > test_design > design > decomposition > research > new)
|
||||
- Within the same priority level, alphabetical by task name
|
||||
4. Output:
|
||||
```
|
||||
Next available task for agent 'implementer':
|
||||
Task: fix-login-bug
|
||||
Phase: implement
|
||||
Phase priority: 7 (high — close to completion)
|
||||
Status: unclaimed
|
||||
|
||||
To claim: python ~/.automaton/scripts/status.py --claim --task fix-login-bug --agent implementer
|
||||
```
|
||||
5. If no tasks available:
|
||||
- Output: "No tasks available for agent '{agent-id}'. All tasks are claimed by other agents, in phases this agent cannot work on, or completed."
|
||||
|
||||
**Work queue (list all available):**
|
||||
```
|
||||
python ~/.automaton/scripts/status.py --available --agent {agent-id} [--project {project-path}]
|
||||
```
|
||||
|
||||
Same logic as `--next-available` but returns ALL matching tasks, not just the first:
|
||||
```
|
||||
Available tasks for agent 'implementer':
|
||||
1. Task: fix-login-bug | Phase: implement | Status: unclaimed
|
||||
2. Task: add-payment-api | Phase: implement | Status: claimed by 'other-agent' (expires: 2026-06-14T15:30:00Z)
|
||||
```
|
||||
|
||||
### 5. Lock File Details
|
||||
|
||||
**Format:**
|
||||
```
|
||||
agent: {agent-id}
|
||||
phase: {current-phase-from-state-file}
|
||||
claimed: 2026-06-14T14:30:00Z
|
||||
expires: 2026-06-14T15:00:00Z
|
||||
```
|
||||
|
||||
**Atomic writes:** Same pattern as `.state` — write to `.state.lock.tmp`, then rename to `.state.lock`.
|
||||
|
||||
**File classification:** `.state.lock` is metadata (like `.state`), not a phase deliverable. It is excluded from `--validate-folder` checks and artifact heuristics.
|
||||
|
||||
**Git:** `.state.lock` should be in `.gitignore` (it's ephemeral agent state, not project state).
|
||||
|
||||
**Lock expiry check:** Every `status.py` command that reads `.state.lock` must also check expiry. Expired locks are treated as non-existent (available for claiming).
|
||||
|
||||
### 6. Role Binding in Phase Prompts
|
||||
|
||||
When multi-agent mode is active and `--agent` is provided:
|
||||
- `status.py --task {task}` output includes the agent's allowed phases:
|
||||
```
|
||||
Task: add-user-auth
|
||||
Phase: research (from .state)
|
||||
Agent: researcher
|
||||
Agent allowed phases: research, decomposition, design, test_design
|
||||
|
||||
ALLOWED for this agent:
|
||||
- Read project files, ask questions, write SPEC.md
|
||||
- Transition to decompose, design (if agent is configured for those phases)
|
||||
|
||||
FORBIDDEN for this agent:
|
||||
- Edit code (implement phase)
|
||||
- Write BUG_REPORT.md (bug_find phase)
|
||||
- Write ADVERSARIAL_BUG_REPORT.md (adversarial_bug_find phase)
|
||||
- Write DOC_REVIEW.md (doc_review phase)
|
||||
- Write VERDICT.md (referee phase)
|
||||
```
|
||||
- Phase prompts gain an additional section when the agent is role-bound:
|
||||
```markdown
|
||||
## Agent Role
|
||||
You are agent '{agent-id}'. Your allowed phases are: {phases}.
|
||||
You may NOT perform actions from phases not in your allowed list.
|
||||
If the task requires a phase you cannot perform, output: "TASK_HANDOFF — task '{task-name}' requires phase '{required-phase}' which is outside my allowed phases. Agent '{allowed-agent-list}' can continue."
|
||||
```
|
||||
|
||||
**Single-agent mode:** This section is absent. The agent has full access to all phases.
|
||||
|
||||
### 7. Coordinator Role
|
||||
|
||||
The `orchestrator` agent has special privileges:
|
||||
- Can create new task folders
|
||||
- Can transition `.state` between phases (other agents can only request transitions)
|
||||
- Can claim tasks on behalf of other agents (work assignment)
|
||||
- Can release claims from other agents (override)
|
||||
- Can force-transition a task (override validation, with `--force` flag)
|
||||
|
||||
**Coordinator claiming on behalf of another agent:**
|
||||
```
|
||||
python ~/.automaton/scripts/status.py --claim --task {task-name} --agent {target-agent-id} --as-coordinator
|
||||
```
|
||||
|
||||
**Coordinator force-release:**
|
||||
```
|
||||
python ~/.automaton/scripts/status.py --release --task {task-name} --as-coordinator
|
||||
```
|
||||
|
||||
**Coordinator force-transition:**
|
||||
```
|
||||
python ~/.automaton/scripts/status.py --transition {phase} --task {task-name} --force
|
||||
```
|
||||
|
||||
These are ONLY available when `--agent orchestrator` (or whatever agent has `role: coordinator`) is provided.
|
||||
|
||||
### 8. Autopilot Mode in Multi-Agent Configuration
|
||||
|
||||
When `Mode: multi-agent` and `Autopilot: Enabled`:
|
||||
- The coordinator agent drives the `drive_all()` loop as before
|
||||
- But instead of executing each phase directly, it:
|
||||
1. Claims the task on behalf of the appropriate agent
|
||||
2. Loads the phase prompt for that agent's role
|
||||
3. Transitions `.state` when the phase produces its artifact
|
||||
4. Releases the claim and moves to the next phase
|
||||
- If the appropriate agent is unavailable, the coordinator can execute the phase itself (coordinator has access to all phases by default)
|
||||
- This preserves the autopilot behavior while respecting agent roles
|
||||
|
||||
When `Mode: multi-agent` and `Autopilot: Disabled`:
|
||||
- Each agent uses `--next-available --agent {my-id}` to find work
|
||||
- Each agent claims, works, transitions, and releases independently
|
||||
- The coordinator monitors progress via `--list` or `--audit`
|
||||
|
||||
### 9. Conflict Resolution
|
||||
|
||||
**Two agents claim simultaneously:**
|
||||
- Atomic write (`.state.lock.tmp` → `.state.lock`) ensures only one wins
|
||||
- The loser gets "ERROR: Task already claimed by agent '{winner}'"
|
||||
- This is the same pattern used by `.state` atomic writes
|
||||
|
||||
**Agent dies mid-phase:**
|
||||
- Lock expires after `Lock timeout` (default 30 min)
|
||||
- Any agent can re-claim after expiry
|
||||
- `status.py --list` shows expired locks with "STALE" status
|
||||
- `status.py --next-available` treats expired locks as unclaimed
|
||||
|
||||
**Phase mismatch after claim:**
|
||||
- Agent claims task in "research" phase
|
||||
- By the time agent starts, another agent transitioned the task to "design"
|
||||
- Agent discovers phase mismatch when it reads `.state` or runs `--validate-folder`
|
||||
- Agent should release the claim and find new work
|
||||
|
||||
**Task completed while claimed:**
|
||||
- If a task reaches "complete" or "human_intervention" state, all locks on that task are automatically released
|
||||
- `status.py --transition complete` and `status.py --transition human_intervention` delete `.state.lock` as part of the transition
|
||||
|
||||
### 10. Status Output with Multi-Agent Info
|
||||
|
||||
`status.py --task {task}` in multi-agent mode adds claim info:
|
||||
```
|
||||
Task: add-user-auth
|
||||
Phase: research (from .state)
|
||||
Agent: researcher (claimed: 2026-06-14T14:30:00Z, expires: 2026-06-14T15:00:00Z)
|
||||
Agent allowed phases: research, decomposition, design, test_design
|
||||
Allowed actions:
|
||||
- Read project files, ask clarifying questions, write SPEC.md
|
||||
Forbidden actions:
|
||||
- Edit code (implement phase)
|
||||
- Write BUG_REPORT.md (bug_find phase)
|
||||
Next artifact needed: SPEC.md
|
||||
Next phase: design or implement
|
||||
```
|
||||
|
||||
`status.py --list` in multi-agent mode adds a "Claimed By" column:
|
||||
```
|
||||
Task Phase Claimed By Expires
|
||||
add-user-auth research researcher 15:00 UTC
|
||||
fix-login-bug implement implementer 15:15 UTC
|
||||
add-payment-api design — —
|
||||
```
|
||||
|
||||
## Acceptance Criteria
|
||||
- [ ] Single-agent mode (no Agent Configuration) has zero behavioral change from current status-script spec
|
||||
- [ ] `Agent Configuration` section in `.agent.md` activates multi-agent mode
|
||||
- [ ] `--claim` creates `.state.lock` atomically; refuses if already claimed; overclaims if expired
|
||||
- [ ] `--release` removes `.state.lock`; refuses if wrong agent
|
||||
- [ ] Automatic lock release on phase transition when claiming agent can't work the next phase
|
||||
- [ ] `--next-available` finds highest-priority unclaimed task for a given agent role
|
||||
- [ ] `--available` lists all unclaimed tasks for a given agent role
|
||||
- [ ] Lock expiry works (default 30 min, configurable)
|
||||
- [ ] Role binding adds agent-specific FORBIDDEN section to `status.py --task` output
|
||||
- [ ] Phase prompts include `## Agent Role` section when agent is role-bound
|
||||
- [ ] `TASK_HANDOFF` signal defined for when agent can't perform a required phase
|
||||
- [ ] Coordinator agent can claim on behalf of others, force-release, force-transition
|
||||
- [ ] Autopilot mode respects agent roles (claims on behalf, delegates)
|
||||
- [ ] Manual mode uses `--next-available` for self-organizing agents
|
||||
- [ ] `.state.lock` is excluded from `--validate-folder` and artifact heuristics
|
||||
- [ ] Completed / human_intervention tasks auto-release locks
|
||||
- [ ] `--list` shows claim info in multi-agent mode
|
||||
- [ ] `--task` shows agent and claim info in multi-agent mode
|
||||
- [ ] Phase validation on claim (agent must be configured for the task's current phase)
|
||||
- [ ] Schema validation for Agent Configuration (every phase covered, only one coordinator)
|
||||
- [ ] Tests in `tests/test_status.py` for all multi-agent commands
|
||||
- [ ] Tests for lock expiry and overclaiming
|
||||
- [ ] Tests for conflict resolution (simultaneous claims, agent death, phase mismatch)
|
||||
|
||||
## Non-Goals
|
||||
- This spec does not cover agent-to-agent messaging or notification (agents discover work via `--next-available` polling or coordinator assignment)
|
||||
- This spec does not cover distributed locking across network filesystems (file-based locks are sufficient for local agent coordination)
|
||||
- This spec does not cover dashboard integration for multi-agent (future work)
|
||||
- This spec does not cover the `.state` file format or `--validate-folder`/`--audit` (covered by state-file-enforcement and status-script specs)
|
||||
- This spec does not cover prompt restructuring (covered by phase-scoped-prompts)
|
||||
- This spec does not cover CI/CD integration for multi-agent pipeline orchestration
|
||||
Reference in New Issue
Block a user