Compare commits

...
14 Commits
Author SHA1 Message Date
Lap Tran f13043d315 feat(pi): 3 new Pi Dev extensions — read-tweet, auto-context, task-lifecycle tools
CI / build (push) Has been cancelled
2026-06-26 22:41:38 -04:00
Lap Tran c2355954b9 feat(pi): pi-automaton.sh startup context printer for Pi Dev
Pi Dev doesn't auto-load automaton's system prompt (unlike opencode),
so the agent has no awareness of tasks, phases, or status.py commands.
This bridges that gap:

- scripts/pi-automaton.sh: NEW — detects scope (framework/project),
  reads active tasks, default model, config snippet, AGENTS.md rules,
  and prints a formatted context block the user pastes as their first
  message to the Pi Dev agent. Does NOT launch pi.
- scripts/register-guards.sh: EXTENDED — after Pi Dev guard install,
  offers to symlink pi-automaton.sh to ~/bin/pi-automaton
  (interactive prompt only when stdin is a terminal).
- README.md: added 'Using Pi Dev with automaton' FAQ subsection
  documenting the context printer workflow.
2026-06-26 20:47:08 -04:00
Lap Tran 0437cbae6c docs: architecture section, FAQ, cross-references, vault-memory update
README.md:
  - Added §1.5 'Architecture: Framework vs Project' with directory tree
    and lifecycle flow diagram
  - Added §2.5 FAQ covering coexistence, hooks, loops, multi-project
  - Updated install section already done in prior commit

AGENTS.md:
  - Added 'Script Cross-References' table mapping install.sh →
    onboard-project.sh → status.py --create-task

scripts/install-hooks.sh:
  - Updated header to reference onboard-project.sh as caller
  - Added 'Next step' line pointing to --create-task

vault-memory CONTEXT.md:
  - Updated test count (518→611)
  - Added new scripts (detect_models.py, onboard-project.sh)
  - Documented install/onboard flow and split architecture
2026-06-26 16:51:39 -04:00
Lap Tran b880f2535a fix(install): make install.sh idempotent + add onboard-project.sh
- install.sh: no longer exits early when ~/.automaton exists.
  Skips the clone but runs all setup (VRAM detection, guards,
  self-improvement loop, virtualenv). Both curl|bash and
  git-clone + ./install.sh now work correctly.
- onboard-project.sh: new script that bootstraps automaton in a
  project — creates .automaton/ skeleton, detects models, writes
  config.md, inits git, installs hooks, adds .gitignore entries.
- README.md: fix install flow docs (curl|bash + clone-then-run),
  add onboard-project.sh as Option A for project setup
2026-06-26 13:54:28 -04:00
Lap Tran f980ccfe27 feat(dashboard): model badges on kanban cards + design doc update
- task.py: Task dataclass gains models: dict[str, str], loaded from
  .state.models in discover_tasks()
- app.py: models dict included in all task API responses
- dashboard.js: model badges rendered between artifacts and subtask
  progress on kanban cards; ROLE_LABELS map for readable tooltips
- styles.css: .task-card-models and .model-badge styles
- design/loops/technical.md: document {model} substitution token,
  per-role model field, and model-divergence brake gate (gate #7)
2026-06-26 13:26:00 -04:00
Lap Tran 35e449b03e feat(model-divergence): full enforcement — manifest, transition, claim, audit, loop gates, detect script
Completes all 3 model-divergence enforcement subtasks:

- scripts/detect_models.py: probes opencode.json + localhost endpoints,
  builds models.json with --json/--write/--force
- scripts/status.py: CONFLICT_MATRIX, --model flag, --transition --model,
  --claim --model, --audit Category 6, model-divergence brake gate in
  --check-gate, helpers for manifest loading and conflict checking
- scripts/loop-runner.py: _role_model() helper + {model} passed via extras
  dict to _invoke_harness for implement, verify, orchestrate roles
- tests/test_model_divergence.py: 33 tests covering all enforcement layers
- Single-LLM mode: record model advisory, no conflict check
- Multi-LLM mode (2+ models): conflict matrix enforced at transition, claim,
  and loop brake gate
- Project-level models.json preferred over global ~/.automaton/models.json
2026-06-26 13:23:17 -04:00
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00
Lap Tran fe43b9e1fc Remove old review system from dashboard (cosmetic only)
CI / build (push) Has been cancelled
The review system (REVIEW.md, approve/changes_requested buttons,
review filter, review badges) was purely cosmetic — only the dashboard
read/wrote it. No workflow component (status.py, autopilot.py,
loop-runner.py, prompts) ever enforced it.

The 'Approve' button in the task detail panel confused users into
thinking it approved the task's phase gate. In reality it only wrote
to REVIEW.md, which had zero effect on transitions.

Removed:
- Review section (buttons, textarea, status badge) from detail panel
- Review badge from task cards
- Review filter from toolbar
- Pending-review counter from header
- All review-related CSS

Users now use the single '🔓 Approve Phase' button in the detail
panel, which calls status.py --approve and actually transitions the
task.
2026-06-25 11:23:55 -04:00
Lap Tran 54f65f9861 Add phase approval endpoint and button to dashboard
CI / build (push) Has been cancelled
The dashboard had no way to run status.py --approve. The existing
'Approve' button only wrote to REVIEW.md (code review status), not
approve the task's phase gate.

Changes:
- /api/approve/{task_name} POST endpoint in app.py calls status.py --approve
- '🔓 Approve Phase' button in detail panel for approval-gated tasks
- approvePhase() JS function sends the POST and refreshes the board
- CSS for the new button (warning-yellow styling)

This enables the full automation flow: user clicks Approve Phase →
autopilot picks up the approved task on next tick → drives it forward.
2026-06-25 11:01:47 -04:00
Lap Tran 7a5aa1f3a4 Add --execute and --install-schedule to autopilot.py
CI / build (push) Has been cancelled
--execute flag runs status.py --transition instead of printing suggestions.
--install-schedule creates a launchd plist that runs autopilot --drive --execute
every 60 seconds, providing full automation for task lifecycle transitions.

The autopilot loop:
- Runs on a schedule (60s default), ticks the most advanced unblocked task
- Stalls at :awaiting_approval gates until user runs --approve
- Handles all phases: new → research → decomposition → design → test_design →
  implement → code_review → bug_find → adversarial_bug_find → doc_review →
  referee → complete
- Fails cleanly when required artifacts are missing (agent must write them)
- Logs stdout/stderr to ~/.automaton/logs/autopilot-*.log

Also fixes test leak in scripts/automaton-cleanup.sh (pytest temp path was
being written into the real stub).
2026-06-25 10:53:14 -04:00
Lap Tran d325963644 Flatten model-divergence subtasks into 3 independent tasks
CI / build (push) Has been cancelled
Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)

Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.

Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.
2026-06-25 07:25:10 -04:00
Lap Tran 7336db282d Bootstrap self-improvement loop, decompose model-divergence-enforcement
CI / build (push) Has been cancelled
Self-improvement loop:
- Created via --create-loop --from-template self-improvement
- Scheduled via launchd (3600s interval)
- State: running

model-divergence-enforcement task:
- Research approved, decomposed into 3 sequential subtasks:
  1. mde-manifest-detection (models.json + detect_models.py)
  2. mde-interactive-enforcement (conflict matrix + --model args + audit)
  3. mde-loop-enforcement (loop.json roles + {model} substitution + check-gate)
- Parent at decomposition:approved (stays active until subtasks complete)
- Each subtask has BRIEF.md with scope, deliverables, acceptance criteria
2026-06-25 07:15:25 -04:00
Lap Tran 715f6f9495 Fix 7 bugs in autopilot.py, add 43 tests
Fixes from functional correctness review:
- PHASE_PRIORITY: add code_review (was missing, caused wrong task selection)
- cmd_drive: implement→code_review (was illegal implement→bug_find)
- cmd_drive: add handlers for code_review and code_review:approved
- _all_tasks: skip tasks/complete/ (was showing completed tasks as phase=None)
- _all_tasks: recurse into subtasks/ (subtasks were invisible to autopilot)
- _all_tasks: use .automaton/tasks for non-framework projects (was project/tasks)
- is_terminal: only True for complete (human_intervention has legal transitions)
- needs_user_input: flag human_intervention as blocked (was auto-driving it)

Tests cover all 8 bugs: PHASE_PRIORITY ordering, every phase transition,
code_review handlers, complete/ exclusion, subtask recursion, project path,
terminal state, user-input detection, human_intervention handling.
2026-06-25 07:11:44 -04:00
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00
702 changed files with 4745 additions and 278 deletions
+1
View File
@@ -3,3 +3,4 @@ __pycache__/
*.egg-info/
.venv/
venv/
logs/
+9
View File
@@ -111,3 +111,12 @@ When the agent receives a trigger command, it must:
1. Read the corresponding template file.
2. Replace all `{placeholders}` with the actual project values.
3. Execute the rendered prompt.
## Backlog
Design backlogs are the framework's outstanding-work store when no active tasks exist. Each design area has its own `BACKLOG.md`:
- `~/.automaton/design/loops/BACKLOG.md` — loop engineering v1.1+ and deferred items.
- `~/.automaton/design/framework/BACKLOG.md` — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement).
The self-improvement loop consumes these automatically when configured with `work_source.kind = "backlog"` and `work_source.area` set to the relevant design area (`"loops"` or `"framework"`). For manual work, read the topmost `- [ ]` item and create a task via `status.py --create-task`.
+18
View File
@@ -47,6 +47,9 @@ Automaton is a **contract-based, state-enforced workflow framework** for LLM age
│ ├── config.py
│ ├── core/
│ └── ui/
├── design/ # Design docs for major features
│ ├── loops/ # Loop engineering system (v1 locked)
│ └── framework/ # Framework agent features (rule agents, Agent tab, model-divergence)
├── tests/ # pytest suite
└── tasks/ # Framework development tasks
```
@@ -133,6 +136,21 @@ Modes:
2. Run `python3 -m pytest tests/test_prompt_paths.py` to ensure task paths are canonical.
3. Update `CHANGELOG.md` under `[unreleased]`.
## Script Cross-References
The framework provides three lifecycle scripts that should be referenced from each other:
| Script | Purpose | Called when | Next step |
|---|---|---|---|
| `scripts/install.sh` | Install framework on a fresh machine | `curl \| bash` or `git clone + bash` | → `scripts/onboard-project.sh` |
| `scripts/onboard-project.sh` | Bootstrap automaton in a project | After framework install, per project | → `status.py --create-task` |
| `scripts/install-hooks.sh` | Install git hooks per project | After onboarding, or manually | see `onboard-project.sh` |
- `install.sh` outputs "Next: onboard-project.sh" at the end.
- `onboard-project.sh` outputs "Next: status.py --create-task" at the end.
- `install-hooks.sh` is called by `onboard-project.sh` automatically.
- `update.sh` does NOT call `onboard-project.sh` — it only updates the framework.
## Adding a New Script
1. Place the script in `scripts/`.
+30
View File
@@ -2,6 +2,36 @@
## [unreleased]
### Fixed — dashboard scroll-reset on auto-refresh
- **`automaton/dashboard/html/dashboard.js`** (`renderBoard`): auto-refresh rebuilt the board via `board.innerHTML = html` every tick (default 2s), destroying each `.column-body`'s `scrollTop` and snapping it back to 0 — so users couldn't scroll the Done group down to review older tasks. Now snapshots each column-body's `scrollTop` (plus the board's `scrollLeft` and the active view's `scrollTop`) before the rebuild and restores them after, matched by index (PHASE_GROUPS order is stable).
- **New tests**: `tests/test_dashboard_ui.py` — Playwright browser smoke test (board renders tasks; column scroll survives an auto-refresh tick). Skipped via `importorskip` when playwright/chromium is absent so CI without a browser stays green. Verified the test fails without the fix (scrollTop resets to 0) and passes with it.
### Fixed — failing plist-isolation test (host bleed false positive)
- **`tests/test_cleanup_done.py`** (`TestInstallCleanupScheduleIsolation.test_plist_written_to_override_dir_not_host`): asserted `not host.exists()`, but the host `~/Library/LaunchAgents/com.automaton.cleanup.plist` legitimately exists from a real `--install-cleanup-schedule` run, causing a false failure. Now snapshots the host plist's `st_mtime_ns` (or absence) before the test run and asserts it's unchanged after — a pre-existing real install no longer fails the test; only an actual write during the run would.
### Changed — bind ornith as the Implement model
- **`config.md`** (Model Configuration): set `Model: omlx/Ornith-1.0-35B-4bit-mlx` and `Override context window: 32768` (matches the opencode.json limit for the local LLM). Interactive autopilot already used ornith via opencode's default model; this makes it explicit so auto-detection can't pick another model. Loop ticks still use the single `harness.command` for all roles — per-role model binding (`{model}` substitution in loop-runner.py) is **not** implemented yet (see model-divergence gap below).
### Changed — README: document the self-improvement loop's scope for new projects
- **`README.md`** (Project Setup): added "The Self-Improvement Loop is framework-scoped" note — the default SI loop targets `~/.automaton/` (the framework), not your project, by design. Documents the leave-running / pause / create-a-project-loop paths.
### Known gap — model-divergence was marked complete but unimplemented
- The `model-divergence-enforcement` parent task and its 3 subtasks (`mde-manifest-detection`, `mde-interactive-enforcement`, `mde-loop-enforcement`) are `.state = complete` but contain only `SPEC.md`/`DECOMPOSITION.md` — no `IMPLEMENTATION.md`, no `VERDICT.md`. The promised code (`status.py` model_divergence audit category, `--transition --model`, `.state.models`, `loop.json` per-role `model` + `{model}` substitution in loop-runner.py, dashboard badges) was never written. Consequence: loop roles (implement/verify/orchestrate) all run the same model, so the D12 conflict-of-interest rule (Verify ≠ Implement session/model) is unenforced. Interactive autopilot is unaffected.
### Added — framework agent features design docs
- **New `design/framework/`** directory: design index, functional design, technical design, and backlog for three framework-level agent features:
- **Model-Divergence Enforcement** — `models.json` manifest, single vs multi-LLM mode detection, conflict matrix (`code_review≠implement`, `bug_find≠implement`, `adversarial_bug_find≠implement+bug_find`, `referee≠implement+bug_find+adversarial_bug_find`, `loop-verify≠loop-implement`), `--transition --model`, `.state.models`, `loop.json` per-role model binding + `{model}` substitution, `--audit` model_divergence category, dashboard badges. Shipped as a manual task (3 subtasks), not a backlog item.
- **Rule Agents** — Rule Proposer (daily scheduled standalone agent, scans completed tasks' `BUG_REPORT.md`/`ADVERSARIAL_BUG_REPORT.md`/`VERDICT.md`, proposes rules to `RULE_PROPOSALS.md`) and Rule Reviewer (monthly scheduled standalone agent, consolidates `.rules.md` → `RULE_REVIEW.md`). Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers. Conflict-of-interest: Reviewer must differ from Proposer; both must differ from tasks they review. Enforcement deferred until model-divergence ships.
- **Agent Tab Redesign** — replace 4 fake `AGENT_TYPE_META` types with two sections: Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) + Scheduled Jobs (real `job.kind`: cleanup, loop, rule-scan, rule-review). New `/api/phase-roles` endpoint.
- **`design/framework/BACKLOG.md`**: 3 v1 items (`agent-tab-real-roles`, `rule-proposer-agent`, `rule-reviewer-agent`) + 3 deferred items. Consumable by the self-improvement loop via `work_source.area = "framework"`.
- **Cross-references**: `design/loops/README.md` and `design/loops/BACKLOG.md` now point to the sibling `design/framework/` area. `AGENTS.md` repo layout updated.
### Added — cross-loop task claim (task `add-claim-loop-task`)
- **`scripts/status.py`**: New `--claim-loop-task <name> --task <taskname> [--project P]` command. Exit 0 = claimed (or already self-claimed, idempotent). Exit 2 = already claimed by another running/paused loop (`task_already_claimed:{other}` on stderr) or untracked loop (`loop_untracked`). Uses `_loop_lock` to serialize writes; cross-loop scan is advisory (self-healing on next tick).
+185 -20
View File
@@ -6,22 +6,28 @@ A contract-based operating system for LLM agents, designed to enforce discipline
Before you can use the framework in any project, you must install the core logic into your local environment.
**Run these commands in your terminal:**
Choose one of the following methods:
### Option A: One-liner (curl pipe, recommended)
```bash
# Clone the framework into the global config directory
git clone <your-git-url> ~/.automaton
# Enter the directory
cd ~/.automaton
# Make the installation script executable and run it
# You must provide the git URL as the first argument
chmod +x install.sh
./install.sh <your-git-url>
curl -fsSL https://raw.githubusercontent.com/<your-org>/automaton/main/scripts/install.sh | bash -s -- <your-git-url>
```
The git URL is required because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling.
This clones the framework to `~/.automaton/`, runs VRAM detection, installs pre-edit guards, creates the self-improvement loop, and sets up the Python virtualenv.
### Option B: Clone first
```bash
git clone <your-git-url> ~/.automaton
bash ~/.automaton/scripts/install.sh
```
The script detects that `~/.automaton` already exists, skips the clone, and runs all setup steps (VRAM detection, guards, loop, virtualenv).
### Both methods do the same thing
The git URL is required on fresh install because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling.
*Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory. A self-improvement loop is created and scheduled by default (see Loop Engineering below).*
@@ -45,11 +51,95 @@ If a project was set up under the old model (with copies of framework files), it
---
## 1.5 Architecture: Framework vs Project
Automaton uses a **split architecture** — one shared framework, many project `./.automaton/` directories:
```
~/.automaton/ ← Framework (installed once per machine)
├── scripts/ ← shared tooling: status.py, loop-runner.py
├── prompts/ ← shared LLM prompts
├── plugins/ ← shared harness plugins
├── templates/ ← shared task & loop templates
├── .automaton/tasks/ ← framework housekeeping tasks (self-improvement)
└── .automaton/loops/ ← framework loops (self-improvement loop)
~/projects/my-app/
└── .automaton/ ← Project (onboarded once per project)
├── tasks/ ← YOUR project's tasks
├── models.json ← YOUR project's model config
├── config.md ← YOUR project's VRAM config
├── project-name.md ← YOUR project's display name
└── loops/ ← YOUR project's loops
```
**Key rules:**
- The agent is **scope-aware**: if you're inside `~/.automaton/`, it operates in **framework mode** (reads framework tasks). If you're inside a project dir, it operates in **project mode** (reads project tasks). They never interfere.
- All framework scripts (`status.py`, etc.) live in `~/.automaton/scripts/` and are shared — never copied into projects.
- Framework prompts live in `~/.automaton/prompts/` — projects reference them by path at runtime.
- Git hooks are **per-project**. Each project installs its own via `bash ~/.automaton/scripts/install-hooks.sh <project-path>`.
- The self-improvement loop targets **only** the framework itself. Your project won't get framework-level tasks in its board.
- You can work on both at the same time in different terminals — independent `.automaton/` directories, shared tooling.
### Lifecycle overview
```
┌─────────────────────────────────────────────────────┐
│ 1. Install Framework (once per machine) │
│ curl .../install.sh | bash -s -- <git-url> │
│ → clones to ~/.automaton/ │
│ → VRAM detection, guards, venv, self-improvement │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 2. Onboard Project (once per project) │
│ bash ~/.automaton/scripts/onboard-project.sh <dir> │
│ → creates .automaton/ skeleton │
│ → probes models, writes config.md │
│ → git init + hooks │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 3. Create Task (per feature) │
│ python3 ~/.automaton/scripts/status.py │
│ --create-task my-feature --project . │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 4. Work Through Phases (per task) │
│ status.py --transition research --task my-feature │
│ → agent writes SPEC.md │
│ status.py --transition implement --task my-feature │
│ → agent writes code + IMPLEMENTATION.md │
│ ... → complete │
└─────────────────────────────────────────────────────┘
```
---
## 2. Project Setup (Per project)
Once the framework is installed globally, you must "onboard" every individual project you work on.
### Option A: The Agent-Driven Way (Recommended)
### Option A: The Onboarding Script (Recommended)
```bash
bash ~/.automaton/scripts/onboard-project.sh /path/to/project
```
This will:
1. Create `.automaton/` skeleton if missing.
2. Run `detect_models.py --write` to probe local models (falls back to a minimal `models.json`).
3. Generate `config.md` with VRAM recommendations.
4. Write `project-name.md` from the directory name.
5. Initialize git if not already a repo.
6. Install git hooks (pre-commit + pre-push).
7. Add automaton entries to `.gitignore`.
8. Run `status.py --audit` to verify the setup.
### Option B: The Agent-Driven Way
If you want the agent to handle the configuration for you, navigate to your project root and run:
> *"Onboard this project into automaton."
@@ -60,17 +150,91 @@ The agent will automatically:
3. Generate your `.agent.md` and `.rules.md` files.
4. Initiate the "Exploration Ritual" to understand your codebase.
### Option B: The Manual Way
If you prefer to set it up manually, create a `.automaton/` directory in your project root and add:
- `.agent.md`: Project-specific configuration (Mode, rules, etc.).
- `.rules.md`: Project-specific constraints and past failure modes.
### Option C: The Manual Way
If you prefer to set it up manually:
Then install the git pre-commit hook:
```bash
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit
# Create the automaton directory
mkdir -p .automaton/tasks .automaton/loops .automaton/design
# Install git hooks
bash ~/.automaton/scripts/install-hooks.sh .
# Configure your model(s)
python3 ~/.automaton/scripts/detect_models.py --write --project .
```
This hook blocks commits when no task is in `implement` or `doc_review` phase, preventing accidental code changes without a tracked task.
Then add `.agent.md` and `.rules.md` for the agent.
### The Self-Improvement Loop is framework-scoped
The self-improvement loop installed by default targets the **framework itself** (`work_source.project: ~/.automaton/`), not your project. It ticks hourly against `status.py --audit` on `~/.automaton/` to keep the framework healthy. This is by design — it improves the framework you depend on while you work on your project.
- **Leave it running** if you want the framework maintained in the background (recommended).
- **Disable it** if you want zero background activity: `python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/`
- **Want a loop on your project too?** Create a separate one targeted at the project root:
```bash
python3 ~/.automaton/scripts/status.py --create-loop my-project-loop \
--from-template self-improvement --project /path/to/project
```
---
## 2.5 FAQ
### Can I work on the framework and a project at the same time?
Yes. They have separate `.automaton/` directories. Open two terminals:
```
Terminal 1: cd ~/.automaton → framework mode
Terminal 2: cd ~/projects/my-app → project mode
```
The agent detects scope from your current directory. Each can have its own tasks, loops, and config. They share the same `~/.automaton/scripts/` binaries.
### Why doesn't `install.sh` need a Git URL when run from the repo?
Because the framework is already cloned. `install.sh` skips cloning when `~/.automaton/` exists and runs all the setup steps (VRAM detection, pip deps, self-improvement loop, guards). The Git URL is only required for a fresh install via `curl | bash`.
### Do I need to run `install.sh` again after pulling updates?
No. `git pull` inside `~/.automaton/` updates the code. The self-improvement loop and guards persist across updates. If you want to re-register guards (e.g. after switching harnesses), run `bash ~/.automaton/scripts/register-guards.sh`.
### How do git hooks work per project?
Each project installs its own hooks via:
```bash
bash ~/.automaton/scripts/install-hooks.sh /path/to/project
```
The pre-commit hook blocks commits when no task is in `implement` or `doc_review` phase. The pre-push hook catches `--no-verify` bypasses. They're independent per repo.
### Can I have multiple projects onboarded at once?
Yes. Each project gets its own `.automaton/` directory. Run `onboard-project.sh` once per project. The shared scripts in `~/.automaton/scripts/` enforce the state machine on whichever project you point `--project` at.
### What about loops on my project?
The self-improvement loop runs only on the framework. To add a loop to your project:
```bash
python3 ~/.automaton/scripts/status.py --create-loop my-loop \
--from-template self-improvement --project /path/to/project
python3 ~/.automaton/scripts/status.py --install-schedule my-loop \
--interval 3600 --project /path/to/project
```
### Using Pi Dev with automaton
Pi Dev has the `automaton-guard-pi` plugin (installed by `register-guards.sh`) which blocks edits outside allowed phases. However, Pi Dev does **not** auto-load automaton's system prompt (unlike opencode). For the agent to understand tasks and phases, provide context manually.
**Before starting a Pi Dev session**, run the context printer:
```bash
bash ~/.automaton/scripts/pi-automaton.sh
```
Or if symlinked to `~/bin/`:
```bash
pi-automaton
```
Copy the output and paste it as your first message to the Pi Dev agent. This tells the agent about:
- The automaton workflow framework and phase lifecycle
- Active tasks in the current project
- The default model and VRAM configuration
- Project-specific rules from `AGENTS.md`
- Which commands to use for transitions and task creation
The `automaton-guard-pi` plugin still blocks edits outside `implement`/`doc_review` even without this context — the context printer just makes the agent *aware* of why it's being blocked and how to use the framework correctly.
---
@@ -250,6 +414,7 @@ The runner resolves prompt files from `loop.json` `roles.*.prompt` (e.g. `loop-i
| `blast_radius.file_scope` | List of paths the loop may edit |
| `blast_radius.use_worktree` | If true, tick runs in a per-loop git worktree |
| `work_source.kind` | `single`, `audit`, or `backlog` |
| `work_source.area` | Design area for `backlog` kind (default `"loops"`; `"framework"` reads `design/framework/BACKLOG.md`) |
| `roles.implement.prompt` | Prompt file for Implement role |
| `roles.verify.prompt` | Prompt file for Verify role |
| `roles.orchestrate.prompt` | Prompt file for Orchestrate role |
+42 -22
View File
@@ -115,6 +115,7 @@ class Task:
name: str
folder_path: Path
state: TaskState
phase_raw: str = ""
artifacts: dict[str, ArtifactStatus] = field(default_factory=dict)
sub_tasks: list[SubTask] = field(default_factory=list)
parent_spec: Optional[str] = None
@@ -125,10 +126,14 @@ class Task:
design_content: Optional[str] = None
spec_content: Optional[str] = None
decomposition_content: Optional[str] = None
code_review_content: Optional[str] = None
test_plan_content: Optional[str] = None
implementation_content: Optional[str] = None
parent_spec_content: Optional[str] = None
vram_config_content: Optional[str] = None
waves: list[WaveGroup] = field(default_factory=list)
is_corrupted: bool = False
models: dict[str, str] = field(default_factory=dict)
@property
def display_name(self) -> str:
@@ -297,7 +302,7 @@ class Task:
@property
def is_approval_gated(self) -> bool:
return self.state in (TaskState.RESEARCH, TaskState.DECOMPOSITION, TaskState.DESIGN, TaskState.TEST_DESIGN, TaskState.CODE_REVIEW)
return self.phase_raw.endswith(":awaiting_approval")
@property
def blocker(self) -> str:
@@ -310,7 +315,7 @@ class Task:
if not artifact.content:
return f"Empty required artifact: {required}"
if self.is_approval_gated:
return "Awaiting user approval — use `status.py --approve` to approve"
return "Awaiting user approval — click Approve Phase in the dashboard"
if self.state == TaskState.BUG_FIND:
if "BUG_REPORT.md" not in self.artifacts or not self.artifacts["BUG_REPORT.md"].content:
return "Agent must generate BUG_REPORT.md"
@@ -484,7 +489,7 @@ def _state_string_to_task_state(phase: str) -> TaskState:
return mapping.get(base, TaskState.BACKLOG)
def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, ArtifactStatus]]:
def determine_task_state(folder_path: Path) -> tuple[TaskState, str, dict[str, ArtifactStatus]]:
artifacts = {}
for filename in ARTIFACTS:
@@ -511,7 +516,7 @@ def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, Artifa
try:
phase = state_file.read_text(encoding="utf-8").strip()
if phase:
return _state_string_to_task_state(phase), artifacts
return _state_string_to_task_state(phase), phase, artifacts
except (OSError, IOError):
pass # Fall through to artifact heuristic
@@ -521,40 +526,40 @@ def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, Artifa
if "VERDICT.md" in artifacts:
verdict_content = artifacts["VERDICT.md"].content
if not verdict_content:
return TaskState.BLOCKED, artifacts
return TaskState.BLOCKED, "blocked", artifacts
verdict_status = parse_verdict_status(verdict_content)
if verdict_status == VERDICT_FAIL or verdict_status == VERDICT_NEEDS_REVIEW:
return TaskState.BLOCKED, artifacts
return TaskState.BLOCKED, "blocked", artifacts
if verdict_status == VERDICT_PASS:
return TaskState.DONE, artifacts
return TaskState.DONE, "done", artifacts
# Verdict exists but status is unparseable — pending referee review
if verdict_status is None:
return TaskState.REFEREE, artifacts
return TaskState.REFEREE, "referee", artifacts
# State machine aligned with orchestrate.md
# Check from most advanced to least advanced
if "DOC_REVIEW.md" in artifacts:
return TaskState.DOC_REVIEW, artifacts
return TaskState.DOC_REVIEW, "doc_review", artifacts
if "ADVERSARIAL_BUG_REPORT.md" in artifacts and "BUG_REPORT.md" in artifacts:
return TaskState.ADV_BUG_FIND, artifacts
return TaskState.ADV_BUG_FIND, "adv_bug_find", artifacts
if "BUG_REPORT.md" in artifacts:
return TaskState.BUG_FIND, artifacts
return TaskState.BUG_FIND, "bug_find", artifacts
if "ADVERSARIAL_BUG_REPORT.md" in artifacts:
return TaskState.BUG_FIND, artifacts
return TaskState.BUG_FIND, "bug_find", artifacts
if "CODE_REVIEW.md" in artifacts:
return TaskState.CODE_REVIEW, artifacts
return TaskState.CODE_REVIEW, "code_review", artifacts
if "IMPLEMENTATION.md" in artifacts:
return TaskState.IMPLEMENT, artifacts
return TaskState.IMPLEMENT, "implement", artifacts
if "TEST_PLAN.md" in artifacts:
return TaskState.TEST_DESIGN, artifacts
return TaskState.TEST_DESIGN, "test_design", artifacts
if "DESIGN.md" in artifacts:
return TaskState.DESIGN, artifacts
return TaskState.DESIGN, "design", artifacts
if "DECOMPOSITION.md" in artifacts:
return TaskState.DECOMPOSITION, artifacts
return TaskState.DECOMPOSITION, "decomposition", artifacts
if "SPEC.md" in artifacts:
return TaskState.RESEARCH, artifacts
return TaskState.RESEARCH, "research", artifacts
return TaskState.BACKLOG, artifacts
return TaskState.BACKLOG, "backlog", artifacts
def parse_sub_tasks(parent_folder: Path) -> list[SubTask]:
@@ -569,7 +574,7 @@ def parse_sub_tasks(parent_folder: Path) -> list[SubTask]:
name = subtask_folder.name
if not _VALID_TASK_NAME_CHARS.issuperset(set(name)):
continue
state, artifacts = determine_task_state(subtask_folder)
state, _, artifacts = determine_task_state(subtask_folder)
verdict_status = None
has_verdict = False
if "VERDICT.md" in artifacts and artifacts["VERDICT.md"].content:
@@ -653,12 +658,12 @@ def discover_tasks(tasks_dir: Path) -> list[Task]:
if not _VALID_TASK_NAME_CHARS.issuperset(set(folder_path.name)):
continue
state, artifacts = determine_task_state(folder_path)
state, phase_raw, artifacts = determine_task_state(folder_path)
sub_tasks = parse_sub_tasks(folder_path)
task = Task(
name=folder_path.name, folder_path=folder_path, state=state,
artifacts=artifacts, sub_tasks=sub_tasks,
phase_raw=phase_raw, artifacts=artifacts, sub_tasks=sub_tasks,
)
# Load specific artifact contents for detail panel
@@ -677,9 +682,24 @@ def discover_tasks(tasks_dir: Path) -> list[Task]:
if "DECOMPOSITION.md" in artifacts and artifacts["DECOMPOSITION.md"].content:
task.decomposition_content = artifacts["DECOMPOSITION.md"].content
task.waves = parse_waves(task.decomposition_content)
if "CODE_REVIEW.md" in artifacts and artifacts["CODE_REVIEW.md"].content:
task.code_review_content = artifacts["CODE_REVIEW.md"].content
if "TEST_PLAN.md" in artifacts and artifacts["TEST_PLAN.md"].content:
task.test_plan_content = artifacts["TEST_PLAN.md"].content
if "IMPLEMENTATION.md" in artifacts and artifacts["IMPLEMENTATION.md"].content:
task.implementation_content = artifacts["IMPLEMENTATION.md"].content
task.parent_spec_content = parse_parent_spec(folder_path)
task.vram_config_content = parse_vram_config(folder_path)
# Load .state.models for model divergence badges
state_models_path = folder_path / ".state.models"
if state_models_path.exists():
try:
import json as _json
task.models = _json.loads(state_models_path.read_text(encoding="utf-8"))
except (OSError, IOError, _json.JSONDecodeError):
task.models = {}
tasks.append(task)
# Sort by state (most advanced first)
+154 -54
View File
@@ -1,7 +1,7 @@
const state = {
scope: 'none', currentView: 'board', theme: 'default', selectedTask: null,
tasks: [], refreshCount: 0, autoRefresh: true, showWaves: true,
filterPhase: 'all', filterReview: 'all', filterWave: 'all', searchQuery: '', filterVisible: false,
filterPhase: 'all', filterWave: 'all', searchQuery: '', filterVisible: false,
refreshInterval: null, projectName: null,
};
@@ -48,14 +48,8 @@ function getPhaseGroupForState(state) {
return group ? group.id : null;
}
// Get display group for a task, accounting for review status.
// Approved planning tasks advance to Design; rejected ones go to Blocked.
// Get display group for a task based on its phase state.
function getTaskDisplayGroup(task) {
const review = task.review ? task.review.status : 'pending';
if ((task.state === 'research' || task.state === 'decomposition' || task.state === 'backlog')) {
if (review === 'approved') return 'design';
if (review === 'changes_requested') return 'blocked';
}
return getPhaseGroupForState(task.state);
}
@@ -89,8 +83,7 @@ function renderHeader() {
wipTasks.textContent = filtered.filter(t => wipStates.includes(t.state)).length;
doneTasks.textContent = filtered.filter(t => t.state === 'done').length;
blockedTasks.textContent = filtered.filter(t => t.state === 'blocked').length;
const pendingReview = document.getElementById('pending-review');
if (pendingReview) pendingReview.textContent = filtered.filter(t => (t.state !== 'done' && t.state !== 'blocked') && (!t.review || t.review.status === 'pending')).length;
// Project name display
const projectName = state.projectName;
if (projectName) {
@@ -117,6 +110,13 @@ async function fetchProjectName() {
function renderBoard() {
const board = document.getElementById('board');
// Preserve scroll positions across re-renders so auto-refresh doesn't
// snap columns back to the top while the user is reviewing older tasks.
const prevBodies = Array.from(board.querySelectorAll('.column-body'));
const savedScrolls = prevBodies.map(el => el.scrollTop);
const savedBoardScrollLeft = board.scrollLeft;
const view = document.querySelector('.view.active');
const savedViewScrollTop = view ? view.scrollTop : 0;
const filtered = getFilteredTasks();
const groups = {};
PHASE_GROUPS.forEach(group => { groups[group.id] = []; });
@@ -147,9 +147,25 @@ function renderBoard() {
</div>`;
}).join('');
board.innerHTML = html;
// Restore scroll positions (matched by index — PHASE_GROUPS order is stable).
const newBodies = board.querySelectorAll('.column-body');
newBodies.forEach((el, i) => { if (savedScrolls[i] != null) el.scrollTop = savedScrolls[i]; });
board.scrollLeft = savedBoardScrollLeft;
if (view) view.scrollTop = savedViewScrollTop;
attachCardListeners();
}
const ROLE_LABELS = {
'implement': 'Implement',
'code_review': 'Code Review',
'bug_find': 'Bug Find',
'adversarial_bug_find': 'Adv Bug Find',
'doc_review': 'Doc Review',
'referee': 'Referee',
'loop-implement': 'Loop Impl',
'loop-verify': 'Loop Verify',
};
const ARTIFACT_LABELS = {
'research': 'SPEC.md', 'decomposition': 'DECOMPOSITION.md',
'design': 'DESIGN.md', 'test_design': 'TEST_PLAN.md',
@@ -163,10 +179,6 @@ function renderTaskCard(task) {
const statusClass = task.state === 'done' ? 'done' : task.state === 'blocked' ? 'blocked' : 'in_progress';
const statusIcon = task.state === 'done' ? '✅' : task.state === 'blocked' ? '❌' : '🔄';
const subLabel = getSubLabel(task.state);
const reviewStatus = task.review ? task.review.status : 'pending';
const reviewBadge = reviewStatus === 'approved' ? '<span class="review-badge approved" title="Approved">✅</span>'
: reviewStatus === 'changes_requested' ? '<span class="review-badge changes" title="Changes requested">❌</span>'
: '<span class="review-badge pending" title="Pending review">🟡</span>';
const progressHtml = task.sub_tasks.length > 0
? `<span class="subtask-progress">${task.sub_tasks.filter(st => st.has_verdict && st.verdict_status === 'PASS').length}/${task.sub_tasks.length}</span>`
: '';
@@ -177,17 +189,24 @@ function renderTaskCard(task) {
return `<div class="subtask-item"><span class="subtask-status ${stStatus}">${stIcon}</span><span>${st.name}</span></div>`;
}).join('')}</div>`
: '';
const artifactsHtml = (reviewStatus === 'pending' || reviewStatus === 'changes_requested')
? `<div class="task-card-artifacts">${COLUMNS.filter(col => task.artifacts[col.id]).map(col => {
const artifactsHtml = `<div class="task-card-artifacts">${COLUMNS.filter(col => task.artifacts[col.id]).map(col => {
const label = ARTIFACT_LABELS[col.id] || col.label;
return `<span class="artifact-badge" title="${col.label}">${label}</span>`;
}).join('')}</div>`;
const modelKeys = Object.keys(task.models || {});
const modelsHtml = modelKeys.length > 0
? `<div class="task-card-models">${modelKeys.map(role => {
const m = task.models[role];
const roleLabel = ROLE_LABELS[role] || role;
return `<span class="model-badge" title="${roleLabel}: ${m}">${m}</span>`;
}).join('')}</div>`
: '';
return `<div class="task-card" data-task="${task.name}" data-status="${statusClass}">
<div class="task-card-header"><span class="task-card-name">${escapeHtml(task.display_name)}</span><span style="display:flex;align-items:center;gap:4px">${reviewBadge}<span class="task-card-status ${statusClass}">${statusIcon}</span></span></div>
<div class="task-card-header"><span class="task-card-name">${escapeHtml(task.display_name)}</span><span class="task-card-status ${statusClass}">${statusIcon}</span></div>
<div class="task-card-sublabel">${subLabel}</div>
${task.status_reason ? `<div class="task-card-reason">${escapeHtml(task.status_reason)}</div>` : ''}
${artifactsHtml}
${modelsHtml}
${progressHtml ? `<div class="task-card-footer"><span class="subtask-progress">${progressHtml}</span></div>` : ''}
${subtasksHtml}
</div>`;
@@ -211,33 +230,72 @@ function renderDetail(task) {
return `<span class="detail-artifact"><span class="${cls}">${icon}</span>${col.label}</span>`;
}).join('');
const phaseGroupHtml = phaseGroup ? `<span class="detail-phase-badge" style="background: ${phaseGroupColor}20; color: ${phaseGroupColor}">${PHASE_GROUPS.find(g => g.id === phaseGroup).label}</span>` : '';
const reviewStatus = task.review ? task.review.status : 'pending';
const reviewStatusText = reviewStatus === 'approved' ? '✅ Approved' : reviewStatus === 'changes_requested' ? '❌ Changes Requested' : '🟡 Pending Review';
const reviewComment = task.review && task.review.comment ? `<p class="review-comment">${escapeHtml(task.review.comment)}</p>` : '';
let reviewActionsHtml;
if (reviewStatus === 'approved') {
reviewActionsHtml = '<button class="review-btn revoke" data-task="' + escapeHtml(task.name) + '" data-status="pending">↩ Revoke Approval</button>';
} else if (reviewStatus === 'changes_requested') {
reviewActionsHtml = '<button class="review-btn revoke" data-task="' + escapeHtml(task.name) + '" data-status="pending">↩ Revoke Changes</button>';
} else {
reviewActionsHtml = '<button class="review-btn approve" data-task="' + escapeHtml(task.name) + '" data-status="approved">✅ Approve</button>' +
'<button class="review-btn changes" data-task="' + escapeHtml(task.name) + '" data-status="changes_requested">❌ Request Changes</button>';
}
// Build approval section for approval-gated phases
const APPROVAL_ARTIFACT_MAP = {
'research': { contentKey: 'spec_content', label: 'SPEC.md', phaseName: 'Research' },
'decomposition': { contentKey: 'decomposition_content', label: 'DECOMPOSITION.md', phaseName: 'Decomposition' },
'design': { contentKey: 'design_content', label: 'DESIGN.md', phaseName: 'Design' },
'test_design': { contentKey: 'test_plan_content', label: 'TEST_PLAN.md', phaseName: 'Test Design' },
'code_review': { contentKey: 'code_review_content', label: 'CODE_REVIEW.md', phaseName: 'Code Review' },
};
const awaitingPhase = task.phase_raw ? task.phase_raw.replace(':awaiting_approval', '') : null;
const approvalInfo = awaitingPhase ? APPROVAL_ARTIFACT_MAP[awaitingPhase] : null;
const approvalContent = approvalInfo && approvalInfo.contentKey ? task[approvalInfo.contentKey] : null;
const approvalHtml = task.is_approval_gated
? `<div class="approval-section">
<div class="approval-header">
<span class="approval-icon">🔒</span>
<span class="approval-title">${approvalInfo ? approvalInfo.phaseName : 'Phase'} — Needs Approval</span>
<button class="approve-phase-btn" data-task="${escapeHtml(task.name)}">Approve Phase</button>
</div>
${task.blocker ? `<p class="approval-blocker">${escapeHtml(task.blocker)}</p>` : ''}
${approvalContent ? `<details class="approval-artifact" open>
<summary>${approvalInfo.label} — review content before approving</summary>
<pre class="detail-content-text">${escapeHtml(approvalContent)}</pre>
</details>` : `<p class="approval-missing">${approvalInfo ? approvalInfo.label : 'Artifact'} not yet written — an agent must create it before this phase can complete.</p>`}
</div>`
: '';
// Build transition section for phases that can advance
const TRANSITION_MAP = {
'research:approved': { target: 'decomposition', label: 'Advance to Decomposition' },
'decomposition:approved': { target: 'design', label: 'Advance to Design' },
'design:approved': { target: 'implement', label: 'Advance to Implementation' },
'test_design:approved': { target: 'implement', label: 'Advance to Implementation' },
'implement': { target: 'code_review', label: 'Advance to Code Review' },
'code_review:approved': { target: 'bug_find', label: 'Advance to Bug Finding' },
'bug_find': { target: 'adv_bug_find', label: 'Advance to Adversarial Bug Finding' },
'adv_bug_find': { target: 'doc_review', label: 'Advance to Document Review' },
'doc_review': { target: 'referee', label: 'Advance to Referee' },
};
const nextTransition = TRANSITION_MAP[task.phase_raw];
const transitionHtml = nextTransition
? `<div class="detail-section detail-actions"><h4>🚀 Phase Actions</h4><button class="transition-btn" data-task="${escapeHtml(task.name)}" data-target="${nextTransition.target}">${nextTransition.label}</button></div>`
: '';
// Build artifact editor for missing required artifacts
const requiredArtifact = task.required_artifact_name;
const hasRequiredArtifact = task.artifacts[requiredArtifact];
const artifactEditorHtml = requiredArtifact && !hasRequiredArtifact && !task.is_approval_gated
? `<div class="detail-section detail-artifact-editor">
<h4>📝 Write ${requiredArtifact}</h4>
<p class="artifact-editor-hint">This artifact is required before the task can advance. Write it below and save.</p>
<textarea class="artifact-editor" data-task="${escapeHtml(task.name)}" data-filename="${requiredArtifact}" rows="12" placeholder="# ${requiredArtifact.replace('.md','')}\n\nWrite content here..."></textarea>
<button class="save-artifact-btn" data-task="${escapeHtml(task.name)}" data-filename="${requiredArtifact}">Save ${requiredArtifact}</button>
</div>`
: '';
content.innerHTML = `
<div class="detail-section"><h4>Status</h4><span class="detail-status-badge ${statusClass}">${statusText}</span>${phaseGroupHtml}${task.is_edit_phase ? '<span class="detail-edit-badge">✏️ Edit Allowed</span>' : ''}${task.is_approval_gated ? '<span class="detail-approval-badge">🔒 Requires Approval</span>' : ''}</div>
<div class="detail-section"><h4>Status</h4><span class="detail-status-badge ${statusClass}">${statusText}</span>${phaseGroupHtml}${task.is_edit_phase ? '<span class="detail-edit-badge">✏️ Edit Allowed</span>' : ''}</div>
${statusReason ? `<div class="detail-section detail-status-reason">${escapeHtml(statusReason)}</div>` : ''}
${task.blocker ? `<div class="detail-section detail-blocker"><h4>⚠️ What's Blocking</h4><p class="detail-blocker-text">${escapeHtml(task.blocker)}</p></div>` : ''}
${approvalHtml}
${task.blocker && !task.is_approval_gated ? `<div class="detail-section detail-blocker"><h4>⚠️ What's Blocking</h4><p class="detail-blocker-text">${escapeHtml(task.blocker)}</p></div>` : ''}
${artifactEditorHtml}
${transitionHtml}
${task.phase_guidance ? `<div class="detail-section detail-guidance"><h4>▶ What's Next</h4><pre class="detail-guidance-text">${escapeHtml(task.phase_guidance)}</pre></div>` : ''}
${task.state === 'blocked' && task.blocked_action_items && task.blocked_action_items.length > 0 ? `<div class="detail-section detail-action-items"><h4>📋 Action Items</h4><ul class="detail-action-list">${task.blocked_action_items.map(item => `<li>${escapeHtml(item)}</li>`).join('')}</ul></div>` : ''}
<div class="detail-section"><h4>Artifacts</h4><div class="detail-artifacts">${artifactsHtml}</div></div>
<div class="detail-section"><h4>Review</h4>
<span class="review-badge ${reviewStatus}">${reviewStatusText}</span>
${reviewComment}
<textarea class="review-textarea" id="review-comment-${task.name}" placeholder="Optional comment..." rows="2"></textarea>
<div class="review-actions">
${reviewActionsHtml}
</div>
</div>
${task.sub_tasks.length > 0 ? `<div class="detail-section"><h4>Sub-tasks (${task.sub_tasks.filter(st => st.has_verdict).length}/${task.sub_tasks.length})</h4>
<ul class="detail-subtask-list">${task.sub_tasks.map(st => {
const stStatus = st.has_verdict && st.verdict_status === 'PASS' ? 'pass' : st.has_verdict && st.verdict_status === 'FAIL' ? 'fail' : 'incomplete';
@@ -519,31 +577,62 @@ async function renderAgentTab() {
_startAgentPoller();
}
async function submitReview(taskName, status) {
const textarea = document.getElementById(`review-comment-${taskName}`);
const comment = textarea ? textarea.value : '';
async function approvePhase(taskName) {
try {
const res = await fetch(`/api/task/${taskName}/review`, {
const res = await fetch(`/api/approve/${taskName}`, { method: 'POST' });
const data = await res.json();
if (data.success) {
closeDetail();
await refreshData();
} else {
console.error('Phase approval failed:', data);
}
} catch (err) {
console.error('Phase approval failed:', err);
}
}
async function transitionTask(taskName, target) {
try {
const res = await fetch(`/api/transition/${taskName}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ status, comment }),
body: JSON.stringify({ target }),
});
const data = await res.json();
if (data.success) {
closeDetail();
await refreshData();
} else {
console.error('Transition failed:', data);
}
} catch (err) {
console.error('Review submission failed:', err);
console.error('Transition failed:', err);
}
}
async function saveArtifact(taskName, filename, content) {
try {
const res = await fetch(`/api/write-artifact/${taskName}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ filename, content }),
});
const data = await res.json();
if (data.success) {
closeDetail();
await refreshData();
} else {
console.error('Save artifact failed:', data);
}
} catch (err) {
console.error('Save artifact failed:', err);
}
}
function getFilteredTasks() {
let filtered = [...state.tasks];
if (state.filterPhase !== 'all') filtered = filtered.filter(t => t.state === state.filterPhase);
if (state.filterReview === 'pending') filtered = filtered.filter(t => !t.review || t.review.status === 'pending');
else if (state.filterReview === 'approved') filtered = filtered.filter(t => t.review && t.review.status === 'approved');
else if (state.filterReview === 'changes_requested') filtered = filtered.filter(t => t.review && t.review.status === 'changes_requested');
if (state.filterWave === 'has-waves') filtered = filtered.filter(t => t.sub_tasks.length > 0);
else if (state.filterWave === 'no-waves') filtered = filtered.filter(t => t.sub_tasks.length === 0);
if (state.searchQuery) {
@@ -622,15 +711,26 @@ function setupUI() {
document.getElementById('btn-close-detail').addEventListener('click', closeDetail);
document.getElementById('detail-overlay').addEventListener('click', (e) => { if (e.target === e.currentTarget) closeDetail(); });
document.getElementById('filter-phase').addEventListener('change', (e) => { state.filterPhase = e.target.value; renderCurrentView(); });
document.getElementById('filter-review').addEventListener('change', (e) => { state.filterReview = e.target.value; renderCurrentView(); });
document.getElementById('filter-wave').addEventListener('change', (e) => { state.filterWave = e.target.value; renderCurrentView(); });
document.getElementById('search-input').addEventListener('input', (e) => { state.searchQuery = e.target.value; renderCurrentView(); });
document.addEventListener('click', (e) => {
const btn = e.target.closest('.review-btn');
if (btn) {
const taskName = btn.dataset.task;
const status = btn.dataset.status;
if (taskName && status) submitReview(taskName, status);
const approveBtn = e.target.closest('.approve-phase-btn');
if (approveBtn) {
const taskName = approveBtn.dataset.task;
if (taskName) approvePhase(taskName);
}
const transitionBtn = e.target.closest('.transition-btn');
if (transitionBtn) {
const taskName = transitionBtn.dataset.task;
const target = transitionBtn.dataset.target;
if (taskName && target) transitionTask(taskName, target);
}
const saveBtn = e.target.closest('.save-artifact-btn');
if (saveBtn) {
const taskName = saveBtn.dataset.task;
const filename = saveBtn.dataset.filename;
const editor = document.querySelector(`.artifact-editor[data-task="${taskName}"][data-filename="${filename}"]`);
if (taskName && filename && editor) saveArtifact(taskName, filename, editor.value);
}
});
}
-10
View File
@@ -26,7 +26,6 @@
<span class="stat-wip">WIP: <strong id="wip-tasks">0</strong></span>
<span class="stat-done">Done: <strong id="done-tasks">0</strong></span>
<span class="stat-blocked">Blocked: <strong id="blocked-tasks">0</strong></span>
<span class="stat-pending">Pending: <strong id="pending-review">0</strong></span>
</div>
<div class="header-controls">
<button class="btn btn-icon" id="btn-refresh" title="Refresh">↻</button>
@@ -50,15 +49,6 @@
<option value="blocked">Blocked</option>
</select>
</div>
<div class="filter-group">
<label>Review:</label>
<select id="filter-review">
<option value="all">All</option>
<option value="pending">Pending</option>
<option value="approved">Approved</option>
<option value="changes_requested">Changes Requested</option>
</select>
</div>
<div class="filter-group">
<label>Waves:</label>
<select id="filter-wave">
+42 -13
View File
@@ -305,21 +305,50 @@ kbd {
::-webkit-scrollbar-track { background: transparent; }
::-webkit-scrollbar-thumb { background: var(--beerus); border-radius: 3px; }
::-webkit-scrollbar-thumb:hover { background: var(--border-active); }
.review-badge { font-size: 12px; margin-left: 4px; }
.review-badge.approved { color: var(--success); }
.review-badge.changes_requested { color: var(--error); }
.review-badge.pending { color: var(--warning); }
.review-actions { display: flex; gap: 10px; margin-top: 10px; }
.review-btn { padding: 7px 16px; border: 1px solid var(--border-color); border-radius: var(--radius-sm); cursor: pointer; font-size: 12px; transition: all 0.15s; background: var(--bg-card); color: var(--text-primary); font-family: inherit; font-weight: 500; }
.review-btn.approve { border-color: var(--success); color: var(--success); }
.review-btn.approve:hover { background: var(--success-bg); }
.review-btn.changes { border-color: var(--error); color: var(--error); }
.review-btn.changes:hover { background: var(--error-bg); }
.review-comment { font-size: 12px; color: var(--text-secondary); padding: 10px 12px; background: var(--bg-primary); border-radius: var(--radius-sm); margin-top: 6px; border-left: 2px solid var(--border-color); }
.review-textarea { width: 100%; margin-top: 8px; padding: 10px; background: var(--bg-primary); border: 1px solid var(--border-color); border-radius: var(--radius-sm); color: var(--text-primary); font-size: 12px; font-family: inherit; resize: vertical; }
.review-textarea:focus { outline: none; border-color: var(--primary); box-shadow: 0 0 0 2px rgba(138,180,248,0.2); }
.approve-phase-btn { display: inline-block; margin-left: 8px; padding: 4px 12px; border: 1px solid var(--warning); border-radius: var(--radius-sm); cursor: pointer; font-size: 11px; font-weight: 500; background: var(--warning-bg); color: var(--warning); font-family: inherit; transition: all 0.15s; }
.approve-phase-btn:hover { filter: brightness(1.1); }
/* Approval section — prominent card for approval-gated tasks */
.approval-section {
margin: 12px 0;
padding: 14px 16px;
border: 2px solid var(--warning);
border-radius: var(--radius-md);
background: var(--warning-bg);
}
.approval-header {
display: flex;
align-items: center;
gap: 8px;
margin-bottom: 8px;
}
.approval-icon { font-size: 18px; }
.approval-title { font-size: 14px; font-weight: 600; flex: 1; }
.approval-blocker { font-size: 12px; color: var(--text-secondary); margin: 0 0 8px 0; padding: 6px 10px; background: var(--bg-card); border-radius: var(--radius-sm); }
.approval-missing { font-size: 12px; color: var(--text-secondary); margin: 8px 0 0 0; font-style: italic; }
.approval-artifact { margin-top: 8px; font-size: 12px; }
.approval-artifact summary { cursor: pointer; font-weight: 500; padding: 4px 0; color: var(--text-primary); }
.approval-artifact summary:hover { color: var(--primary); }
.task-card-artifacts { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 6px; }
.artifact-badge { font-size: 10px; padding: 2px 8px; background: var(--bg-primary); border: 1px solid var(--border-color); border-radius: 4px; color: var(--text-secondary); font-family: 'SF Mono', 'Fira Code', monospace; font-weight: 500; }
.task-card-models { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 6px; }
.model-badge { font-size: 10px; padding: 1px 6px; background: #e3f2fd; border: 1px solid #90caf9; border-radius: 4px; color: #1565c0; font-family: 'SF Mono', 'Fira Code', monospace; font-weight: 500; }
/* Transition button — advance to next phase */
.transition-btn { display: inline-block; padding: 6px 16px; border: 1px solid var(--primary); border-radius: var(--radius-sm); cursor: pointer; font-size: 12px; font-weight: 500; background: rgba(74, 144, 226, 0.1); color: var(--primary); font-family: inherit; transition: all 0.15s; }
.transition-btn:hover { filter: brightness(1.15); background: rgba(74, 144, 226, 0.2); }
.detail-actions { margin: 8px 0; padding: 10px 12px; background: var(--bg-card); border-radius: var(--radius-md); border: 1px solid var(--border-color); }
.detail-actions h4 { margin: 0 0 8px 0; font-size: 12px; color: var(--text-secondary); font-weight: 500; }
/* Artifact editor — inline textarea for writing missing artifacts */
.detail-artifact-editor { margin: 8px 0; padding: 12px; background: var(--bg-card); border: 1px solid var(--primary); border-radius: var(--radius-md); }
.detail-artifact-editor h4 { margin: 0 0 4px 0; font-size: 12px; font-weight: 600; }
.artifact-editor-hint { font-size: 11px; color: var(--text-secondary); margin: 0 0 8px 0; }
.artifact-editor { width: 100%; padding: 10px; border: 1px solid var(--border-color); border-radius: var(--radius-sm); background: var(--bg-primary); color: var(--text-primary); font-family: 'SF Mono', 'Fira Code', monospace; font-size: 12px; line-height: 1.5; resize: vertical; box-sizing: border-box; }
.artifact-editor:focus { outline: none; border-color: var(--primary); }
.save-artifact-btn { display: inline-block; margin-top: 8px; padding: 6px 16px; border: 1px solid var(--success); border-radius: var(--radius-sm); cursor: pointer; font-size: 12px; font-weight: 500; background: rgba(74, 208, 120, 0.1); color: var(--success); font-family: inherit; transition: all 0.15s; }
.save-artifact-btn:hover { filter: brightness(1.15); background: rgba(74, 208, 120, 0.2); }
@media (max-width: 768px) {
.header { flex-wrap: wrap; gap: 8px; }
+147 -2
View File
@@ -51,6 +51,21 @@ TASK_STATE_ARTIFACT = {
TASK_STATES = list(TASK_STATE_ARTIFACT.keys())
def next_transition(phase_raw: str) -> str | None:
transitions = {
"research:approved": "decomposition",
"decomposition:approved": "design",
"design:approved": "implement",
"test_design:approved": "implement",
"implement": "code_review",
"code_review:approved": "bug_find",
"bug_find": "adv_bug_find",
"adv_bug_find": "doc_review",
"doc_review": "referee",
"referee": "complete",
}
return transitions.get(phase_raw)
MAX_POST_BODY = 65536 # 64KB
MAX_REVIEW_COMMENT_LENGTH = 4096
CACHE_TTL = 1.0 # seconds
@@ -117,6 +132,15 @@ class DashboardHandler(SimpleHTTPRequestHandler):
self._send_error(413, "Payload too large")
return
self._handle_review(task_name)
elif self.path.startswith("/api/approve/"):
task_name = unquote(self.path.split("/api/approve/")[1])
self._handle_phase_approval(task_name)
elif self.path.startswith("/api/transition/"):
task_name = unquote(self.path.split("/api/transition/")[1])
self._handle_transition(task_name)
elif self.path.startswith("/api/write-artifact/"):
task_name = unquote(self.path.split("/api/write-artifact/")[1])
self._handle_write_artifact(task_name)
else:
self._send_error(404, "Not found")
@@ -193,6 +217,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"name": t.name,
"display_name": t.display_name,
"state": t.state.value,
"phase_raw": t.phase_raw,
"status_reason": t.status_reason,
"artifacts": {s: TASK_STATE_ARTIFACT[s] in t.artifacts for s in TASK_STATES},
"sub_tasks": [
@@ -201,8 +226,14 @@ class DashboardHandler(SimpleHTTPRequestHandler):
],
"verdict_content": t.verdict_content,
"bug_report_content": t.bug_report_content,
"adversarial_bug_report_content": t.adversarial_bug_report_content,
"doc_review_content": t.doc_review_content,
"design_content": t.design_content,
"spec_content": t.spec_content,
"decomposition_content": t.decomposition_content,
"code_review_content": t.code_review_content,
"test_plan_content": t.test_plan_content,
"implementation_content": t.implementation_content,
"parent_spec_content": t.parent_spec_content,
"vram_config_content": t.vram_config_content,
"blocked_action_items": t.blocked_action_items,
@@ -214,7 +245,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"is_approval_gated": t.is_approval_gated,
"blocker": t.blocker,
"waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in t.waves],
"review": self._get_review_status(t.name),
"models": t.models,
}
for t in tasks
]
@@ -299,6 +330,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"name": task.name,
"display_name": task.display_name,
"state": task.state.value,
"phase_raw": task.phase_raw,
"status_reason": task.status_reason,
"artifacts": {s: TASK_STATE_ARTIFACT[s] in task.artifacts for s in TASK_STATES},
"sub_tasks": [
@@ -307,12 +339,22 @@ class DashboardHandler(SimpleHTTPRequestHandler):
],
"verdict_content": task.verdict_content,
"bug_report_content": task.bug_report_content,
"adversarial_bug_report_content": task.adversarial_bug_report_content,
"doc_review_content": task.doc_review_content,
"design_content": task.design_content,
"spec_content": task.spec_content,
"decomposition_content": task.decomposition_content,
"parent_spec_content": task.parent_spec_content,
"vram_config_content": task.vram_config_content,
"blocked_action_items": task.blocked_action_items,
"unblock_instructions": task.unblock_instructions,
"phase_guidance": task.phase_guidance,
"required_artifact_name": task.required_artifact_name,
"next_phase_name": task.next_phase_name,
"is_edit_phase": task.is_edit_phase,
"is_approval_gated": task.is_approval_gated,
"blocker": task.blocker,
"waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in task.waves],
"review": self._get_review_status(task.name),
}
self._send_json(task_data)
@@ -388,6 +430,109 @@ class DashboardHandler(SimpleHTTPRequestHandler):
except json.JSONDecodeError:
self._send_error(400, "Invalid JSON")
def _handle_phase_approval(self, task_name: str):
"""Approve a task's current phase (--approve) via status.py."""
if not self._validate_task_name(task_name):
self._send_error(400, "Invalid task name")
return
project_root = self.project_root
if not project_root:
self._send_error(503, "Not in automaton project")
return
status_py = Path.home() / ".automaton" / "scripts" / "status.py"
if not status_py.exists():
self._send_error(500, "status.py not found")
return
import subprocess
result = subprocess.run(
[sys.executable, str(status_py), "--approve", "--task", task_name,
"--project", str(project_root)],
capture_output=True, text=True, timeout=30,
)
_invalidate_task_cache()
if result.returncode == 0:
self._send_json({"success": True, "message": result.stdout.strip()})
else:
self._send_error(400, result.stderr.strip() or result.stdout.strip())
def _handle_transition(self, task_name: str):
if not self._validate_task_name(task_name):
self._send_error(400, "Invalid task name")
return
project_root = self.project_root
if not project_root:
self._send_error(503, "Not in automaton project")
return
status_py = Path.home() / ".automaton" / "scripts" / "status.py"
if not status_py.exists():
self._send_error(500, "status.py not found")
return
import subprocess
content_length = int(self.headers.get('Content-Length', 0))
target = None
if content_length > 0:
try:
body = json.loads(self.rfile.read(content_length))
target = body.get("target")
except (json.JSONDecodeError, UnicodeDecodeError):
pass
if not target:
tasks = _get_cached_tasks(project_root)
task = next((t for t in tasks if t.name == task_name), None)
if not task:
self._send_error(404, "Task not found")
return
target = next_transition(task.phase_raw)
if not target:
self._send_json({"success": False, "message": "Cannot determine next transition from current phase"})
return
result = subprocess.run(
[sys.executable, str(status_py), "--transition", target, "--task", task_name,
"--project", str(project_root)],
capture_output=True, text=True, timeout=30,
)
_invalidate_task_cache()
if result.returncode == 0:
self._send_json({"success": True, "message": f"Transitioned to {target}: {result.stdout.strip()}"})
else:
self._send_json({"success": False, "message": result.stderr.strip() or result.stdout.strip()})
def _handle_write_artifact(self, task_name: str):
if not self._validate_task_name(task_name):
self._send_error(400, "Invalid task name")
return
project_root = self.project_root
if not project_root:
self._send_error(503, "Not in automaton project")
return
content_length = int(self.headers.get('Content-Length', 0))
if content_length > MAX_POST_BODY:
self._send_error(413, "Payload too large")
return
if content_length == 0:
self._send_error(400, "Empty request body")
return
try:
body = json.loads(self.rfile.read(content_length))
filename = body.get("filename", "")
content = body.get("content", "")
if not filename or not filename.endswith(".md"):
self._send_error(400, "Invalid filename — must be a .md file")
return
tasks = _get_cached_tasks(project_root)
task = next((t for t in tasks if t.name == task_name), None)
if not task:
self._send_error(404, "Task not found")
return
artifact_path = Path(task.folder_path) / filename
artifact_path.write_text(content, encoding="utf-8")
_invalidate_task_cache()
self._send_json({"success": True, "message": f"Written {filename}"})
except (json.JSONDecodeError, UnicodeDecodeError):
self._send_error(400, "Invalid JSON")
except (OSError, IOError) as e:
self._send_error(500, f"Failed to write file: {e}")
def _serve_review_summary(self):
project_root = self.project_root
if not project_root:
+13 -2
View File
@@ -33,8 +33,8 @@ To disable auto-detection and use manual values:
Settings for the LLM model being used.
- **Model**: auto # Use auto-detection from API config, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
- **Model**: omlx/Ornith-1.0-35B-4bit-mlx # Local LLM (opencode provider); used as the Implement role
- **Override context window**: 32768 # Matches opencode.json limit.context for ornith
### Auto-detection
@@ -51,6 +51,17 @@ To disable auto-detection and use manual values:
- **Override context window**: 128k
```
## Available Models
Models available for model-divergence enforcement. This file is managed by `scripts/detect_models.py`. In single-LLM mode (0-1 models), no hard blocks are enforced. In multi-LLM mode (2+ models), the conflict matrix enforces role-model separation.
- **Default**: omlx/Ornith-1.0-35B-4bit-mlx # Used when no role-specific binding is set
- **Advised**: true # Recommend a second model in single-LLM mode
No additional models are configured in the manifest. To add models:
1. Run `python3 ~/.automaton/scripts/detect_models.py --write` to auto-detect from opencode.json and localhost endpoints.
2. Or manually create `~/.automaton/models.json` (see `design/framework/technical.md` §2 for schema).
## System Requirements
Requirements for the environment the framework runs in.
+33
View File
@@ -0,0 +1,33 @@
# Framework Backlog
Status: **v1 draft 2026-06-25**. These items are captured for the self-improvement loop (or manual pickup) once the self-improvement loop is running and the design docs are in place.
This backlog mirrors the pattern of `design/loops/BACKLOG.md` but for framework-level agent features. The self-improvement loop's `work_source.area` can be set to `"framework"` to pull from here.
---
## v1 (Priority)
| ID | Item | Description | Dependencies |
|---|---|---|---|
| FW-1 | **agent-tab-real-roles** | Redesign the dashboard Agent tab: replace 4 fake `AGENT_TYPE_META` types with two sections — (1) Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) showing active/inactive based on current task phases; (2) Scheduled Jobs (cleanup, loop-tick, rule-scan, rule-review) using real `job.kind` from `/api/scheduled`. New `/api/phase-roles` endpoint. | None |
| FW-2 | **rule-proposer-agent** | Daily scheduled standalone agent that scans completed tasks since last run, reads their `BUG_REPORT.md` / `ADVERSARIAL_BUG_REPORT.md` / `VERDICT.md`, proposes new rules with concrete examples to `RULE_PROPOSALS.md`. Uses direct harness invocation (reuses `loop-runner._invoke_harness`). State in `.state.rule-scan`. Requires different LLM than the tasks it reviews (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding) |
| FW-3 | **rule-reviewer-agent** | Monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, missing examples — writes `RULE_REVIEW.md`. Uses direct harness invocation. State in `.state.rule-review`. Must use different LLM than Rule Proposer (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding), `rule-proposer-agent` |
---
## Deferred / Tier 2 (Post-v1)
| ID | Item | Description | Notes |
|---|---|---|---|
| FW-4 | **model-divergence-enforcement** | Manifest (`models.json`), mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model + `{model}` substitution, audit category, dashboard badges. This is a **manual task**, not a backlog item — tracked separately. | See separate task `model-divergence-enforcement` |
| FW-5 | **dashboard-agent-tab-v2** | Enhance Agent tab with: role details (click → task list), schedule management (enable/disable from UI), last-run timestamps, per-agent logs. | Requires FW-1 |
| FW-6 | **rule-enforcement-gate** | Pre-edit hook (`--can-edit --rule <name>`) that enforces rules from `.rules.md` before edits. Separate from rule agents (which only propose/review). | Requires FW-2, FW-3 |
---
## Notes
- **Dependency on `model-divergence-enforcement`**: FW-2 and FW-3 are designed with model-binding in their configs (per-role `model` in their schedule config), but hard-block enforcement is deferred until the model-divergence feature ships. The design docs note the dependency; implementation tasks in the backlog carry a "depends on" annotation.
- **Self-improvement loop**: Once the self-improvement loop is running with `work_source.area = "framework"`, it will pick up items from this backlog automatically. The loop template `templates/loops/self-improvement/loop.json` already supports `work_source.area`.
- **FW-4 is NOT in this backlog** — it's a manual parent task created via `status.py --create-task model-divergence-enforcement` and decomposed into subtasks.
+67
View File
@@ -0,0 +1,67 @@
# Framework Agent Features — Design Index
Status: **v1 draft 2026-06-25**. Design for three framework-level agent features: rule agents (proposer + reviewer), Agent tab redesign, and model-divergence enforcement.
## What this is
The Automaton framework's agent-feature layer: scheduled standalone agents for rule maintenance (Rule Proposer, Rule Reviewer), a redesigned dashboard Agent tab showing real phase roles instead of fake types, and a model-divergence enforcement system that prevents conflict-of-interest by ensuring different LLM roles use different models.
Built on top of the existing `status.py` phase machine and loop infrastructure — no second enforcement surface.
## Why
- `.rules.md` is maintained manually; failure patterns from completed tasks are not systematically captured.
- The dashboard Agent tab shows 4 made-up agent types (`completed_task_archiver`, `single`, `audit`, `backlog`) — not the real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator).
- Conflict-of-interest: two sessions with the same LLM playing bug finder and adversarial bug finder (or implementer and reviewer) don't yield useful results. The framework needs computational enforcement of model divergence.
## Documents
- [`functional.md`](functional.md) — what v1 does, roles, triggers, outputs, schedules, conflict-of-interest, success criteria. **Read this first.**
- [`technical.md`](technical.md) — the implementation contract: file map, state schemas, `status.py` flags, runner flows, dashboard API changes, test coverage. **Read this if you're implementing.**
- [`BACKLOG.md`](BACKLOG.md) — v1 work queue (3 items) and deferred items. The self-improvement loop's work queue when `work_source.area = "framework"`.
## v1 scope (draft)
Three feature areas, each with a clear boundary:
1. **Model-Divergence Enforcement** — Foundational layer. `models.json` manifest, mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, audit category, dashboard badges. Enables the other two features.
2. **Rule Agents** — Two standalone scheduled agents (NOT task lifecycle phases):
- **Rule Proposer**: Daily scan of completed tasks → `RULE_PROPOSALS.md` with concrete examples.
- **Rule Reviewer**: Monthly consolidation of `.rules.md` → `RULE_REVIEW.md`.
- Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers (mirror `--install-cleanup-schedule`).
3. **Agent Tab Redesign** — Dashboard `/agent` view:
- **Phase Roles section**: 6 roles from `.agent.md`, active/inactive based on current task phases.
- **Scheduled Jobs section**: Real job kinds from `/api/scheduled` (cleanup, loop, rule-scan, rule-review), replacing 4 fake `AGENT_TYPE_META` types.
- New `/api/phase-roles` endpoint.
## Locked decisions (referenced as `(Fn)` in functional/technical)
| ID | Decision |
|---|---|
| F1 | Rule agents are **standalone scheduled agents**, NOT task lifecycle phases. They don't block task completion. |
| F2 | Rule Proposer runs **daily**; Rule Reviewer runs **monthly**. OS-native schedulers (launchd/cron/schtasks). |
| F3 | Rule agents use **direct harness invocation** (reuse `loop-runner._invoke_harness`), not loop infrastructure. |
| F4 | Conflict-of-interest: Rule Reviewer **must use different LLM** than Rule Proposer. Rule agents must use different LLM than tasks they review. Enforcement **deferred** until model-divergence ships (F5). |
| F5 | Model-divergence enforcement is a **foundational layer** shipped first (manual task, not backlog). Rule agents designed with model-binding config but hard-block enforced later. |
| F6 | Agent tab has **two sections**: Phase Roles (from `.agent.md` + task state) + Scheduled Jobs (from `/api/scheduled`). |
| F7 | Phase Roles section needs **new `/api/phase-roles` endpoint** mapping active tasks to their phase roles. |
| F8 | Scheduled Jobs section uses **real `job.kind`** (cleanup, loop, rule-scan, rule-review) — remove `AGENT_TYPE_META` fake types. |
| F9 | State files: `.state.rule-scan` (last scanned task, timestamp, proposed rules), `.state.rule-review` (last run timestamp). |
| F10 | New `status.py` flags: `--rule-scan`, `--install-rule-scan-schedule`, `--rule-review`, `--install-rule-review-schedule` (mirror `--cleanup-done` / `--install-cleanup-schedule`). |
## Relationship to loops
- `design/loops/` is the loop engineering system (state-enforced unattended work).
- `design/framework/` is framework-level agent features (rule maintenance, observability, model governance).
- The self-improvement loop (`templates/loops/self-improvement/`) can drive `design/framework/` work by setting `work_source.area = "framework"` — no code change needed (see `loop-runner.py:_find_work_backlog`).
- Model-divergence enforcement (F5) is a prerequisite for hard-blocking conflict-of-interest in rule agents and loops.
## Post-v1
After v1 lands and the self-improvement loop is running:
- Rule enforcement gate (`--can-edit --rule <name>`) — separate feature.
- Dashboard Agent tab v2: role details, schedule management, per-agent logs.
- Rule agents gain comprehension-debt tracking (`last-read-sha` per rule file).
+335
View File
@@ -0,0 +1,335 @@
# Framework Agent Features — Functional Design
Status: v1 (draft 2026-06-25). Supersedes any prior informal discussions of rule agents or Agent tab redesign.
Audience: framework maintainers (currently: the human and one AI assistant). After handoff the self-improvement loop is also an audience — designs must be legible to a fresh-context LLM verifier.
## 1. Problem
Three gaps in the framework's agent layer:
1. **`.rules.md` is maintained manually.** Failure patterns from completed tasks (`BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md`) are not systematically captured into rules. The Self-Improvement section of `.rules.md` (lines 33-36) says "Add one rule per observed failure mode with a concrete example" and "Consolidate contradictions monthly" — but no agent does this. It relies on the human or a session agent remembering.
2. **The dashboard Agent tab shows fake agent types.** `AGENT_TYPE_META` in `dashboard.js:330-334` defines 4 types (`completed_task_archiver`, `single`, `audit`, `backlog`) that are loop work-source kinds, not real agent roles. The 6 real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) are not surfaced.
3. **No conflict-of-interest enforcement on model divergence.** `status.py:1432-1436` enforces that the reviewer session differs from the implementer session (via `.state.implementer`), but there is no enforcement that the *model* playing bug-finder differs from the model playing adversarial bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping.
## 2. Goals
v1 — **the framework maintains its own rules, surfaces real agent roles, and computationally enforces model divergence**:
1. **Rule Proposer**: a daily scheduled standalone agent that scans completed tasks since its last run, reads their failure artifacts, and proposes new rules with concrete examples to `RULE_PROPOSALS.md`.
2. **Rule Reviewer**: a monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, rules missing examples — and writes `RULE_REVIEW.md`.
3. **Agent Tab Redesign**: replace 4 fake `AGENT_TYPE_META` types with two sections — Phase Roles (6 roles from `.agent.md`, active/inactive based on current task phases) and Scheduled Jobs (real `job.kind` from `/api/scheduled`).
4. **Model-Divergence Enforcement**: a `models.json` manifest, mode detection (single vs multi-LLM), a conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, an audit category, and dashboard badges.
5. **Conflict-of-interest for rule agents**: Rule Reviewer must use a different LLM than Rule Proposer. Both must use different LLMs than the tasks they review. Enforcement is *designed now* but *hard-blocked only after model-divergence ships* (F4, F5).
## 3. Non-Goals (v1)
- **Rule enforcement during task execution.** Rule agents only *propose* and *review* rules; they do not block edits. A future `--can-edit --rule <name>` gate (FW-6) is separate.
- **Auto-approving rules.** A human always reviews `RULE_PROPOSALS.md` and `RULE_REVIEW.md` before rules are merged into `.rules.md`. No auto-merge path in v1.
- **Rule agents as task lifecycle phases.** Rule agents are *standalone scheduled agents* (F1). They do not block task completion. They are not phases in the state machine.
- **A new agent harness.** Rule agents use direct harness invocation (reusing `loop-runner._invoke_harness`). No new runtime.
- **Network-fetched dependencies.** New code is Python stdlib only. No new pip installs.
- **Model capability inspection.** The framework never inspects model capability, provider, or size (D8). It only tracks *which* model fills *which* role and enforces the conflict matrix.
## 4. Model-Divergence Enforcement (foundational layer)
Shipped first as a manual task (`model-divergence-enforcement`), not a backlog item. Enables conflict-of-interest hard-blocking for rule agents and loops.
### 4.1 Manifest: `models.json`
A new file at `~/.automaton/models.json` (or `{project}/.automaton/models.json`):
```json
{
"default": "glm-4.6",
"advised": true,
"models": [
{"name": "glm-4.6", "provider": "opencode", "context_window": 131072, "location": "remote"},
{"name": "qwen3-coder", "provider": "opencode", "context_window": 131072, "location": "remote"},
{"name": "llama-3.3-70b", "provider": "localhost", "context_window": 32768, "location": "http://localhost:8080"}
]
}
```
- `default`: the model used when no role-specific binding is set.
- `advised`: if `true`, the framework prints a one-time advisory in single-LLM mode recommending a second model for conflict-of-interest roles, then goes silent.
- `models[]`: the roster. `location` is `"remote"` or a localhost URL for probing.
### 4.2 Mode Detection
- **0-1 models** in `models.json` (or file missing) → **single-LLM mode**. Advisory once (if `advised: true`), then silent. No hard blocks.
- **2+ models** → **multi-LLM mode**. Hard-block on conflict-matrix violations. Auto-assign next-available non-conflicting model on conflict; refuse only if no non-conflicting model exists.
- **Missing file** → single-LLM mode (backward compatible). Existing behavior preserved.
### 4.3 Conflict Matrix (locked)
| Role | Must differ from |
|---|---|
| `code_review` | `implement` |
| `bug_find` | `implement` |
| `adversarial_bug_find` | `implement`, `bug_find` |
| `referee` | `implement`, `bug_find`, `adversarial_bug_find` |
| `loop-verify` | `loop-implement` |
`doc_review`, `code_review`, and `bug_find` are independent of each other (not conflicts). Only `bug_find` ↔ `adversarial_bug_find` conflicts (they are adversary pairs).
### 4.4 Auto-Assignment (multi-LLM mode)
1. Default model → assigned to `implement` (and `loop-implement`).
2. On conflict, pick the next-available model from `models[]` that does not conflict.
3. User override: `loop.json` `roles.<role>.model` or `status.py --transition --model <name>`.
4. Refuse only if no non-conflicting model exists.
### 4.5 State: `.state.models`
Each task gets `{task}/.state.models` recording which model filled which role:
```json
{"implement": "glm-4.6", "code_review": "qwen3-coder", "bug_find": "qwen3-coder", "adversarial_bug_find": "llama-3.3-70b", "referee": "llama-3.3-70b"}
```
`status.py --transition --model <name>` records the model for the role being transitioned into. `--claim` in multi-LLM mode checks the conflict matrix against `.state.models` and refuses on violation.
### 4.6 Loop Integration
`loop.json` gains per-role `model` and `harness.command` with `{model}` substitution:
```json
"roles": {
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
}
```
`loop-runner.py:_invoke_harness` substitutes `{model}` into the harness command. `--check-gate` enforces `loop-verify` ≠ `loop-implement` model in multi-LLM mode.
### 4.7 Audit + Dashboard
- `status.py --audit` gains a `model_divergence` category: flags tasks where `.state.models` violates the conflict matrix.
- Dashboard task cards show model badges (one per role filled).
## 5. Rule Proposer
A **standalone scheduled agent** (F1) that proposes new rules from completed-task failure patterns.
### 5.1 Trigger
Daily, via OS-native scheduler (mirrors `--install-cleanup-schedule`). `status.py --install-rule-scan-schedule [--interval 86400]` installs the schedule unit. Manual: `status.py --rule-scan`.
### 5.2 Inputs
- `.state.rule-scan`: state file tracking the last scanned task and timestamp.
- Completed tasks (in `tasks/complete/` or tasks with `.state` phase `complete`) that were completed since the last scan.
- For each such task: `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md` (if present).
- Current `.rules.md` (to deduplicate against existing rules).
### 5.3 Output
`RULE_PROPOSALS.md` (at `~/.automaton/RULE_PROPOSALS.md` in framework mode, or `{project}/.automaton/RULE_PROPOSALS.md` in project mode). Format:
```markdown
# Rule Proposals — {date}
## Proposed Rule: {title}
**Source**: tasks/{task-name}/VERDICT.md
**Pattern**: {one-line description of the failure mode}
**Example**:
{concrete code/config snippet from the task}
**Proposed rule text**:
{the rule as it would appear in .rules.md}
---
```
Proposals are *appended* per run. A human reviews and merges accepted rules into `.rules.md`. No auto-merge (Non-Goal).
### 5.4 LLM Session
Direct harness invocation (F3): `status.py --rule-scan` reuses `loop-runner._invoke_harness` to spawn one LLM session with a prompt that includes the failure artifacts and current rules. The session proposes rules in the `RULE_PROPOSALS.md` format.
### 5.5 Conflict-of-Interest
The Rule Proposer's LLM must differ from the implementer + bug-hunter + adversarial-bug-hunter models of the tasks it scans. This prevents the model that made the bug from proposing the rule about its own bug.
**Enforcement deferred** (F4): until model-divergence ships, the Rule Proposer runs with the default model. The design includes a `model` field in its schedule config; hard-blocking activates once `models.json` exists and multi-LLM mode is detected.
### 5.6 State: `.state.rule-scan`
```json
{"last_scan_at": "2026-06-25T10:00:00Z", "last_scanned_task": "fix-context-sizing", "proposals_count": 3}
```
## 6. Rule Reviewer
A **standalone scheduled agent** (F1) that consolidates `.rules.md` periodically.
### 6.1 Trigger
Monthly, via OS-native scheduler. `status.py --install-rule-review-schedule [--interval 2592000]` installs the schedule unit. Manual: `status.py --rule-review`.
### 6.2 Inputs
- `.state.rule-review`: state file tracking the last run timestamp.
- Current `.rules.md` (full file).
- Recent `RULE_PROPOSALS.md` entries (since last review).
- Recent completed-task summaries (last 30 days) for context on stale rules.
### 6.3 Output
`RULE_REVIEW.md` (at `~/.automaton/RULE_REVIEW.md` or project equivalent). Format:
```markdown
# Rule Review — {date}
## Contradictions Found
- Rule A ("...") contradicts Rule B ("..."). Suggested resolution: {merge/drop/keep A}.
## Stale Rules (no observed instance in last 30 days)
- Rule C ("..."). Suggested action: drop or annotate as low-priority.
## Rules Missing Examples
- Rule D ("..."). Suggested example: {from a recent task}.
## Merge Candidates
- Rules E and F overlap. Suggested merged text: {...}.
---
```
A human reviews and applies accepted changes to `.rules.md`. No auto-apply (Non-Goal).
### 6.4 LLM Session
Direct harness invocation (F3), same as Rule Proposer. One LLM session with a prompt that includes the full `.rules.md` and recent proposals.
### 6.5 Conflict-of-Interest
The Rule Reviewer's LLM **must differ from the Rule Proposer's LLM** (F4). The Proposer proposes (bias toward adding); the Reviewer consolidates (bias toward pruning). Same model = self-review = rubber-stamping.
**Enforcement deferred** until model-divergence ships.
### 6.6 State: `.state.rule-review`
```json
{"last_review_at": "2026-06-25T10:00:00Z", "contradictions_found": 2, "stale_rules": 5, "merges_suggested": 1}
```
## 7. Agent Tab Redesign
### 7.1 Current State (to be replaced)
`dashboard.js:330-334` defines `AGENT_TYPE_META` with 4 fake types:
- `completed_task_archiver` — actually the cleanup scheduled job.
- `single` — actually a loop with `work_source.kind = "single"`.
- `audit` — actually a loop with `work_source.kind = "audit"`.
- `backlog` — actually a loop with `work_source.kind = "backlog"`.
These are loop work-source kinds, not agent roles. They conflate two different concepts.
### 7.2 New Design: Two Sections
**Section 1 — Phase Roles**
Shows the 6 phase roles from `.agent.md` Agent Configuration:
| Role ID | Phases | Active when |
|---|---|---|
| `researcher` | research, decomposition, design, test_design | A task is in one of these phases |
| `implementer` | implement | A task is in `implement` phase |
| `code-reviewer` | code_review | A task is in `code_review` phase |
| `bug-hunter` | bug_find, adversarial_bug_find | A task is in one of these phases |
| `referee` | referee | A task is in `referee` phase |
| `orchestrator` | new, complete, human_intervention | A task is in one of these phases |
Each role card shows: role icon, role label, status (Active/Idle — based on whether any task is in that role's phases), and the count of tasks in that role's phases. Clicking a role filters the task list to tasks in that role's phases.
**Data source**: new `/api/phase-roles` endpoint. Returns:
```json
{
"roles": [
{"id": "researcher", "label": "Researcher", "icon": "🔬", "phases": ["research", "decomposition", "design", "test_design"], "active_tasks": 2, "status": "active"},
{"id": "implementer", "label": "Implementer", "icon": "⚙️", "phases": ["implement"], "active_tasks": 1, "status": "active"},
...
]
}
```
**Section 2 — Scheduled Jobs**
Shows real scheduled jobs from `/api/scheduled`, using `job.kind` (not fake agent types):
| Job Kind | Icon | Label | Source |
|---|---|---|---|
| `cleanup` | 🧹 | Cleanup Archiver | `com.automaton.cleanup` |
| `loop` | 🔄 | Loop: {name} | `com.automaton.loop.{name}` |
| `rule-scan` | 📝 | Rule Proposer | `com.automaton.rule-scan` (new) |
| `rule-review` | 📋 | Rule Reviewer | `com.automaton.rule-review` (new) |
Each job card shows: job label, status (Enabled/Disabled/Misconfigured), next-run interval, runtime state (for loops: iteration count, halt status; for rule agents: last-scan/review timestamp). `AGENT_TYPE_META` is removed entirely; rendering uses `job.kind` directly.
### 7.3 Self-Documenting Names
Per `.rules.md` "Self-Documenting UI Names" (lines 53-64), all schedule unit names and stub filenames are self-documenting:
- `com.automaton.rule-scan` (launchd label)
- `automaton-rule-scan.sh` (stub filename)
- `com.automaton.rule-review` / `automaton-rule-review.sh`
## 8. Schedules
Rule agents use OS-native schedulers, mirroring the `--install-cleanup-schedule` pattern (`status.py:2188-2260`):
| Agent | Flag | Default Interval | Launchd Label | Stub |
|---|---|---|---|---|
| Rule Proposer | `--install-rule-scan-schedule` | 86400s (daily) | `com.automaton.rule-scan` | `automaton-rule-scan.sh` |
| Rule Reviewer | `--install-rule-review-schedule` | 2592000s (monthly) | `com.automaton.rule-review` | `automaton-rule-review.sh` |
Platform dispatch via `platform.system()`:
- **Darwin**: `~/Library/LaunchAgents/com.automaton.rule-scan.plist` with `StartInterval`.
- **Linux**: crontab line via `_install_cron_block_generic`.
- **Windows**: `schtasks /create /tn "AutomatonRuleScan" ...`.
Stub scripts are 3-line bash/bat files that call `python3 status.py --rule-scan` (or `--rule-review`).
## 9. Conflict-of-Interest
### 9.1 Dependency Chain
```
model-divergence-enforcement (shipped first, manual task)
↓ enables hard-block
rule-proposer-agent (FW-2)
↓ conflict-of-interest
rule-reviewer-agent (FW-3) — must differ from Proposer
```
### 9.2 Design Now, Enforce Later (F4, F5)
- Rule agents are *designed* with `model` fields in their schedule config.
- The design docs specify the conflict-of-interest rules.
- Hard-block enforcement *activates* when `models.json` exists and multi-LLM mode is detected.
- Until then, rule agents run with the default model (single-LLM mode, advisory only).
### 9.3 Why Different Models
- **Proposer vs Reviewer**: Proposer has a bias toward *adding* rules (more is better). Reviewer has a bias toward *pruning* (less is better). Same model = self-review = the proposer's rules never get pruned.
- **Rule agent vs scanned tasks**: The model that introduced a bug should not propose the rule about its own bug — it has a blind spot for that failure mode.
## 10. Success Criteria for v1
1. **Model-divergence**: `--transition --model` records the model in `.state.models`; `--claim` refuses conflict-matrix violations in multi-LLM mode; `--audit` flags violations; dashboard shows model badges.
2. **Rule Proposer**: `status.py --rule-scan` reads completed tasks since last scan, proposes rules to `RULE_PROPOSALS.md`, updates `.state.rule-scan`. Test with a seeded completed task containing a `VERDICT.md`.
3. **Rule Reviewer**: `status.py --rule-review` reads `.rules.md` + recent proposals, writes `RULE_REVIEW.md`, updates `.state.rule-review`. Test with a seeded `.rules.md` containing a contradiction.
4. **Agent Tab**: `/api/phase-roles` returns 6 roles with active-task counts; dashboard renders Phase Roles + Scheduled Jobs sections; `AGENT_TYPE_META` is removed; `job.kind` drives rendering.
5. **Schedules**: `--install-rule-scan-schedule` and `--install-rule-review-schedule` install OS-native units with self-documenting names.
6. **Tests**: `pytest tests/ -v` is green; new tests cover state schemas, scan flows, dashboard API, and conflict-matrix enforcement.
7. **No regression**: pre-existing test suite passes unchanged.
## 11. Locked Decision Index
All decisions referenced by `(Fn)` above are recorded in the v1 design conversation (this session). They are non-negotiable for v1 implementation. Changes require a design doc update and a new `[unreleased]` changelog entry.
See `README.md` § "Locked decisions" for the full table.
+536
View File
@@ -0,0 +1,536 @@
# Framework Agent Features — Technical Design
Companion to `functional.md`. This file is the implementation contract: every line here is what the implementation tasks build. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update.
## 1. File Map (what v1 adds)
```
~/.automaton/
├── models.json # NEW — model manifest (see §2)
├── scripts/
│ ├── detect_models.py # NEW — probes opencode.json + localhost endpoints
│ └── status.py # EXTENDED — new flags (see §4)
├── prompts/
│ ├── rule-proposer.md # NEW — Rule Proposer session prompt
│ ├── rule-reviewer.md # NEW — Rule Reviewer session prompt
│ └── onboarding.md # EXTENDED — Step 2e (backlog check)
├── automaton/
│ └── dashboard/
│ ├── html/dashboard.js # EXTENDED — remove AGENT_TYPE_META, two-section render
│ └── ui/app.py # EXTENDED — /api/phase-roles endpoint
├── .automaton/ # (framework self-hosting: this is ~/.automaton/.automaton/)
│ ├── .state.rule-scan # NEW — Rule Proposer state (see §3)
│ ├── .state.rule-review # NEW — Rule Reviewer state (see §3)
│ ├── RULE_PROPOSALS.md # NEW — Rule Proposer output (append-per-run)
│ ├── RULE_REVIEW.md # NEW — Rule Reviewer output (append-per-run)
│ └── automaton-rule-scan.sh # NEW — generated by --install-rule-scan-schedule
│ automaton-rule-review.sh # NEW — generated by --install-rule-review-schedule
└── tests/
├── test_model_divergence.py # NEW — manifest, conflict matrix, --transition --model
├── test_rule_agents.py # NEW — scan flows, state files, output schemas
└── test_dashboard_phase_roles.py # NEW — /api/phase-roles, two-section render
```
Per-project paths mirror the loop convention: `{project}/.automaton/.state.rule-scan`, `{project}/.automaton/RULE_PROPOSALS.md`, etc. For framework self-hosting, the project is `~/.automaton/` itself.
## 2. `models.json` Schema
```json
{
"schema_version": 1,
"default": "glm-4.6",
"advised": true,
"models": [
{
"name": "glm-4.6",
"provider": "opencode",
"context_window": 131072,
"location": "remote"
},
{
"name": "qwen3-coder",
"provider": "opencode",
"context_window": 131072,
"location": "remote"
},
{
"name": "llama-3.3-70b",
"provider": "localhost",
"context_window": 32768,
"location": "http://localhost:8080"
}
]
}
```
- `default`: model name used when no role-specific binding exists. Must be present in `models[]`.
- `advised`: bool. If `true`, single-LLM mode prints a one-time advisory recommending a second model, then goes silent.
- `models[]`: roster. `name` is the unique key. `provider` is informational. `context_window` is informational (framework never inspects capability, D8). `location` is `"remote"` or a localhost URL (for `detect_models.py` probing).
- **Missing file** → single-LLM mode (backward compatible). All model-divergence commands are no-ops.
- **0-1 models** → single-LLM mode. Advisory once if `advised: true`.
- **2+ models** → multi-LLM mode. Hard-block on conflict matrix.
### 2.1 `detect_models.py`
```
python3 scripts/detect_models.py [--json]
```
1. Parse `opencode.json` (or `opencode.jsonc`) for provider+model entries.
2. Probe localhost endpoints: `http://localhost:8080/v1/models`, `http://localhost:11434/api/tags` (Ollama), `http://localhost:1234/v1/models` (LM Studio), `http://localhost:8000/v1/models` (vLLM).
3. Merge results, emit a candidate `models.json` to stdout (or write if `--json` not set).
4. Used by `install.sh` / `update.sh` / `upgrade.sh` to bootstrap or refresh `models.json`.
## 3. State Schemas
### 3.1 `.state.rule-scan`
```json
{
"schema_version": 1,
"last_scan_at": "2026-06-25T10:00:00Z",
"last_scanned_task": "fix-context-sizing",
"proposals_count": 3,
"scanned_tasks_count": 12
}
```
- `last_scanned_task`: the most recent task name scanned. Next scan starts after this task (alphabetical or mtime order).
- `proposals_count`: cumulative count of proposals written to `RULE_PROPOSALS.md`.
- Stored at `{project}/.automaton/.state.rule-scan`. Missing file → first run scans all completed tasks.
### 3.2 `.state.rule-review`
```json
{
"schema_version": 1,
"last_review_at": "2026-06-25T10:00:00Z",
"contradictions_found": 2,
"stale_rules": 5,
"merges_suggested": 1
}
```
- Stored at `{project}/.automaton/.state.rule-review`. Missing file → first run reviews all rules.
### 3.3 `.state.models` (per-task)
```json
{
"schema_version": 1,
"implement": "glm-4.6",
"code_review": "qwen3-coder",
"bug_find": "qwen3-coder",
"adversarial_bug_find": "llama-3.3-70b",
"referee": "llama-3.3-70b",
"doc_review": null
}
```
- Stored at `{task}/.state.models`. One file per task.
- Written by `--transition --model <name>` when entering a phase.
- Read by `--claim` (conflict-matrix check) and `--audit` (violation detection).
- Roles not yet filled are `null` or absent.
## 4. `status.py` New Flags
All model-divergence and rule-agent commands route through `status.py` — no second enforcement surface.
```
# Model-divergence
status.py --transition <phase> --task <t> [--model <name>] Records model in .state.models; checks conflict matrix
status.py --claim --task <t> --agent <a> [--model <name>] Refuses if model conflicts with filled roles (multi-LLM mode)
status.py --audit EXTENDED — +model_divergence category
status.py --can-edit [...] UNCHANGED
# Rule agents
status.py --rule-scan [--project <p>] [--dry-run] Scan completed tasks, propose rules to RULE_PROPOSALS.md
status.py --install-rule-scan-schedule [--interval S] Install OS-native unit for --rule-scan (default daily)
status.py --rule-review [--project <p>] [--dry-run] Consolidate .rules.md, write RULE_REVIEW.md
status.py --install-rule-review-schedule [--interval S] Install OS-native unit for --rule-review (default monthly)
```
### 4.1 `--transition --model` flow
1. Load `models.json`. If missing or single-LLM mode → record model (advisory), no conflict check.
2. If multi-LLM mode: load `.state.models` for the task. Check the role being entered against the conflict matrix (§5).
3. If `--model` not provided: auto-assign next-available non-conflicting model from `models[]`. Refuse if none available.
4. If `--model` provided: verify it's in `models[]`. Check conflict matrix. Refuse on violation.
5. Write `role: model` to `.state.models`. Transition the phase.
### 4.2 `--claim --model` flow
1. Load `models.json`. If single-LLM mode → existing claim logic, no model check.
2. If multi-LLM mode: load `.state.models`. Determine the role for the phase being claimed. Check conflict matrix against already-filled roles.
3. Refuse if the claiming agent's model conflicts. Error message names the conflicting role and model.
### 4.3 `--audit` extension
New audit category `model_divergence`:
- For each task with `.state.models`: check all filled roles against the conflict matrix.
- Flag violations as `severity: high` (conflict-of-interest is a correctness issue, not a style issue).
- Output format mirrors existing audit categories.
## 5. Conflict Matrix (implementation)
```python
CONFLICT_MATRIX = {
"code_review": {"implement"},
"bug_find": {"implement"},
"adversarial_bug_find": {"implement", "bug_find"},
"referee": {"implement", "bug_find", "adversarial_bug_find"},
"loop-verify": {"loop-implement"},
}
```
- Key = role being entered. Value = set of roles that must have a different model.
- `doc_review`, `code_review`, `bug_find` are NOT in conflict with each other (only `bug_find` ↔ `adversarial_bug_find` conflicts).
- Check function: `def _check_conflict(state_models: dict, role: str, model: str, matrix: dict) -> Optional[str]` — returns the conflicting role name or `None`.
## 6. Rule Proposer Flow (`--rule-scan`)
```
1. Load .state.rule-scan (or init if missing).
2. Find completed tasks since last_scanned_task:
- Scan tasks/complete/ and tasks with .state phase=complete
- Filter by mtime > last_scan_at (or all if first run)
- Sort by mtime ascending
3. For each task:
a. Read BUG_REPORT.md, ADVERSARIAL_BUG_REPORT.md, VERDICT.md (skip if none exist)
b. Read current .rules.md (for dedup context — capped at 4k tokens)
c. Build proposer prompt (see §7)
d. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_proposer_model)
e. Parse LLM output for proposed rules (expect RULE_PROPOSALS.md format)
f. Append proposals to RULE_PROPOSALS.md
g. Update .state.rule-scan (last_scanned_task, proposals_count)
4. Write final .state.rule-scan with last_scan_at = now.
```
- `--dry-run`: list tasks that would be scanned, do not invoke harness.
- `--project`: scope to a project (default: framework dir).
- Errors during a single task scan do not abort the run; the scan continues to the next task and logs the error.
## 7. Rule Proposer Prompt Shape (`rule-proposer.md`)
```
# Rule Proposer — {date}
You are scanning completed tasks for failure patterns that should become rules.
## Current rules (read-only, for dedup)
{current_rules} # .rules.md content, capped at 4k tokens
## Task failure artifacts
{bug_report} # BUG_REPORT.md content, capped at 2k tokens
{adversarial_report} # ADVERSARIAL_BUG_REPORT.md, capped at 2k tokens
{verdict} # VERDICT.md, capped at 2k tokens
## What to do
For each distinct failure pattern you observe:
1. Check if a rule already exists in .rules.md that covers it. If so, skip.
2. If no existing rule covers it, propose a new rule with:
- A concrete example from the task artifacts
- The proposed rule text as it would appear in .rules.md
## Output (strict markdown, no JSON)
## Proposed Rule: {title}
**Source**: tasks/{task-name}/VERDICT.md
**Pattern**: {one-line description}
**Example**:
{concrete snippet}
**Proposed rule text**:
{rule text}
---
```
No `{model}` token in the prompt — the model is selected by the caller and passed to `_invoke_harness`.
## 8. Rule Reviewer Flow (`--rule-review`)
```
1. Load .state.rule-review (or init if missing).
2. Read .rules.md (full file).
3. Read recent RULE_PROPOSALS.md entries (since last_review_at).
4. Read recent completed-task summaries (last 30 days) for staleness context.
5. Build reviewer prompt (see §9).
6. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_reviewer_model).
7. Parse LLM output for review sections (contradictions, stale, missing examples, merges).
8. Append to RULE_REVIEW.md.
9. Update .state.rule-review.
```
- `--dry-run`: report what would be reviewed, do not invoke harness.
## 9. Rule Reviewer Prompt Shape (`rule-reviewer.md`)
```
# Rule Reviewer — {date}
You are consolidating .rules.md for contradictions, staleness, and missing examples.
## Current rules (full)
{rules_content} # .rules.md, full file
## Recent proposals (since last review)
{recent_proposals} # RULE_PROPOSALS.md entries since last_review_at
## Recent completed tasks (last 30 days, for staleness context)
{task_summaries} # one-line per task: name + phase + completion date
## What to check
1. Contradictions: rules that conflict with each other.
2. Stale rules: no observed instance in last 30 days.
3. Rules missing examples: any rule without a concrete example.
4. Merge candidates: overlapping rules that could be consolidated.
## Output (strict markdown, no JSON)
## Contradictions Found
- ...
## Stale Rules
- ...
## Rules Missing Examples
- ...
## Merge Candidates
- ...
```
## 10. Harness Invocation (direct, not loop)
Rule agents reuse `loop-runner._invoke_harness` directly — they are NOT loops. The function signature (from `loop-runner.py:366-404`):
```python
def _invoke_harness(harness_command: str, prompt_content: str, cwd: str, env: dict = None) -> str:
```
`status.py --rule-scan` calls this as:
```python
from loop_runner import _invoke_harness
output = _invoke_harness(
harness_command=rule_harness_command, # from schedule config or default
prompt_content=resolved_prompt, # rule-proposer.md with tokens substituted
cwd=str(project_dir),
env={"AUTOMATON_RULE_ROLE": "proposer"}
)
```
`{model}` substitution: if the harness command contains `{model}`, it's replaced with the rule agent's configured model. Until model-divergence ships, this is the default model.
### 10.1 Schedule Config for Rule Agents
Rule agents do not use `loop.json`. Their config is embedded in the schedule stub:
```bash
#!/usr/bin/env bash
cd "<project_root>"
python3 "<framework>/scripts/status.py" --rule-scan --model <name>
```
The `--model` flag is optional and ignored in single-LLM mode. In multi-LLM mode it sets the rule agent's model (subject to conflict-of-interest checks once enforced).
## 11. Scheduler Unit Generation
Mirrors `cmd_install_cleanup_schedule` (`status.py:2188-2260`) exactly:
### 11.1 `--install-rule-scan-schedule`
```python
def cmd_install_rule_scan_schedule(args) -> int:
interval = args.interval if args.interval else 86400 # daily
# 1. Write stub: automaton-rule-scan.sh
# 2. Platform dispatch:
# Darwin → ~/Library/LaunchAgents/com.automaton.rule-scan.plist
# Linux → crontab block via _install_cron_block_generic
# Windows → schtasks /create /tn "AutomatonRuleScan"
```
### 11.2 `--install-rule-review-schedule`
```python
def cmd_install_rule_review_schedule(args) -> int:
interval = args.interval if args.interval else 2592000 # monthly
# Same pattern, labels: com.automaton.rule-review / AutomatonRuleReview
```
### 11.3 `_list_scheduled_jobs` extension
`_list_scheduled_jobs` (`status.py:2282`) gains recognition for new labels:
```python
def _launchd_label_kind(label: str) -> tuple[str, str]:
if label.startswith("com.automaton.loop."):
return ("loop", label[len("com.automaton.loop."):])
if label in ("com.automaton.cleanup",):
return ("cleanup", "")
if label in ("com.automaton.rule-scan",):
return ("rule-scan", "")
if label in ("com.automaton.rule-review",):
return ("rule-review", "")
...
```
This makes rule-scan and rule-review jobs appear in `/api/scheduled` with their real `kind`, which the Agent tab renders directly.
## 12. Agent Tab Data Flow
### 12.1 New endpoint: `/api/phase-roles`
`app.py` gains a handler:
```python
elif self.path == "/api/phase-roles":
self._serve_phase_roles()
```
```python
def _serve_phase_roles(self):
# 1. Parse .agent.md Agent Configuration for role definitions
# 2. Load all tasks via status.py module
# 3. For each role, count tasks in that role's phases
# 4. Return JSON:
{
"roles": [
{"id": "researcher", "label": "Researcher", "icon": "🔬",
"phases": ["research", "decomposition", "design", "test_design"],
"active_tasks": 2, "status": "active"},
...
],
"available": True
}
```
Role icons (self-documenting, per `.rules.md` Self-Documenting UI Names):
| Role | Icon |
|---|---|
| researcher | 🔬 |
| implementer | ⚙️ |
| code-reviewer | 👁️ |
| bug-hunter | 🐛 |
| referee | ⚖️ |
| orchestrator | 🎯 |
### 12.2 `dashboard.js` changes
**Remove**: `AGENT_TYPE_META` (lines 330-334), `AGENT_TYPES` (337), `AGENT_TYPE_META_FALLBACK` (338), `_resolveAgentType` (351-354).
**Replace `renderAgentTab`** with a two-section render:
```javascript
async function renderAgentTab() {
const panel = document.getElementById('agent-panel');
panel.innerHTML = '<div class="bg-loading">Loading…</div>';
const [rolesRes, schedRes] = await Promise.all([
fetch('/api/phase-roles').then(r => r.json()).catch(() => ({roles: [], available: false})),
fetchSchedule(),
]);
// Section 1: Phase Roles
const rolesHtml = rolesRes.available ? renderPhaseRoles(rolesRes.roles)
: '<div class="bg-empty">Phase roles require .agent.md Agent Configuration.</div>';
// Section 2: Scheduled Jobs
const jobsHtml = renderScheduledJobs(schedRes.jobs || []);
panel.innerHTML = `
<div class="agent-section">
<h3>Phase Roles</h3>
<div class="bg-grid">${rolesHtml}</div>
</div>
<div class="agent-section">
<h3>Scheduled Jobs</h3>
<div class="bg-grid">${jobsHtml}</div>
</div>`;
}
```
**`renderScheduledJobs`** uses `job.kind` directly (no fake type resolution):
```javascript
const JOB_META = {
cleanup: { icon: '🧹', label: 'Cleanup Archiver' },
loop: { icon: '🔄', label: (j) => `Loop: ${j.name}` },
'rule-scan': { icon: '📝', label: 'Rule Proposer' },
'rule-review':{ icon: '📋', label: 'Rule Reviewer' },
};
```
## 13. Loop Integration (model-divergence)
### 13.1 `loop.json` per-role model
```json
"roles": {
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
}
```
- `model` is optional. If absent, uses `models.json` `default`.
- `loop-verify` model is checked against `loop-implement` model in `--check-gate` (multi-LLM mode).
### 13.2 `{model}` substitution in `_invoke_harness`
`loop-runner.py:366-404` `_invoke_harness` gains `{model}` token substitution:
```python
def _invoke_harness(harness_command, prompt_content, cwd, env=None, model=None):
if model and "{model}" in harness_command:
harness_command = harness_command.replace("{model}", model)
...
```
The caller passes `model` from the role config. If the harness command has no `{model}` token, the model is informational only (the harness picks its own).
### 13.3 `--check-gate` model-divergence check
In multi-LLM mode, `--check-gate` adds:
- Load `loop.json` roles. Compare `verify.model` vs `implement.model`.
- If same model and multi-LLM mode → halt as `model_conflict` (new halt reason, or reuse `human_intervention` with a descriptive message).
## 14. Test Coverage
### 14.1 `test_model_divergence.py`
- `test_models_json_missing_single_llm_mode` — no file → advisory, no blocks.
- `test_single_model_advisory_once` — 1 model, `advised: true` → advisory printed once, then silent.
- `test_multi_llm_conflict_matrix` — 2+ models, `--transition --model` records, `--claim` refuses conflict.
- `test_auto_assign_next_available` — no `--model` flag → auto-assigns non-conflicting model.
- `test_auto_assign_exhausted` — all models conflict → refuse.
- `test_audit_model_divergence` — `--audit` flags conflict-matrix violations.
- `test_loop_verify_neq_implement` — `--check-gate` halts on same model in multi-LLM mode.
### 14.2 `test_rule_agents.py`
- `test_rule_scan_finds_completed_tasks` — seeded completed task with VERDICT.md → proposal written.
- `test_rule_scan_state_tracking` — `.state.rule-scan` updated with last_scanned_task + count.
- `test_rule_scan_dedup` — existing rule in `.rules.md` → not re-proposed.
- `test_rule_scan_dry_run` — no harness invocation, lists candidates.
- `test_rule_review_finds_contradictions` — seeded `.rules.md` with contradiction → review written.
- `test_rule_review_state_tracking` — `.state.rule-review` updated.
- `test_install_rule_scan_schedule` — stub + plist created with correct labels.
- `test_install_rule_review_schedule` — stub + plist created with correct labels.
### 14.3 `test_dashboard_phase_roles.py`
- `test_api_phase_roles` — `/api/phase-roles` returns 6 roles with correct phases.
- `test_phase_roles_active_count` — tasks in phases → correct active_tasks count.
- `test_scheduled_jobs_new_kinds` — rule-scan and rule-review jobs appear with correct `kind`.
- `test_agent_type_meta_removed` — `AGENT_TYPE_META` no longer in dashboard.js (grep test).
## 15. Rollout (3 sequential tasks for model-divergence)
The model-divergence-enforcement parent task decomposes into 3 subtasks:
1. **manifest+detection**: `models.json` schema, `detect_models.py`, `install.sh`/`update.sh`/`upgrade.sh` integration, `config.md` section, onboarding Step 2d.
2. **interactive enforcement+audit**: `.state.models`, `--transition --model`, `--claim --model` conflict check, `--audit` model_divergence category, dashboard badges.
3. **loop enforcement+dashboard**: `loop.json` per-role model, `{model}` substitution, `--check-gate` model check, loop dashboard badges.
Rule agents (FW-2, FW-3) and Agent tab (FW-1) are backlog items, picked up after model-divergence ships (for FW-2/FW-3) or independently (for FW-1).
## 16. Locked Decision Index
All decisions referenced by `(Fn)` are in `README.md` § "Locked decisions". Implementation must conform. Deviations require a design doc update + `[unreleased]` CHANGELOG entry.
+2
View File
@@ -6,6 +6,8 @@ Work queue for the self-improvement loop after task 7 lands. Items not assigned
A loop configured with `work_source.kind = "backlog"` reads this file, picks the topmost `[ ]` item, drafts an implementation, transitions through phases, hands off to a human reviewer. Mark items `[x]` when complete; move items to `DONE.md` (created later) on closure.
**Sibling backlog**: `design/framework/BACKLOG.md` covers framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). A loop with `work_source.area = "framework"` reads that file instead of this one.
## v1.1 — framework manages its own docs
- [ ] **design-update-loop-template** — `templates/loops/design-update/` that keeps `design/<area>/*.md` in sync with the code it documents. Triggered by `last-read-sha` drift detection.
+1
View File
@@ -15,6 +15,7 @@ Per-session manual driving doesn't scale against the framework's growing backlog
- [`functional.md`](functional.md) — what v1 does, roles, the five deaths, blast radius, schedules, success criteria. **Read this first.**
- [`technical.md`](technical.md) — the implementation contract: file map, `.state.loop` schema, `status.py` flags, gate checks, runner flow, test coverage. **Read this if you're implementing.**
- [`BACKLOG.md`](BACKLOG.md) — v1.1 and deferred items (Scope 2 design-update loop, Scope 3 self-designing, parallel mode, dashboard panel). The self-improvement loop's work queue.
- **Sibling design**: [`../framework/`](../framework/) — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). Has its own `BACKLOG.md` consumable via `work_source.area = "framework"`.
## v1 scope (locked)
+5
View File
@@ -103,6 +103,7 @@ Gate checks, in order:
4. **Task phase** — if `current_task` is set, that task's `.state` must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to `human_intervention` by some other path), halt as `human_intervention`.
5. **Worktree drift** — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via `git diff --name-only main...HEAD` restricted to `file_scope`), halt as `drift_detected`.
6. **Score plateau** — last N entries in `score_history` are flat or monotonically decreasing (where N = `score_plateau_window`). Trip → halt as `verifier_failed`.
7. **Model divergence** — in multi-LLM mode (2+ models in `models.json`), checks that the loop's implement and verify roles use different models. If they share the same model, halt as `human_intervention` (this prevents same-model verification / rubber-stamping within a loop tick). Single-LLM mode is exempt. Model is resolved from `roles[<role>].model` if set, otherwise the manifest default.
All halts atomically set `status=halted`, `halt_reason=<reason>`, write to `.state.log`, and call `--pause-loop`'s schedule-disable step (see §6).
@@ -258,11 +259,13 @@ To bound `outputs/` directory growth (O5 from `add-loop-runner/BUG_REPORT.md`),
v1.1's default `harness.command` is `opencode run` -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the `command` list (single argv element per token, no shell expansion):
- `{model}` -- the model assigned to the role (from `roles[<role>].model` or manifest default). Passed via `extras["model"]`. If the role has no model assignment, the token is left unsubstituted.
- `{prompt}` -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path.
- `{prompt_content}` -- the resolved prompt file's text content as a single argv element. Safe under `subprocess.run` list mode; no shell quoting needed. Used by the default command since `opencode run` takes the message as a positional argument and has no `--prompt-file` flag.
- `{cwd}` -- the working directory the harness should run in (the loop's project root or worktree).
- `{output}`, `{artifact}` -- role-specific extras (the implement output path handed to verify).
- `{verdict}`, `{current_task}`, `{current_phase}`, etc. -- other runtime extras; see `_resolve_prompt` below.
- `{model}` -- the model assigned to the role being invoked (from `roles[<role>].model` in `loop.json`, or the manifest default). The runner passes it via the `extras["model"]` key. If the role has no explicit model, `{model}` is left as-is (no substitution). This allows per-role model pinning without hardcoding the model name in `harness.command`.
The default command does NOT hardcode a `--model` flag; the spawned `opencode run` inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set `harness.command` in their `loop.json`:
@@ -336,6 +339,8 @@ This means the harness receives a fully-resolved prompt file with all context ba
}
```
Each role in `roles` accepts an optional `"model"` field to pin a specific model for that role (e.g. `"implement": {"prompt": "loop-implement.md", "tier": 16000, "model": "model-a"}`). When set, the runner passes `model=<value>` in the harness extras for that role, enabling `{model}` substitution in `harness.command`. This is how multi-LLM loops prevent same-model verification — see `CONFLICT_MATRIX` in `status.py`.
Installs default-on at `install.sh` time: `status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` then `status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"`. Both commands use `|| true` so the framework works even if loop creation fails. `update.sh` bootstraps the loop idempotently for existing users (checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`). Disabling: `status.py --pause-loop self-improvement --project ~/.automaton/`.
## 10. Tests (`tests/test_loops.py`)
+14
View File
@@ -0,0 +1,14 @@
{
"current_task": null,
"halt_reason": null,
"iteration_count": 0,
"last_tick_at": null,
"last_verdict": null,
"name": "self-improvement",
"resumed_count": 0,
"schema_version": 1,
"score_history": [],
"status": "running",
"worktree_branch": null,
"worktree_path": null
}
+3
View File
@@ -0,0 +1,3 @@
#!/usr/bin/env bash
cd "/Users/laptran/.automaton"
python3 "/Users/laptran/.automaton/scripts/loop-runner.py" --mode tick --loop "self-improvement"
+38
View File
@@ -0,0 +1,38 @@
{
"name": "self-improvement",
"description": "Ticks against status.py --audit on the framework's own repo",
"schedule": {
"interval_seconds": 3600
},
"brakes": {
"max_iterations": 10,
"max_budget_usd": null,
"score_plateau_window": 3
},
"blast_radius": {
"file_scope": ["scripts/", "prompts/", "tests/", "design/"],
"base_branch": "main",
"use_worktree": true
},
"work_source": {
"kind": "audit",
"project": "~/.automaton/"
},
"acceptance_criteria": [
"Audit findings resolved (no outstanding Cat-1/Cat-2/Cat-4 violations on the resolved task)",
"All R-numbers from the task SPEC.md implemented",
"Tests pass with no regressions",
"Pipeline driven to complete"
],
"roles": {
"implement": {"prompt": "loop-implement.md"},
"verify": {"prompt": "loop-verifier.md"},
"orchestrate": {"prompt": "loop-orchestrate.md"}
},
"outputs": {
"retention": 20
},
"harness": {
"command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]
}
}
+12 -1
View File
@@ -53,4 +53,15 @@ Picked from BUG_REPORTs of the loop tasks and from `design/loops/BACKLOG.md` def
## NEXT
Resume task 2 (`add-state-loop-lock`): SPEC is already written, state is `research:awaiting_approval`. Approve it, transition to implement, write the `_loop_lock` helper + wrap callsites in `status.py` and `loop-runner.py`, add `tests/test_state_loop_lock.py` (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order.
Resume task 2 (`add-state-loop-lock`): SPEC is already written, state is `research:awaiting_approval`. Approve it, transition to implement, write the `_loop_lock` helper + wrap callsites in `status.py` and `loop-runner.py`, add `tests/test_state_loop_lock.py` (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order.
## Session 2026-06-25 — Framework agent features design
- Created `design/framework/` with `README.md`, `functional.md`, `technical.md`, `BACKLOG.md` — design for three framework-level agent features: model-divergence enforcement, rule agents (Proposer + Reviewer), Agent tab redesign.
- **Model-divergence enforcement** is a manual task (`model-divergence-enforcement`), decomposed into 3 subtasks: (1) manifest+detection, (2) interactive enforcement+audit, (3) loop enforcement+dashboard. Not a backlog item.
- **Rule agents** (FW-2, FW-3) are backlog items depending on model-divergence shipping first (conflict-of-interest LLM binding). Rule Proposer runs daily, Rule Reviewer runs monthly. Both use direct harness invocation (reuse `loop-runner._invoke_harness`), not loop infrastructure.
- **Agent tab redesign** (FW-1) is a backlog item with no dependencies. Replaces 4 fake `AGENT_TYPE_META` types with Phase Roles (6 roles from `.agent.md`) + Scheduled Jobs (real `job.kind`).
- Decision: rule agents use **direct harness invocation** (not loops, not standalone status.py commands).
- Decision: model-divergence conflict-of-interest is **designed now, enforced later** — rule agents carry `model` fields in config but hard-blocking activates only when `models.json` exists and multi-LLM mode is detected.
- Cross-references updated: `AGENTS.md` repo layout, `README.md` loop config table, `design/loops/README.md`, `design/loops/BACKLOG.md`, `CHANGELOG.md`, `.onboarding.md` (Backlog section), `prompts/onboarding.md` (Step 2d).
- Next: switch on self-improvement loop (`--create-loop self-improvement --from-template self-improvement` + `--install-schedule`), then create `model-divergence-enforcement` parent task and decompose into 3 subtasks.
+141
View File
@@ -0,0 +1,141 @@
import type { ExtensionAPI, ExtensionContext, BeforeAgentStartEventResult } from "@earendil-works/pi-coding-agent";
import { existsSync, readFileSync, readdirSync, statSync } from "fs";
import { basename, join } from "path";
import { homedir } from "os";
const AUTOMATON_HOME = join(homedir(), ".automaton");
const STALE_MINUTES = 30;
interface TaskInfo {
name: string;
phase: string;
mtime: Date;
}
function getTasks(autoDir: string): TaskInfo[] {
const tasksDir = join(autoDir, ".automaton", "tasks");
if (!existsSync(tasksDir)) return [];
try {
const entries = readdirSync(tasksDir);
const tasks: TaskInfo[] = [];
for (const entry of entries) {
const stateFile = join(tasksDir, entry, ".state");
if (existsSync(stateFile)) {
try {
const phase = readFileSync(stateFile, "utf-8").trim();
if (!phase) continue;
const stat = statSync(stateFile);
tasks.push({ name: entry, phase, mtime: stat.mtime });
} catch {
// skip unreadable
}
}
}
tasks.sort((a, b) => b.mtime.getTime() - a.mtime.getTime());
return tasks;
} catch {
return [];
}
}
function getLoopInfo(autoDir: string): string[] {
const loopsDir = join(autoDir, ".automaton", "loops");
if (!existsSync(loopsDir)) return [];
try {
const entries = readdirSync(loopsDir);
const lines: string[] = [];
for (const entry of entries) {
const stateFile = join(loopsDir, entry, ".state.loop");
if (existsSync(stateFile)) {
try {
const content = readFileSync(stateFile, "utf-8").trim();
const state = JSON.parse(content);
const taskRef = state.current_task ? `, active task: ${state.current_task}` : "";
lines.push(`Loop "${entry}": ${state.status || "unknown"}${taskRef}`);
} catch {
// skip unparseable
}
}
}
return lines;
} catch {
return [];
}
}
export default function (pi: ExtensionAPI) {
pi.on("before_agent_start", async (_event, ctx): Promise<BeforeAgentStartEventResult | undefined> => {
const cwd = ctx.cwd;
const inFramework = cwd === AUTOMATON_HOME || cwd.startsWith(AUTOMATON_HOME + "/");
let baseDir: string | null = null;
let scopeLabel: string;
if (inFramework) {
baseDir = AUTOMATON_HOME;
scopeLabel = "framework (~/.automaton/)";
} else {
const projectAuto = join(cwd, ".automaton");
if (existsSync(projectAuto)) {
baseDir = cwd;
const projectName = basename(cwd) || "project";
scopeLabel = `project (${projectName}/.automaton/)`;
} else {
return; // not in automaton context
}
}
const tasks = getTasks(baseDir);
const editTasks = tasks.filter((t) => t.phase === "implement" || t.phase === "doc_review");
const currentTask = editTasks.length > 0 ? editTasks[0] : tasks.length > 0 ? tasks[0] : null;
const lines: string[] = [];
lines.push(`Automaton scope: ${scopeLabel}`);
lines.push("");
lines.push("How Automaton works:");
lines.push("- Automaton enforces a task phase state machine. File edits are ONLY allowed when a task is in 'implement' or 'doc_review' phase.");
lines.push("- When the user asks you to build, fix, or change something, you MUST first create a task (use automaton_create_task) and transition it to 'implement' before you can edit any files.");
lines.push("- Without a task in 'implement', the automaton-guard-pi extension will BLOCK all file edits.");
lines.push(`- Available tools: automaton_create_task (create+transition), automaton_transition (change phase), automaton_status (check state).`);
lines.push("");
if (currentTask) {
lines.push(`State: Active task "${currentTask.name}" is in "${currentTask.phase}" phase.`);
if (currentTask.phase === "implement" || currentTask.phase === "doc_review") {
lines.push(`You CAN edit files under this task.`);
const ageMinutes = Math.round((Date.now() - currentTask.mtime.getTime()) / 60000);
if (ageMinutes > STALE_MINUTES) {
lines.push(
`Note: task has been in ${currentTask.phase} for ${ageMinutes} min and is considered stale. ` +
`Run automaton_transition (or touch via status.py) if still active.`,
);
}
} else {
lines.push(`File edits are BLOCKED. Call automaton_transition to move it to 'implement' before editing.`);
}
} else {
lines.push(`State: No active tasks. When the user makes a work request, call automaton_create_task first.`);
}
const byPhase = new Map<string, number>();
for (const t of tasks) {
byPhase.set(t.phase, (byPhase.get(t.phase) || 0) + 1);
}
const summary = Array.from(byPhase.entries())
.sort((a, b) => b[1] - a[1])
.map(([p, c]) => `${p} (${c})`)
.join(", ");
lines.push(`All tasks: ${summary || "none"}`);
const loopInfo = getLoopInfo(baseDir);
for (const l of loopInfo) {
lines.push(l);
}
return {
systemPrompt: `${_event.systemPrompt}\n\n## Automaton Context\n\n${lines.join("\n")}`,
};
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-automaton-context",
"version": "1.0.0",
"description": "Auto-injects Automaton framework context into Pi Dev system prompt — no more manual copy-paste of pi-automaton.sh output",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+171
View File
@@ -0,0 +1,171 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "@sinclair/typebox";
import { execSync } from "child_process";
import { existsSync, readFileSync, readdirSync } from "fs";
import { join } from "path";
import { homedir } from "os";
const STATUS_SCRIPT = join(homedir(), ".automaton", "scripts", "status.py");
function resolveProjectDir(cwd: string): string | null {
const inFramework = cwd === join(homedir(), ".automaton") || cwd.startsWith(join(homedir(), ".automaton") + "/");
if (inFramework) return homedir() + "/.automaton";
if (existsSync(join(cwd, ".automaton"))) return cwd;
return null;
}
function getTasks(projectDir: string) {
const tasksDir = join(projectDir, ".automaton", "tasks");
if (!existsSync(tasksDir)) return [];
const tasks: { name: string; phase: string }[] = [];
for (const entry of readdirSync(tasksDir)) {
const stateFile = join(tasksDir, entry, ".state");
if (existsSync(stateFile)) {
try {
const phase = readFileSync(stateFile, "utf-8").trim();
if (phase) tasks.push({ name: entry, phase });
} catch {}
}
}
tasks.sort((a, b) => a.name.localeCompare(b.name));
return tasks;
}
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "automaton_create_task",
label: "Create Automaton Task",
description:
"Create a new automaton task and transition it to the specified phase. " +
"Use this when the user asks you to do work that requires file edits — you need a task in 'implement' or 'doc_review' phase before you can edit files.",
promptSnippet: "Create automaton tasks for work management",
promptGuidelines: [
"When the user asks you to build, implement, fix, or change something, first call automaton_create_task to create a task and transition it to 'implement' phase",
"Only after the task is in 'implement' can you edit files — the guard will block edits otherwise",
"Use a descriptive task name based on what the user wants (e.g., 'add-login-page', 'fix-api-timeout')",
"Default phase is 'implement' — omit phase for most cases",
],
parameters: Type.Object({
name: Type.String({ minLength: 1, description: "Task name (kebab-case, e.g. 'add-login-page')" }),
phase: Type.Optional(Type.String({ default: "implement", description: "Phase to transition to after creation" })),
}),
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { name, phase = "implement" } = params;
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project. No .automaton/ directory found." }],
details: {},
};
}
try {
const createOut = execSync(
`python3 ${STATUS_SCRIPT} --create-task ${JSON.stringify(name)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
const transitionOut = execSync(
`python3 ${STATUS_SCRIPT} --transition ${phase} --task ${JSON.stringify(name)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
return {
content: [{ type: "text", text: `Created task "${name}" and transitioned to "${phase}".` }],
details: { task: name, phase },
};
} catch (e: any) {
return {
content: [{
type: "text",
text: `Failed to create task: ${e.stderr || e.message || String(e)}`,
}],
details: { error: e.stderr || e.message, task: name },
};
}
},
});
pi.registerTool({
name: "automaton_transition",
label: "Transition Automaton Task",
description:
"Transition an existing automaton task to a new phase. " +
"Valid phases: research, decomposition, design, implement, test_design, testing, doc_review, complete. " +
"Use this to move a task forward (e.g. from research to implement).",
promptSnippet: "Transition automaton tasks between phases",
parameters: Type.Object({
task: Type.String({ minLength: 1, description: "Task name" }),
phase: Type.String({ minLength: 1, description: "Target phase" }),
}),
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { task, phase } = params;
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project." }],
details: {},
};
}
try {
execSync(
`python3 ${STATUS_SCRIPT} --transition ${phase} --task ${JSON.stringify(task)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
return {
content: [{ type: "text", text: `Task "${task}" transitioned to "${phase}".` }],
details: { task, phase },
};
} catch (e: any) {
return {
content: [{
type: "text",
text: `Failed to transition task: ${e.stderr || e.message || String(e)}`,
}],
details: { error: e.stderr || e.message, task, phase },
};
}
},
});
pi.registerTool({
name: "automaton_status",
label: "Automaton Status",
description:
"Show current automaton project status — all tasks, their phases, and any running loops. " +
"Call this to check what state things are in before deciding what to do.",
promptSnippet: "Check automaton project status",
parameters: Type.Object({}),
async execute(_toolCallId, _params, _signal, _onUpdate, _ctx) {
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project." }],
details: {},
};
}
const tasks = getTasks(projectDir);
const editTasks = tasks.filter((t) => t.phase === "implement" || t.phase === "doc_review");
const currentTask = editTasks.length > 0 ? editTasks[0] : tasks.length > 0 ? tasks[0] : null;
const lines: string[] = [];
lines.push(`Project: ${projectDir.split("/").pop()}`);
if (currentTask) {
lines.push(`Current: "${currentTask.name}" (${currentTask.phase})`);
} else {
lines.push("No tasks. Create one with automaton_create_task.");
}
lines.push("");
for (const t of tasks) {
const marker = currentTask && t.name === currentTask.name ? ">" : " ";
lines.push(`${marker} ${t.name.padEnd(35)} ${t.phase}`);
}
return {
content: [{ type: "text", text: lines.join("\n") }],
details: { tasks: tasks.length, current: currentTask?.name || null },
};
},
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-automaton-tools",
"version": "1.0.0",
"description": "Registers automaton_create_task and automaton_transition tools so the agent can manage task lifecycle automatically",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+91
View File
@@ -0,0 +1,91 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "@sinclair/typebox";
const BIRD_BIN = "/opt/homebrew/bin/bird";
const TweetParams = Type.Object({
url: Type.String({ minLength: 1, description: "Tweet URL or numeric ID" }),
mode: Type.Optional(
Type.Union([
Type.Literal("read"),
Type.Literal("thread"),
Type.Literal("replies"),
]),
),
});
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "read_tweet",
label: "Read Tweet",
description:
"Read an X/Twitter tweet or thread using the local bird CLI. " +
"Returns tweet author, text, media, metrics, and timestamps. " +
"Use this instead of webfetch for x.com/twitter.com URLs.",
promptSnippet: "Read X/Twitter tweets with bird CLI",
promptGuidelines: [
"Use read_tweet when the user shares an x.com or twitter.com URL and asks what it says",
"Use mode='thread' for full conversation threads",
"Use mode='replies' to fetch replies to a tweet",
"If fetch fails with auth errors, ask the user to sign in to x.com in their browser and retry",
],
parameters: TweetParams,
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { url, mode = "read" } = params;
try {
const { execSync } = await import("child_process");
const cmd = `${BIRD_BIN} ${mode} ${JSON.stringify(url)} --json`;
const stdout = execSync(cmd, { encoding: "utf-8", timeout: 15000 });
const parsed = JSON.parse(stdout);
const author = parsed.author?.name || parsed.author?.screen_name || "unknown";
const text = parsed.text || parsed.content || "";
const createdAt = parsed.created_at || "";
const retweetCount = parsed.metrics?.retweet_count ?? parsed.metrics?.retweets ?? 0;
const likeCount = parsed.metrics?.like_count ?? parsed.metrics?.likes ?? 0;
const replyCount = parsed.metrics?.reply_count ?? parsed.metrics?.replies ?? 0;
const mediaCount = parsed.media?.length ?? 0;
const threadCount = parsed.thread?.tweets?.length ?? 0;
const repliesCount = parsed.replies?.length ?? 0;
const lines: string[] = [];
lines.push(`Author: ${author}`);
if (createdAt) lines.push(`Posted: ${createdAt}`);
lines.push("");
lines.push(text);
if (retweetCount || likeCount || replyCount) {
lines.push("");
lines.push(`Retweets: ${retweetCount} Likes: ${likeCount} Replies: ${replyCount}`);
}
if (mediaCount) lines.push(`Media: ${mediaCount} attachment(s)`);
if (threadCount) lines.push(`Thread: ${threadCount} tweets total`);
if (repliesCount) lines.push(`Replies fetched: ${repliesCount}`);
return {
content: [
{ type: "text", text: lines.join("\n") },
{ type: "text", text: `\n--- raw ---\n${JSON.stringify(parsed, null, 2)}` },
],
details: {
author,
text: text.slice(0, 500),
url,
mode,
tweetCount: threadCount || repliesCount || 1,
},
};
} catch (e: any) {
const errMsg = e.stderr || e.message || String(e);
const hint = errMsg.includes("auth")
? " Sign in to x.com in your browser and retry."
: "";
return {
content: [{ type: "text", text: `Failed to fetch tweet: ${errMsg}${hint}` }],
details: { error: errMsg, url, mode },
};
}
},
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-read-tweet",
"version": "1.0.0",
"description": "Read X/Twitter tweets using the local `bird` CLI — registered as a read_tweet tool for Pi Dev",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+7
View File
@@ -75,6 +75,13 @@ Install the automaton pre-commit hook to block commits when no task is in an edi
4. Verify the hook: `python ~/.automaton/scripts/status.py --can-edit --project {project}` should return exit code 1 (DENIED) since no tasks exist yet.
5. Note in the onboarding report whether the hook was installed.
### Step 2d: Backlog Check
Check for outstanding design work that may need attention:
1. Read `~/.automaton/design/loops/BACKLOG.md` for loop engineering work queue items.
2. Read `~/.automaton/design/framework/BACKLOG.md` for framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement).
3. Note in the onboarding report whether there are unchecked `- [ ]` items in either backlog that the project owner may want to pick up manually or via the self-improvement loop (`work_source.area = "loops"` or `"framework"`).
## Output
Create or update the following inside {project}/.automaton/:
+1 -1
View File
@@ -1,2 +1,2 @@
#!/usr/bin/env bash
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/private/var/folders/f5/yv0dzbnx47x3yp8sc_2519gh0000gn/T/pytest-of-laptran/pytest-130/test_uninstall_via_disabled_re0"
+214 -61
View File
@@ -7,14 +7,18 @@ the next action in an autopilot workflow.
Usage:
python autopilot.py --project /path/to/project Drive one step forward
python autopilot.py --project . --loop Run continuous loop
python autopilot.py --project . --execute Drive and execute transitions
python autopilot.py --project . --loop Run continuous drive loop
python autopilot.py --project . --detect-stuck Detect stuck tasks
python autopilot.py --project . --summary Show autopilot summary
python autopilot.py --install-schedule [--interval N] Install OS scheduler unit
"""
from __future__ import annotations
import argparse
import platform
import subprocess
import sys
import time
from pathlib import Path
@@ -24,15 +28,17 @@ AUTOMATON_DIR = Path.home() / ".automaton"
PHASE_PRIORITY = {
"referee": 12,
"doc_review": 10,
"adversarial_bug_find": 8,
"bug_find": 6,
"implement": 5,
"test_design": 4,
"design": 3,
"decomposition": 2,
"research": 1,
"new": 0,
"doc_review": 11,
"adversarial_bug_find": 10,
"bug_find": 9,
"code_review": 8,
"implement": 7,
"test_design": 6,
"design": 5,
"decomposition": 4,
"research": 3,
"new": 2,
"human_intervention": 1,
}
@@ -47,6 +53,23 @@ def _find_project_dir(project_arg: Optional[str]) -> Path:
return AUTOMATON_DIR
def _run_transition(task_name: str, target: str, project_dir: Path) -> int:
"""Execute a status.py --transition and print output."""
status_py = str(AUTOMATON_DIR / "scripts" / "status.py")
cmd = [
sys.executable, status_py,
"--transition", target,
"--task", task_name,
"--project", str(project_dir),
]
res = subprocess.run(cmd, capture_output=True, text=True)
if res.stdout:
print(res.stdout, end="")
if res.stderr:
print(res.stderr, end="", file=sys.stderr)
return res.returncode
def _base_phase(phase: str) -> str:
return phase.split(":")[0]
@@ -60,13 +83,28 @@ def _read_state(task_path: Path) -> Optional[str]:
def _all_tasks(project_dir: Path) -> list[tuple[str, Path]]:
tasks_dir = project_dir / "tasks"
if not tasks_dir.is_dir():
"""Scan all tasks (including subtasks) for a project.
Mirrors ``status.py:_all_task_dirs``: skips ``tasks/complete/``,
recurses into ``subtasks/``, uses ``.automaton/tasks`` for non-framework
projects.
"""
if project_dir == AUTOMATON_DIR:
base = AUTOMATON_DIR / "tasks"
else:
base = project_dir / ".automaton" / "tasks"
if not base.is_dir():
return []
result = []
for subdir in sorted(tasks_dir.iterdir()):
if subdir.is_dir():
result.append((subdir.name, subdir))
for entry in sorted(base.iterdir()):
if not entry.is_dir() or entry.name.startswith(".") or entry.name == "complete":
continue
result.append((entry.name, entry))
subtasks = entry / "subtasks"
if subtasks.exists():
for sub in sorted(subtasks.iterdir()):
if sub.is_dir() and not sub.name.startswith("."):
result.append((f"{entry.name}/{sub.name}", sub))
return result
@@ -87,11 +125,16 @@ def scan_all_tasks(project_dir: Path) -> list[dict]:
def is_terminal(task: dict) -> bool:
"""Check if a task is in a terminal state."""
"""Check if a task is in a terminal state.
Only ``complete`` is terminal. ``human_intervention`` has legal
transitions (→ referee, → complete) so the autopilot can still
suggest next steps for it.
"""
phase = task.get("phase")
if phase is None:
return False
return phase in ("complete", "human_intervention")
return phase == "complete"
def needs_user_input(task: dict) -> bool:
@@ -101,6 +144,8 @@ def needs_user_input(task: dict) -> bool:
return False
if phase.endswith(":awaiting_approval"):
return True
if phase == "human_intervention":
return True
task_path = Path(task["path"])
verdict_file = task_path / "VERDICT.md"
if verdict_file.exists():
@@ -209,6 +254,24 @@ def cmd_summary(args):
return 0
def _suggest_or_execute(args, task_name: str, target: str, project_dir: Path,
description: str = "") -> int:
"""Print the suggested transition, or execute it if --execute is set."""
if description:
print(f"→ {description}")
if getattr(args, "execute", False):
rc = _run_transition(task_name, target, project_dir)
if rc == 0:
print(f"✓ {task_name}: transitioned to {target}")
else:
print(f"✗ {task_name}: transition to {target} failed")
return rc
print(
f" python ~/.automaton/scripts/status.py --transition {target} --task {task_name} --project {project_dir}"
)
return 0
def cmd_drive(args):
"""Drive one step: find the best task to advance and output instructions."""
project_dir = _find_project_dir(args.project)
@@ -246,6 +309,13 @@ def cmd_drive(args):
print(
f" python ~/.automaton/scripts/status.py --approve --task {t['name']} --project {project_dir}"
)
elif t["phase"] == "human_intervention":
print(
f" Review {t['name']} — transition to referee or complete:"
)
print(
f" python ~/.automaton/scripts/status.py --transition referee --task {t['name']} --project {project_dir}"
)
else:
print(f" Review {t['name']}/VERDICT.md and take action")
return 0
@@ -259,88 +329,98 @@ def cmd_drive(args):
print()
phase = task["phase"]
name = task["name"]
base = _base_phase(phase) if phase else "unknown"
if base == "new":
print("→ Transition to research:")
print(
f" python ~/.automaton/scripts/status.py --transition research --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "research", project_dir,
"Transition to research"
)
elif base == "research":
if phase == "research":
print("→ Generate SPEC.md, then transition to awaiting_approval:")
print(
f" python ~/.automaton/scripts/status.py --transition research:awaiting_approval --task {task['name']} --project {project_dir}"
print("→ Generate SPEC.md, then request approval:")
_suggest_or_execute(
args, name, "research:awaiting_approval", project_dir
)
elif phase == "research:approved":
print("→ Transition to next phase (decomposition/design/implement):")
print(
f" python ~/.automaton/scripts/status.py --transition decomposition --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "decomposition", project_dir,
"Transition to decomposition"
)
elif base == "decomposition":
if phase == "decomposition":
print("→ Generate DECOMPOSITION.md, then transition to awaiting_approval:")
print(
f" python ~/.automaton/scripts/status.py --transition decomposition:awaiting_approval --task {task['name']} --project {project_dir}"
print("→ Generate DECOMPOSITION.md, then request approval:")
_suggest_or_execute(
args, name, "decomposition:awaiting_approval", project_dir
)
elif phase == "decomposition:approved":
print("→ Create sub-tasks from DECOMPOSITION.md, then complete parent:")
print(
f" python ~/.automaton/scripts/status.py --transition complete --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "complete", project_dir,
"All subtasks complete — mark parent done"
)
elif base == "design":
if phase == "design":
print("→ Generate DESIGN.md, then transition to awaiting_approval:")
print(
f" python ~/.automaton/scripts/status.py --transition design:awaiting_approval --task {task['name']} --project {project_dir}"
print("→ Generate DESIGN.md, then request approval:")
_suggest_or_execute(
args, name, "design:awaiting_approval", project_dir
)
elif phase == "design:approved":
print("→ Transition to test_design or implement:")
print(
f" python ~/.automaton/scripts/status.py --transition test_design --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "test_design", project_dir,
"Transition to test_design"
)
elif base == "test_design":
if phase == "test_design":
print("→ Generate TEST_PLAN.md, then transition to awaiting_approval:")
print(
f" python ~/.automaton/scripts/status.py --transition test_design:awaiting_approval --task {task['name']} --project {project_dir}"
print("→ Generate TEST_PLAN.md, then request approval:")
_suggest_or_execute(
args, name, "test_design:awaiting_approval", project_dir
)
elif phase == "test_design:approved":
print("→ Transition to implement:")
print(
f" python ~/.automaton/scripts/status.py --transition implement --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "implement", project_dir,
"Transition to implement"
)
elif base == "implement":
print("→ Write implementation, generate IMPLEMENTATION.md, then transition:")
print(
f" python ~/.automaton/scripts/status.py --transition bug_find --task {task['name']} --project {project_dir}"
print("→ Write implementation, then request code review:")
_suggest_or_execute(
args, name, "code_review", project_dir
)
elif base == "code_review":
if phase == "code_review":
print("→ Generate CODE_REVIEW.md, then request approval:")
_suggest_or_execute(
args, name, "code_review:awaiting_approval", project_dir
)
elif phase == "code_review:approved":
_suggest_or_execute(
args, name, "bug_find", project_dir,
"Transition to bug_find"
)
elif base == "bug_find":
print("→ Generate BUG_REPORT.md, then transition:")
print(
f" python ~/.automaton/scripts/status.py --transition adversarial_bug_find --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "adversarial_bug_find", project_dir
)
elif base == "adversarial_bug_find":
print("→ Generate ADVERSARIAL_BUG_REPORT.md, then transition:")
print(
f" python ~/.automaton/scripts/status.py --transition doc_review --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "doc_review", project_dir
)
elif base == "doc_review":
print("→ Generate DOC_REVIEW.md, then transition:")
print(
f" python ~/.automaton/scripts/status.py --transition referee --task {task['name']} --project {project_dir}"
_suggest_or_execute(
args, name, "referee", project_dir
)
elif base == "referee":
print("→ Generate VERDICT.md, then transition to complete:")
print(
f" python ~/.automaton/scripts/status.py --transition complete --task {task['name']} --project {project_dir}"
print("→ Generate VERDICT.md, then complete:")
_suggest_or_execute(
args, name, "complete", project_dir
)
elif base == "human_intervention":
print(
"→ Task needs human intervention. Review and transition to referee or complete:"
)
print(
f" python ~/.automaton/scripts/status.py --transition referee --task {task['name']} --project {project_dir}"
print("→ Task needs human intervention. Transition to referee or complete:")
_suggest_or_execute(
args, name, "referee", project_dir
)
return 0
@@ -348,6 +428,9 @@ def cmd_drive(args):
def cmd_loop(args):
"""Run the drive loop continuously until blocked or complete."""
if not getattr(args, "execute", False):
args.execute = True
print("(auto-enabling --execute for loop mode)")
max_iterations = args.max_iterations or 100
for iteration in range(1, max_iterations + 1):
print(f"--- Iteration {iteration}/{max_iterations} ---")
@@ -382,17 +465,85 @@ def cmd_stuck(args):
return 0
def cmd_install_schedule(args):
"""Install an OS scheduler unit that runs autopilot --drive --execute periodically."""
project_dir = _find_project_dir(args.project)
interval = getattr(args, "interval", None) or 60
autopilot_path = AUTOMATON_DIR / "scripts" / "autopilot.py"
system = platform.system()
if system == "Darwin":
plist_dir = Path.home() / "Library" / "LaunchAgents"
plist_dir.mkdir(parents=True, exist_ok=True)
plist_path = plist_dir / "com.automaton.autopilot.plist"
plist = (
f"<?xml version=\"1.0\" encoding=\"UTF-8\"?>\n"
f"<!DOCTYPE plist PUBLIC \"-//Apple//DTD PLIST 1.0//EN\" "
f"\"http://www.apple.com/DTDs/PropertyList-1.0.dtd\">\n"
f"<plist version=\"1.0\">\n"
f"<dict>\n"
f" <key>Label</key><string>com.automaton.autopilot</string>\n"
f" <key>ProgramArguments</key>\n"
f" <array>\n"
f" <string>{sys.executable}</string>\n"
f" <string>{autopilot_path}</string>\n"
f" <string>--project</string>\n"
f" <string>{project_dir}</string>\n"
f" <string>--drive</string>\n"
f" <string>--execute</string>\n"
f" </array>\n"
f" <key>StartInterval</key><integer>{interval}</integer>\n"
f" <key>RunAtLoad</key><true/>\n"
f" <key>StandardOutPath</key><string>{AUTOMATON_DIR}/logs/autopilot-stdout.log</string>\n"
f" <key>StandardErrorPath</key><string>{AUTOMATON_DIR}/logs/autopilot-stderr.log</string>\n"
f"</dict>\n"
f"</plist>\n"
)
plist_path.write_text(plist)
(AUTOMATON_DIR / "logs").mkdir(parents=True, exist_ok=True)
print(f"Installed launchd unit: {plist_path}")
print(f"Interval: {interval}s")
print(f"Logs: {AUTOMATON_DIR / 'logs' / 'autopilot-*.log'}")
print("Load with: launchctl load ~/Library/LaunchAgents/com.automaton.autopilot.plist")
print("Unload with: launchctl unload ~/Library/LaunchAgents/com.automaton.autopilot.plist")
return 0
elif system == "Linux":
cron_line = f"*/{max(1, interval // 60)} * * * * {sys.executable} {autopilot_path} --project {project_dir} --drive --execute >> {AUTOMATON_DIR}/logs/autopilot.log 2>&1"
(AUTOMATON_DIR / "logs").mkdir(parents=True, exist_ok=True)
print("Add to crontab:")
print(f" {cron_line}")
return 0
elif system == "Windows":
print("Windows: use Task Scheduler to run:")
print(f" {sys.executable} {autopilot_path} --project {project_dir} --drive --execute")
return 0
print(f"ERROR: unsupported platform '{system}'")
return 1
def main():
parser = argparse.ArgumentParser(description="Automaton autopilot runtime")
parser.add_argument("--project", help="Project root directory")
parser.add_argument(
"--drive", action="store_true", help="Drive one step forward (default)"
)
parser.add_argument(
"--execute", action="store_true",
help="Execute transitions instead of printing suggestions"
)
parser.add_argument("--summary", action="store_true", help="Show autopilot summary")
parser.add_argument(
"--stuck", action="store_true", dest="detect_stuck", help="Detect stuck tasks"
)
parser.add_argument("--loop", action="store_true", help="Run continuous drive loop")
parser.add_argument(
"--install-schedule", action="store_true",
help="Install OS scheduler unit that runs --drive --execute periodically"
)
parser.add_argument(
"--interval", type=int, default=60,
help="Tick interval in seconds for --install-schedule (default: 60)"
)
parser.add_argument(
"--max-iterations", type=int, help="Max iterations for --loop (default: 100)"
)
@@ -407,6 +558,8 @@ def main():
args = parser.parse_args()
if args.install_schedule:
return cmd_install_schedule(args)
if args.summary:
return cmd_summary(args)
if args.detect_stuck:
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env python3
"""Probe opencode.json and localhost endpoints to produce a candidate models.json.
Usage:
python3 scripts/detect_models.py [--json] [--write]
Without --json, prints a human-readable report.
With --json, emits the candidate models.json to stdout as the last JSON line.
With --write, writes the candidate to ~/.automaton/models.json (idempotent,
never overwrites an existing file unless --force is also given).
Probing strategy (stdlib only):
1. Parse opencode.json (or opencode.jsonc) for configured provider+model pairs.
2. Probe localhost endpoints to find locally-running LLM servers:
- http://localhost:8080/v1/models (llama.cpp / generic OpenAI-compatible)
- http://localhost:11434/api/tags (Ollama)
- http://localhost:1234/v1/models (LM Studio)
- http://localhost:8000/v1/models (vLLM)
3. Merge results into a candidate models.json.
"""
from __future__ import annotations
import json
import os
import re
import sys
from pathlib import Path
from typing import Optional
AUTOMATON_DIR = Path.home() / ".automaton"
# ---------------------------------------------------------------------------
# opencode.json parsing
# ---------------------------------------------------------------------------
def _find_opencode_json() -> Optional[Path]:
"""Locate the opencode config file (opencode.json or opencode.jsonc)."""
candidates = [
Path.cwd() / "opencode.json",
Path.cwd() / "opencode.jsonc",
AUTOMATON_DIR / "opencode.json",
AUTOMATON_DIR / "opencode.jsonc",
Path.home() / ".opencode.json",
Path.home() / ".config" / "opencode" / "opencode.json",
Path.home() / ".config" / "opencode" / "opencode.jsonc",
]
for p in candidates:
if p.exists():
return p
return None
def _parse_opencode_models(config_path: Path) -> list[dict]:
"""Extract model entries from an opencode.json config.
Expected structure (common patterns):
{
"providers": {
"opencode": { "model": "glm-4.6", ... },
...
}
}
or a flatter:
{
"model": "glm-4.6",
...
}
"""
try:
content = config_path.read_text(encoding="utf-8")
except OSError:
return []
# Strip JSONC comments (// line comments only, sufficient for our use)
content = re.sub(r"//.*", "", content)
try:
data = json.loads(content)
except json.JSONDecodeError:
return []
if not isinstance(data, dict):
return []
models: list[dict] = []
seen: set[str] = set()
# Check top-level "model" field (single-model config)
single = data.get("model")
if isinstance(single, str) and single not in seen:
seen.add(single)
models.append({"name": single, "provider": "opencode", "context_window": None, "location": "remote"})
# Check providers dict
providers = data.get("providers") or {}
for prov_name, prov_cfg in providers.items():
if isinstance(prov_cfg, dict):
model_name = prov_cfg.get("model")
if isinstance(model_name, str) and model_name not in seen:
seen.add(model_name)
models.append({"name": model_name, "provider": prov_name, "context_window": None, "location": "remote"})
# Check "models" list (explicit model roster)
model_list = data.get("models")
if isinstance(model_list, list):
for entry in model_list:
if isinstance(entry, dict):
name = entry.get("name") or entry.get("model")
if isinstance(name, str) and name not in seen:
seen.add(name)
models.append({
"name": name,
"provider": entry.get("provider", "opencode"),
"context_window": entry.get("context_window"),
"location": entry.get("location", "remote"),
})
return models
# ---------------------------------------------------------------------------
# Localhost probing
# ---------------------------------------------------------------------------
def _fetch_json(url: str, timeout: int = 5) -> Optional[dict]:
"""Fetch a JSON response from a URL using urllib (stdlib)."""
import urllib.request
import urllib.error
try:
req = urllib.request.Request(url, method="GET")
with urllib.request.urlopen(req, timeout=timeout) as resp:
body = resp.read().decode("utf-8")
return json.loads(body)
except (OSError, urllib.error.URLError, json.JSONDecodeError, ValueError):
return None
def _probe_ollama() -> list[dict]:
"""Probe Ollama: GET http://localhost:11434/api/tags → models[].name"""
data = _fetch_json("http://localhost:11434/api/tags")
if not data:
return []
models_list = data.get("models") or []
return [
{"name": m.get("name"), "provider": "ollama", "context_window": None, "location": "http://localhost:11434"}
for m in models_list
if isinstance(m, dict) and isinstance(m.get("name"), str)
]
def _probe_openai_compatible(url: str, provider: str) -> list[dict]:
"""Probe an OpenAI-compatible /v1/models endpoint."""
data = _fetch_json(url)
if not data:
return []
model_list = data.get("data") or []
return [
{"name": m.get("id"), "provider": provider, "context_window": None, "location": url}
for m in model_list
if isinstance(m, dict) and isinstance(m.get("id"), str)
]
_ENDPOINTS = [
("http://localhost:8080/v1/models", "llama.cpp"),
("http://localhost:11434/api/tags", "ollama"), # handled separately above
("http://localhost:1234/v1/models", "lm-studio"),
("http://localhost:8000/v1/models", "vllm"),
]
def _probe_localhost() -> list[dict]:
"""Probe all known localhost endpoints and merge results."""
seen_names: set[str] = set()
models: list[dict] = []
for url, provider in _ENDPOINTS:
if provider == "ollama":
result = _probe_ollama()
else:
result = _probe_openai_compatible(url, provider)
for m in result:
n = m.get("name")
if isinstance(n, str) and n not in seen_names:
seen_names.add(n)
models.append(m)
return models
# ---------------------------------------------------------------------------
# Merge & write
# ---------------------------------------------------------------------------
def build_candidate_models(probe_local: bool = True) -> dict:
"""Build a candidate models.json dict.
1. Parse models from opencode.json
2. Optionally probe localhost endpoints
3. Merge: opencode config models come first; local probes fill in gaps.
4. Build result with default, advised, models[].
"""
opencode_path = _find_opencode_json()
config_models: list[dict] = []
if opencode_path:
config_models = _parse_opencode_models(opencode_path)
local_models: list[dict] = []
if probe_local:
local_models = _probe_localhost()
# Merge: key by name, config models take priority (unordered)
merged: dict[str, dict] = {}
for m in config_models:
n = m["name"]
if n not in merged:
merged[n] = m
for m in local_models:
n = m.get("name")
if n and n not in merged:
merged[n] = m
models_list = list(merged.values())
# Determine default: first config model, or first local model, or empty
default_name: Optional[str] = None
if config_models:
default_name = config_models[0].get("name")
elif local_models:
default_name = local_models[0].get("name")
# Determine advised: if only 0-1 models, set advised=true; else false
advised = len(models_list) <= 1
result: dict = {
"schema_version": 1,
"default": default_name,
"advised": advised,
"models": models_list,
}
return result
def write_models_file(candidate: dict, force: bool = False) -> bool:
"""Write candidate models.json to AUTOMATON_DIR.
Never overwrites an existing file unless force=True.
Returns True if written, False if skipped.
"""
target = AUTOMATON_DIR / "models.json"
if target.exists() and not force:
return False
target.write_text(json.dumps(candidate, indent=2) + "\n")
return True
def format_report(candidate: dict) -> str:
"""Human-readable report of the candidate models."""
lines = []
lines.append("=== Model Detection Report ===")
lines.append("")
source = "No opencode.json found" if not _find_opencode_json() else f"Config: {_find_opencode_json()}"
lines.append(f"Source: {source}")
lines.append("")
models = candidate.get("models", [])
if not models:
lines.append("No models detected.")
else:
lines.append(f"Detected {len(models)} model(s):")
for m in models:
loc = m.get("location", "unknown")
prov = m.get("provider", "?")
ctx = m.get("context_window")
ctx_str = f", context: {ctx}" if ctx else ""
lines.append(f" - {m['name']} ({prov}, {loc}{ctx_str})")
lines.append("")
lines.append(f"Default: {candidate.get('default', 'none')}")
lines.append(f"Advised: {candidate.get('advised', False)}")
lines.append(f"Mode: {'multi-LLM' if len(models) >= 2 else 'single-LLM'}")
lines.append("")
target = AUTOMATON_DIR / "models.json"
if target.exists():
lines.append(f"models.json already exists at {target} (use --force to overwrite)")
else:
lines.append(f"Ready to write to {target} (use --write to create)")
return "\n".join(lines)
def main() -> int:
import argparse
parser = argparse.ArgumentParser(description="Detect available LLM models and write models.json")
parser.add_argument("--json", action="store_true", help="Output candidate JSON on last line")
parser.add_argument("--write", action="store_true", help="Write candidate models.json to ~/.automaton/ (idempotent)")
parser.add_argument("--force", action="store_true", help="Overwrite existing models.json")
parser.add_argument("--no-probe", action="store_true", help="Skip localhost endpoint probing")
args = parser.parse_args()
candidate = build_candidate_models(probe_local=not args.no_probe)
if args.write:
written = write_models_file(candidate, force=args.force)
if written:
print(f"Written models.json to {AUTOMATON_DIR / 'models.json'}")
else:
print(f"Skipped: {AUTOMATON_DIR / 'models.json'} already exists (use --force to overwrite)")
if args.json:
print(json.dumps(candidate))
else:
print(format_report(candidate))
return 0
if __name__ == "__main__":
sys.exit(main())
+5
View File
@@ -3,9 +3,14 @@
#
# Usage: bash ~/.automaton/scripts/install-hooks.sh [project-path]
#
# Called automatically by onboard-project.sh. Can also be run manually
# after framework updates to refresh hooks.
#
# Installs pre-commit and pre-push hooks. The pre-commit hook blocks
# commits when no task is in implement/doc_review. The pre-push hook
# blocks pushes in the same condition, catching --no-verify bypasses.
#
# Next step: python3 ~/.automaton/scripts/status.py --create-task --project .
set -euo pipefail
+34 -28
View File
@@ -3,27 +3,39 @@ set -e
FRAMEWORK_DIR="$HOME/.automaton"
if [ -d "$FRAMEWORK_DIR" ]; then
echo "automaton already installed at $FRAMEWORK_DIR"
echo "Run './update.sh' to update."
exit 0
fi
GIT_URL="${1:-}"
if [ -z "$GIT_URL" ]; then
echo "ERROR: Git URL required."
echo "Usage: ./install.sh <git-url>"
echo "Example: ./install.sh https://github.com/user/automaton.git"
if [ ! -d "$FRAMEWORK_DIR" ]; then
# Fresh install — need a Git URL to clone
if [ -z "$GIT_URL" ]; then
echo "ERROR: Git URL required for fresh install."
echo ""
echo "Usage:"
echo " curl -fsSL https://raw.githubusercontent.com/<user>/automaton/main/scripts/install.sh | bash -s -- <git-url>"
echo ""
echo " --- or ---"
echo ""
echo " git clone <git-url> ~/.automaton"
echo " bash ~/.automaton/scripts/install.sh"
echo ""
echo "The framework is cloned to ~/.automaton. Choose your URL carefully"
echo "as it cannot be changed later without reinstalling."
exit 1
fi
echo "Cloning automaton from $GIT_URL to $FRAMEWORK_DIR..."
git clone "$GIT_URL" "$FRAMEWORK_DIR"
echo ""
else
echo "automaton already installed at $FRAMEWORK_DIR — running setup steps..."
if [ -n "$GIT_URL" ]; then
echo "Note: Git URL argument ignored because ~/.automaton already exists."
echo "To update, run: cd ~/.automaton && ./update.sh"
fi
echo ""
echo "The framework is cloned to ~/.automaton. Choose your URL carefully"
echo "as it cannot be changed later without reinstalling."
exit 1
fi
echo "Cloning automaton from $GIT_URL to $FRAMEWORK_DIR..."
git clone "$GIT_URL" "$FRAMEWORK_DIR"
# --- Everything below is idempotent and runs on both fresh and existing installs ---
echo ""
echo "=== VRAM / Context Detection ==="
echo "Detecting your system's VRAM to recommend task decomposition settings..."
echo ""
@@ -31,7 +43,7 @@ echo ""
# Run VRAM detection script if it exists
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
# Run in project-dir context so it can read framework overhead
detection_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>&1)
detection_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>&1 || true)
# Extract JSON output (the block after "=== JSON Output ===")
json_output=$(echo "$detection_output" | sed -n '/=== JSON Output ===/,$p' | tail -n +2)
@@ -40,12 +52,9 @@ if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
echo "$detection_output"
# Extract key values from JSON using Python
recommended_k=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["recommended_k"])')
max_peak_kb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["max_peak_context_kb"])')
headroom=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["headroom"])')
gpu_vram=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["gpu_vram_gb"])')
ram_gb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["ram_gb"])')
model_context=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["model_context_kb"])')
recommended_k=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["recommended_k"])' 2>/dev/null || echo "?")
max_peak_kb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["max_peak_context_kb"])' 2>/dev/null || echo "?")
headroom=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["headroom"])' 2>/dev/null || echo "?")
echo ""
echo "=== Recommended VRAM Configuration ==="
@@ -72,11 +81,8 @@ echo ""
echo "Installation complete."
echo ""
echo "Next steps:"
echo " 1. cd into a project and run the onboarding prompt"
echo " 2. In each project that uses git, install the automaton hooks:"
echo " bash ~/.automaton/scripts/install-hooks.sh /path/to/project"
echo ""
echo "These hooks block commits and pushes when no task is in an edit-allowed phase."
echo " 1. Onboard a project: bash ~/.automaton/scripts/onboard-project.sh /path/to/project"
echo " 2. Or tell your agent: 'Onboard this project into automaton'"
echo ""
# Register pre-edit guards for detected harnesses
+54 -15
View File
@@ -375,6 +375,9 @@ def _invoke_harness(
"""Build the harness command from loop.json and invoke it. Returns stdout.
extras: substitution tokens specific to this role ({artifact}, {verdict}, etc).
If extras contains a "model" key, the ``{model}`` token in the harness
command is substituted. The caller is responsible for passing the model
via extras (extracted from loop.json role config or manifest default).
"""
resolved_prompt = prompt_path
if loop_path is not None:
@@ -524,6 +527,24 @@ def _role_prompt(cfg: dict, role: str) ->Optional[str]:
return role_cfg.get("prompt")
def _role_model(cfg: dict, role: str) -> Optional[str]:
"""Get the model configured for a role in loop.json, or the manifest default."""
roles = cfg.get("roles") or {}
role_cfg = roles.get(role) or {}
model = role_cfg.get("model")
if model:
return model
models_file = AUTOMATON_DIR / "models.json"
if models_file.exists():
try:
import json as _mj
manifest = _mj.loads(models_file.read_text())
return manifest.get("default")
except (OSError, _mj.JSONDecodeError):
pass
return None
# ---------------------------------------------------------------------------
# Work sources (task add-goal-mode)
# ---------------------------------------------------------------------------
@@ -809,27 +830,39 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
out_dir = _outputs_dir(loop_path)
tick_num = state.get('iteration_count', 0) + 1
impl_output = str(out_dir / f"tick{tick_num}-implement.json")
impl_model = _role_model(cfg, "implement")
implement_extras = {
"output": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint,
}
if impl_model:
implement_extras["model"] = impl_model
implement_stdout = _invoke_harness(
harness_cfg, "implement", implement_prompt, cwd,
extras={"output": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
extras=implement_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(impl_output)).write_text(implement_stdout)
# Step 6: spawn Verify
verify_prompt = _role_prompt(cfg, "verify") or ""
verify_output = str(out_dir / f"tick{tick_num}-verify.json")
verify_model = _role_model(cfg, "verify")
verify_extras = {
"output": verify_output,
"artifact": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint,
}
if verify_model:
verify_extras["model"] = verify_model
verify_stdout = _invoke_harness(
harness_cfg, "verify", verify_prompt, cwd,
extras={"output": verify_output,
"artifact": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
extras=verify_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(verify_output)).write_text(verify_stdout)
@@ -854,12 +887,18 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
# Step 9: spawn Orchestrate
orch_prompt = _role_prompt(cfg, "orchestrate") or ""
orch_output = str(out_dir / f"tick{tick_num}-orchestrate.json")
orch_model = _role_model(cfg, "orchestrate")
orch_extras = {
"output": orch_output,
"verdict": json.dumps(verdict),
"current_task": current_task,
"current_phase": state.get("current_phase", ""),
}
if orch_model:
orch_extras["model"] = orch_model
orch_stdout = _invoke_harness(
harness_cfg, "orchestrate", orch_prompt, cwd,
extras={"output": orch_output,
"verdict": json.dumps(verdict),
"current_task": current_task,
"current_phase": state.get("current_phase", "")},
extras=orch_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(orch_output)).write_text(orch_stdout)
+184
View File
@@ -0,0 +1,184 @@
#!/usr/bin/env bash
# onboard-project.sh — Bootstrap automaton in a new or existing project.
#
# Usage:
# bash ~/.automaton/scripts/onboard-project.sh /path/to/project
#
# Creates .automaton/ skeleton, detects models, creates config, inits git,
# installs hooks, and verifies everything works.
set -euo pipefail
FRAMEWORK_DIR="$HOME/.automaton"
PROJECT_DIR="${1:-}"
# Colors
RED='\033[0;31m'
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
BLUE='\033[0;34m'
NC='\033[0m' # No Color
info() { echo -e "${BLUE}INFO:${NC} $1"; }
ok() { echo -e "${GREEN}OK:${NC} $1"; }
warn() { echo -e "${YELLOW}WARN:${NC} $1"; }
error() { echo -e "${RED}ERROR:${NC} $1"; }
# --- Argument checks ---
if [ -z "$PROJECT_DIR" ]; then
error "Usage: bash ~/.automaton/scripts/onboard-project.sh /path/to/project"
exit 1
fi
PROJECT_DIR="$(cd "$PROJECT_DIR" 2>/dev/null && pwd)" || true
if [ -z "$PROJECT_DIR" ] || [ ! -d "$PROJECT_DIR" ]; then
echo ""
error "'$1' does not exist."
echo " Create it first: mkdir -p '$1'"
echo " Then re-run this script."
exit 1
fi
if [ ! -d "$FRAMEWORK_DIR/scripts" ]; then
error "Framework not found at $FRAMEWORK_DIR."
echo " Install the framework first:"
echo " curl -fsSL https://raw.githubusercontent.com/<user>/automaton/main/scripts/install.sh | bash -s -- <git-url>"
echo " Or: git clone <git-url> ~/.automaton && bash ~/.automaton/scripts/install.sh"
exit 1
fi
echo ""
echo "========================================"
echo " Automaton Project Onboarding"
echo " Project: $PROJECT_DIR"
echo "========================================"
echo ""
# --- Step 1: Create .automaton/ skeleton ---
AUTO_DIR="$PROJECT_DIR/.automaton"
if [ -d "$AUTO_DIR" ]; then
warn "$AUTO_DIR already exists — skipping skeleton creation"
else
info "Creating .automaton/ skeleton..."
mkdir -p "$AUTO_DIR/tasks" "$AUTO_DIR/loops" "$AUTO_DIR/design"
ok "Created $AUTO_DIR/"
fi
# --- Step 2: Models ---
MODELS_FILE="$AUTO_DIR/models.json"
if [ -f "$MODELS_FILE" ]; then
warn "$MODELS_FILE already exists — skipping model detection"
else
info "Probing local models..."
if python3 "$FRAMEWORK_DIR/scripts/detect_models.py" --write --project "$PROJECT_DIR" 2>/dev/null; then
ok "Detected models written to $MODELS_FILE"
else
info "Auto-detection failed. Creating minimal models.json..."
cat > "$MODELS_FILE" <<- 'EOF'
{
"models": [
{"name": "default-model", "provider": "local", "context": 32768}
],
"default": "default-model"
}
EOF
warn "Edit $MODELS_FILE to set your actual model(s)."
fi
fi
# --- Step 3: Config ---
CONFIG_FILE="$AUTO_DIR/config.md"
if [ -f "$CONFIG_FILE" ]; then
warn "$CONFIG_FILE already exists — skipping"
else
info "Creating config.md..."
PROJECT_NAME="$(basename "$PROJECT_DIR")"
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
vram_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>/dev/null || true)
json_part=$(echo "$vram_output" | sed -n '/=== JSON Output ===/,$p' | tail -n +2)
recommended=$(echo "$json_part" | python3 -c 'import json,sys; print(json.load(sys.stdin).get("recommended_k","16"))' 2>/dev/null || echo "16")
else
recommended="16"
fi
cat > "$CONFIG_FILE" <<- EOF
# $PROJECT_NAME — Automaton Configuration
## VRAM Configuration
- **Auto-detect**: Yes
- **Target context**: ${recommended}k tokens
- **Headroom**: 25%
- **Max peak context per sub-task**: $((recommended * 3 / 4))k tokens
## Model Configuration
# Uses models.json for model divergence enforcement.
# Default model is read from models.json's "default" key.
EOF
ok "Created $CONFIG_FILE"
fi
# --- Step 4: Project name ---
NAME_FILE="$AUTO_DIR/project-name.md"
if [ -f "$NAME_FILE" ]; then
warn "$NAME_FILE already exists — skipping"
else
PROJECT_NAME="$(basename "$PROJECT_DIR")"
echo "$PROJECT_NAME" > "$NAME_FILE"
ok "Created $NAME_FILE ($PROJECT_NAME)"
fi
# --- Step 5: Git ---
GIT_DIR="$PROJECT_DIR/.git"
if [ -d "$GIT_DIR" ]; then
ok "Git repository already initialized"
else
info "Initializing git repository..."
cd "$PROJECT_DIR" && git init
ok "Git initialized"
fi
# --- Step 6: Git hooks ---
if [ -d "$GIT_DIR" ]; then
info "Installing git hooks..."
bash "$FRAMEWORK_DIR/scripts/install-hooks.sh" "$PROJECT_DIR"
fi
# --- Step 7: .gitignore ---
GITIGNORE="$PROJECT_DIR/.gitignore"
if [ -f "$GITIGNORE" ]; then
if ! grep -q ".automaton/tasks/" "$GITIGNORE" 2>/dev/null; then
echo "" >> "$GITIGNORE"
echo "# Automaton" >> "$GITIGNORE"
echo ".automaton/tasks/" >> "$GITIGNORE"
echo ".automaton/loops/*/worktree/" >> "$GITIGNORE"
warn "Added automaton entries to .gitignore"
fi
else
cat > "$GITIGNORE" <<- 'EOF'
# Automaton
.automaton/tasks/
.automaton/loops/*/worktree/
.automaton/loops/*/outputs/
EOF
ok "Created .gitignore with automaton entries"
fi
# --- Step 8: Verify ---
info "Verifying setup..."
cd "$FRAMEWORK_DIR"
python3 "$FRAMEWORK_DIR/scripts/status.py" --project "$PROJECT_DIR" --audit 2>&1 | head -5 || true
echo ""
echo "========================================"
echo -e "${GREEN} Onboarding complete!${NC}"
echo "========================================"
echo ""
echo " Project: $PROJECT_DIR"
echo " Config: $CONFIG_FILE"
echo " Models: $MODELS_FILE"
echo ""
echo " Next steps:"
echo " 1. cd $PROJECT_DIR"
echo " 2. Create a task:"
echo " python3 ~/.automaton/scripts/status.py --create-task my-first-task --project ."
echo " 3. Start working with your agent."
echo ""
+174
View File
@@ -0,0 +1,174 @@
#!/usr/bin/env bash
# pi-automaton.sh — Pi Dev automaton context printer.
#
# Prints automaton project/framework context for the user to paste as their
# first message to a Pi Dev agent. Does NOT launch pi.
#
# Usage:
# cd /path/to/project
# bash ~/.automaton/scripts/pi-automaton.sh
#
# Copy the output and paste it as your first message in Pi Dev.
set -euo pipefail
FRAMEWORK_DIR="$HOME/.automaton"
STATUS_PY="$FRAMEWORK_DIR/scripts/status.py"
# Colors
BOLD='\033[1m'
DIM='\033[2m'
NC='\033[0m'
info() { echo -e " $1"; }
dim() { echo -e " ${DIM}$1${NC}"; }
dim_nl(){ echo -e "${DIM}$1${NC}"; }
# --- Scope detection ---
CWD="$(pwd)"
if [ "$CWD" = "$FRAMEWORK_DIR" ] || [ "${CWD##"$FRAMEWORK_DIR"}" != "$CWD" ]; then
SCOPE="framework"
SCOPE_LABEL="framework mode (automaton itself)"
PROJECT_DIR="$FRAMEWORK_DIR"
elif [ -d "$CWD/.automaton" ]; then
SCOPE="project"
SCOPE_LABEL="project mode ($(basename "$CWD"))"
PROJECT_DIR="$CWD"
else
echo ""
echo "No automaton project detected in $CWD"
echo ""
echo "To onboard this project:"
echo " bash $FRAMEWORK_DIR/scripts/onboard-project.sh ."
echo ""
exit 1
fi
TASKS_DIR="$PROJECT_DIR/.automaton/tasks"
MODELS_FILE="$PROJECT_DIR/.automaton/models.json"
CONFIG_FILE="$PROJECT_DIR/.automaton/config.md"
AGENTS_FILE="$PROJECT_DIR/.automaton/AGENTS.md"
# --- Collect data ---
# Active tasks via status.py --audit --json
TASKS_JSON=""
if [ -f "$STATUS_PY" ] && [ -d "$TASKS_DIR" ]; then
TASKS_JSON=$(python3 "$STATUS_PY" --audit --json --project "$PROJECT_DIR" 2>/dev/null || true)
fi
# Default model
DEFAULT_MODEL=""
if [ -f "$MODELS_FILE" ]; then
DEFAULT_MODEL=$(python3 -c "
import json
with open('$MODELS_FILE') as f:
m = json.load(f)
print(m.get('default', ''))
" 2>/dev/null || echo "")
fi
# VRAM context snippet from config.md
CONFIG_SNIPPET=""
if [ -f "$CONFIG_FILE" ]; then
CONFIG_SNIPPET=$(grep -i 'context\|headroom\|target' "$CONFIG_FILE" 2>/dev/null | head -3 | sed 's/^/ /')
fi
# Project rules from AGENTS.md
RULES_TEXT=""
if [ -f "$AGENTS_FILE" ]; then
RULES_TEXT=$(grep -v -E '^#|^$' "$AGENTS_FILE" 2>/dev/null | head -10)
fi
# Can-edit status
CAN_EDIT_OUTPUT=""
if [ -f "$STATUS_PY" ]; then
CAN_EDIT_OUTPUT=$(python3 "$STATUS_PY" --can-edit --json --project "$PROJECT_DIR" 2>/dev/null || true)
fi
# --- Build task list ---
TASK_LINES=""
TASK_COUNT=0
if [ -n "$TASKS_JSON" ]; then
while IFS=$'\t' read -r name state edit_phase; do
if [ -n "$name" ]; then
FLAG=""
if [ "$edit_phase" = "true" ]; then
FLAG=" (edit allowed)"
elif [ "$state" != "backlog" ] && [ "$state" != "done" ] && [ "$state" != "blocked" ]; then
FLAG=" (read-only)"
fi
TASK_LINES+=" * $name\t\t$state$FLAG\n"
TASK_COUNT=$((TASK_COUNT + 1))
fi
done < <(echo "$TASKS_JSON" | python3 -c "
import json, sys
data = json.load(sys.stdin)
tasks = []
for t in data.get('tasks', []):
tasks.append((t['name'], t['state']))
# Map edit-eligible phases
EDIT_PHASES = {'implement', 'doc_review'}
for name, state in tasks:
ep = 'true' if state in EDIT_PHASES else 'false'
print(f'{name}\t{state}\t{ep}')
" 2>/dev/null || true)
fi
# --- Render ---
LINE="══════════════════════════════════════════════════"
SEP="──────────────────────────────────────────────────"
echo ""
echo -e "${BOLD}${LINE}${NC}"
echo -e "${BOLD} Automaton Context — $SCOPE_LABEL${NC}"
echo -e "${BOLD}${LINE}${NC}"
echo ""
echo -e "${BOLD}This project uses the automaton workflow framework.${NC}"
info "Tasks are tracked in $(basename "$PROJECT_DIR")/.automaton/tasks/"
info "and flow through phases:"
info "backlog → research → implement → code_review → bug_find → ... → complete"
echo ""
if [ -n "$TASK_LINES" ]; then
echo -e "${BOLD}Active tasks:${NC}"
echo -e "$TASK_LINES"
echo ""
fi
if [ -n "$DEFAULT_MODEL" ]; then
echo -e "${BOLD}Default model:${NC} $DEFAULT_MODEL"
echo ""
fi
if [ -n "$CONFIG_SNIPPET" ]; then
echo -e "${BOLD}Configuration:${NC}"
echo "$CONFIG_SNIPPET"
echo ""
fi
echo -e "${BOLD}To work on a task:${NC}"
info "python3 ~/.automaton/scripts/status.py --transition <phase> --task <name>"
echo ""
echo -e "${BOLD}To create a new task:${NC}"
info "python3 ~/.automaton/scripts/status.py --create-task <name>"
echo ""
if [ -n "$RULES_TEXT" ]; then
echo -e "${BOLD}Project rules (from AGENTS.md):${NC}"
echo "$RULES_TEXT" | head -5
echo ""
fi
echo -e "${BOLD}Important:${NC}"
info "Only modify files when a task is in ${BOLD}implement${NC} or ${BOLD}doc_review${NC} phase"
info "All phase transitions go through status.py"
info "The automaton-guard-pi plugin blocks edits outside allowed phases"
echo ""
echo -e "${DIM}$SEP${NC}"
dim_nl "Copy this entire block and paste it as your first message"
dim_nl "to the Pi Dev agent to provide automaton context."
echo -e "${DIM}$SEP${NC}"
echo ""
+44 -11
View File
@@ -54,17 +54,50 @@ else
echo "OpenCode: not detected (no ~/.config/opencode/opencode.json or .jsonc)"
fi
# Pi Dev guard
PI_SOURCE="$FRAMEWORK_DIR/plugins/automaton-guard-pi"
# Pi Dev extensions
PI_EXTENSIONS=(
"$FRAMEWORK_DIR/plugins/automaton-guard-pi"
"$FRAMEWORK_DIR/plugins/pi-read-tweet"
"$FRAMEWORK_DIR/plugins/pi-automaton-context"
"$FRAMEWORK_DIR/plugins/pi-automaton-tools"
)
if command -v pi &>/dev/null; then
INSTALLED=$(pi list 2>/dev/null | grep -c "automaton-guard-pi" || true)
if [ "$INSTALLED" -gt 0 ]; then
echo "Pi Dev: already installed"
else
echo "Pi Dev: installing guard extension..."
pi install "$PI_SOURCE" 2>&1 | sed 's/^/ /'
INSTALLED_PI=true
echo "Pi Dev: installed"
INSTALLED_PI=true
for ext in "${PI_EXTENSIONS[@]}"; do
ext_name=$(basename "$ext")
FOUND=$(pi list 2>/dev/null | grep -c "$ext_name" || true)
if [ "$FOUND" -gt 0 ]; then
echo "Pi Dev: $ext_name already installed"
else
echo "Pi Dev: installing $ext_name..."
pi install "$ext" 2>&1 | sed 's/^/ /'
echo "Pi Dev: $ext_name installed"
fi
done
# Offer pi-automaton startup wrapper
PI_WRAPPER="$FRAMEWORK_DIR/scripts/pi-automaton.sh"
if [ -f "$PI_WRAPPER" ]; then
echo ""
echo "Pi Dev context wrapper available at:"
echo " $PI_WRAPPER"
echo ""
echo "Before starting a Pi Dev session, run this script to print"
echo "automaton project context that you can paste as your first message."
echo ""
echo " bash ~/.automaton/scripts/pi-automaton.sh"
echo ""
# Only prompt interactively if stdin is a terminal
BIN_DIR="$HOME/bin"
if [ -t 0 ] && [ ! -f "$BIN_DIR/pi-automaton" ]; then
echo -n "Symlink to ~/bin/pi-automaton for easier access? [Y/n] "
read -r REPLY
if [ -z "$REPLY" ] || [ "$REPLY" = "y" ] || [ "$REPLY" = "Y" ]; then
mkdir -p "$BIN_DIR"
ln -sf "$PI_WRAPPER" "$BIN_DIR/pi-automaton"
echo " Created $BIN_DIR/pi-automaton → $PI_WRAPPER"
echo " (ensure ~/bin is in your PATH)"
fi
fi
fi
else
echo "Pi Dev: not detected (pi not in PATH)"
@@ -77,7 +110,7 @@ fi
if ! $INSTALLED_OPENCODE && ! $INSTALLED_PI; then
echo "No harness detected. To install a guard manually:"
echo " OpenCode: add '\"plugin\": [\"$OPENCODE_SOURCE\"]' to ~/.config/opencode/opencode.json"
echo " Pi Dev: pi install $PI_SOURCE"
echo " Pi Dev: pi install $FRAMEWORK_DIR/plugins/automaton-guard-pi"
echo ""
echo "Without a pre-edit guard, git hooks (pre-commit + pre-push)"
echo "provide enforcement at commit/push time instead."
+261 -2
View File
@@ -153,7 +153,19 @@ FORBIDDEN_ARTIFACTS = {
}
NON_ARTIFACT_FILES = {".state", ".state.tmp", ".state.lock", ".state.approvals",
".state.implementer", ".state.lastedit", "VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"}
".state.implementer", ".state.lastedit", ".state.models",
"VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"}
# Model-divergence enforcement
CONFLICT_MATRIX = {
"code_review": {"implement"},
"bug_find": {"implement"},
"adversarial_bug_find": {"implement", "bug_find"},
"referee": {"implement", "bug_find", "adversarial_bug_find"},
"loop-verify": {"loop-implement"},
}
MODELS_JSON_FILE = "models.json"
PHASE_PRIORITY = {
"referee": 12, "doc_review": 11, "adversarial_bug_find": 10,
@@ -438,6 +450,124 @@ def _lock_timeout_seconds(project: Optional[str] = None) -> int:
return val
# ---------------------------------------------------------------------------
# Model-divergence enforcement helpers
# ---------------------------------------------------------------------------
def _load_models_manifest(project: Optional[str] = None) -> Optional[dict]:
"""Load the models.json manifest for the given project.
Searches:
1. project/.automaton/models.json
2. ~/.automaton/models.json (fallback)
Returns None if no models.json exists (single-LLM mode, backward compatible).
"""
project_dir = _find_project_dir(project)
candidates = [
project_dir / ".automaton" / MODELS_JSON_FILE,
AUTOMATON_DIR / MODELS_JSON_FILE,
]
for path in candidates:
if path.exists():
try:
return json.loads(path.read_text())
except (OSError, json.JSONDecodeError):
return None
return None
def _get_model_mode(manifest: Optional[dict]) -> str:
"""Determine the model mode: 'single' or 'multi-llm'.
- Missing manifest → single-LLM (backward compatible)
- 0-1 models → single-LLM
- 2+ models → multi-LLM
"""
if manifest is None:
return "single"
models = manifest.get("models") or []
if len(models) >= 2:
return "multi-llm"
return "single"
def _check_conflict(state_models: dict, role: str, model: str, matrix: Optional[dict] = None) -> Optional[str]:
"""Check if the given model conflicts with already-filled roles.
state_models: dict of {role: model_name} from .state.models
role: the role being entered (e.g. 'code_review')
model: the model name being assigned
matrix: conflict matrix (defaults to CONFLICT_MATRIX)
Returns the name of the conflicting role, or None if no conflict.
"""
if matrix is None:
matrix = CONFLICT_MATRIX
if role not in matrix:
return None
conflicting_roles = matrix[role]
for filled_role, filled_model in state_models.items():
if filled_model == model and filled_role in conflicting_roles:
return filled_role
return None
def _read_state_models(task_path: Path) -> dict:
"""Read .state.models from the task directory. Returns {} if missing."""
f = task_path / ".state.models"
if not f.exists():
return {}
try:
data = json.loads(f.read_text())
if isinstance(data, dict):
return data
except (OSError, json.JSONDecodeError):
pass
return {}
def _write_state_models(task_path: Path, state_models: dict) -> None:
"""Write .state.models atomically."""
tmp = task_path / ".state.models.tmp"
tmp.write_text(json.dumps(state_models, indent=2, sort_keys=True) + "\n")
tmp.replace(task_path / ".state.models")
def _model_divergence_violations(project: Optional[str] = None) -> list[dict]:
"""Scan all tasks for model-divergence violations.
Returns list of violation dicts:
{"task": str, "message": str, "severity": "high", "resolved": False}
"""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode == "single":
return []
violations = []
tasks = _all_task_dirs(project)
for name, path in tasks:
sm = _read_state_models(path)
if not sm:
continue
for role, model in sm.items():
if model is None:
continue
# D8: doc_review, code_review, bug_find have no cross-conflicts
# with each other; only conflicts documented in CONFLICT_MATRIX apply.
conflict = _check_conflict(sm, role, str(model))
if conflict:
violations.append({
"task": name,
"severity": "high",
"message": f"model-divergence: role '{role}' uses model '{model}' "
f"which conflicts with role '{conflict}' (same model)",
"resolved": False,
})
return violations
# --- Command implementations ---
def cmd_show_task(args):
@@ -593,6 +723,60 @@ def cmd_transition(args):
return 1
if current == "human_intervention" and target == "complete":
_auto_update_verdict_on_complete(task_path)
# Model-divergence enforcement (Subtask 2)
# When entering a phase that maps to a role, record the model
target_base = _base_phase(target)
ROLE_PHASES = {"implement", "code_review", "bug_find", "adversarial_bug_find", "doc_review", "referee"}
if target_base in ROLE_PHASES and current != target:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
state_models = _read_state_models(task_path)
model_arg = getattr(args, "model", None)
if model_arg:
# --model explicitly provided — record advisory in single mode, check in multi
state_models[target_base] = model_arg
if mode == "multi-llm":
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 1
conflict = _check_conflict(state_models, target_base, model_arg)
if conflict:
# Remove the entry we just added
del state_models[target_base]
print(f"ERROR: Model '{model_arg}' assigned to role '{target_base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model> to specify a different model.")
return 1
elif mode == "multi-llm":
# Auto-assign: try default, then next-available non-conflicting
default = (manifest or {}).get("default")
assigned = False
if default and default in model_names:
conflict = _check_conflict(state_models, target_base, default)
if not conflict:
state_models[target_base] = default
assigned = True
if not assigned:
for m_name in model_names:
if m_name == default:
continue
conflict = _check_conflict(state_models, target_base, m_name)
if not conflict:
state_models[target_base] = m_name
assigned = True
break
if not assigned:
print(f"ERROR: Cannot auto-assign a model for role '{target_base}'. "
f"All available models conflict with already-filled roles. "
f"Use --model <name> to override.")
return 1
if model_arg or mode == "multi-llm":
_write_state_models(task_path, state_models)
if current == "implement" and target == "code_review":
lock_file = task_path / ".state.lock"
if lock_file.exists():
@@ -980,6 +1164,14 @@ def _audit_collect(args):
"halt_reason": lhalt, "current_task": ltask,
"violation": is_violation, "message": msg})
# Model-divergence violations (Category 6)
for mv in _model_divergence_violations(args.project):
violations.append({
"category": 6, "severity": mv["severity"],
"task": mv["task"], "message": mv["message"],
"resolved": False,
})
return {"violations": violations,
"loops": loops,
"total_tasks": len(tasks),
@@ -1118,7 +1310,16 @@ def cmd_audit(args):
if stuck_found == 0:
print(f"[PASS] No stuck tasks (threshold: {stuck_threshold} min)")
print("\n=== Category 6: Loops ===")
print("\n=== Category 6: Model-Divergence Violations ===")
md_violations = _model_divergence_violations(args.project)
if md_violations:
for v in md_violations:
print(f"[FAIL] {v['task']}: {v['message']}")
violations += 1
else:
print("[PASS] No model-divergence violations found")
print("\n=== Category 7: Loops ===")
violations += _audit_loops_block(args)
print(f"\n=== Summary ===")
@@ -1435,6 +1636,26 @@ def cmd_claim(args):
if implementer == args.agent:
print(f"ERROR: Agent '{args.agent}' implemented this task and cannot claim the code_review phase. Reviewer must be different from implementer.")
return 1
# Model-divergence check on claim (Subtask 2, multi-LLM only)
model_arg = getattr(args, "model", None)
if model_arg:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
if mode == "multi-llm":
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 2
state_models = _read_state_models(task_path)
conflict = _check_conflict(state_models, base, model_arg)
if conflict:
print(f"ERROR: Model '{model_arg}' for role '{base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model>.")
return 1
lock_file = task_path / ".state.lock"
timeout_sec = _lock_timeout_seconds(args.project)
if lock_file.exists():
@@ -2600,6 +2821,42 @@ def _gate_worktree_drift(state: dict, cfg: dict, project: Optional[str]) -> Opti
return None
def _gate_model_divergence(state: dict, cfg: dict, project: Optional[str]) -> Optional[dict]:
"""Model-divergence brake: in multi-LLM mode, verify and implement
roles must use different models. This prevents same-model verification
(rubber-stamping) within a loop tick."""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode != "multi-llm":
return None
roles = cfg.get("roles") or {}
impl_model = None
verify_model = None
impl_cfg = roles.get("implement") or {}
verify_cfg = roles.get("verify") or {}
impl_model = impl_cfg.get("model")
verify_model = verify_cfg.get("model")
# Fall back to manifest default if role has no explicit model
if not impl_model or not verify_model:
default = (manifest or {}).get("default")
if not impl_model:
impl_model = default
if not verify_model:
verify_model = default
if impl_model and verify_model and impl_model == verify_model:
return {
"ok": False,
"reason": "halted:model_conflict",
"halt_reason": "human_intervention",
"remaining_iterations": None,
"remaining_budget_usd": None,
"task_phase": None,
"task_in_halt_loop": True,
"out_of_scope_files": [],
}
return None
def _gate_score_plateau(state: dict, cfg: dict) -> Optional[dict]:
window = int(cfg.get("brakes", {}).get("score_plateau_window", 0))
if window <= 0:
@@ -2648,6 +2905,7 @@ def cmd_check_gate(args) -> int:
_gate_task_phase(state, cfg, args.project),
_gate_worktree_drift(state, cfg, args.project),
_gate_score_plateau(state, cfg),
_gate_model_divergence(state, cfg, args.project),
]
failure = next((g for g in gates if g is not None), None)
if failure is None:
@@ -2864,6 +3122,7 @@ def main():
parser.add_argument("--days", type=int, help="Cleanup age threshold in days (default 7, used with --cleanup-done / --install-cleanup-schedule)")
parser.add_argument("--dry-run", action="store_true", help="With --cleanup-done, list candidates without moving them")
parser.add_argument("--version", action="store_true", help="Print framework version and exit")
parser.add_argument("--model", metavar="NAME", help="Model name for model-divergence enforcement (used with --transition, --claim)")
args = parser.parse_args()

Some files were not shown because too many files have changed in this diff Show More