Compare commits

..
7 Commits
Author SHA1 Message Date
Lap Tran f13043d315 feat(pi): 3 new Pi Dev extensions — read-tweet, auto-context, task-lifecycle tools
CI / build (push) Has been cancelled
2026-06-26 22:41:38 -04:00
Lap Tran c2355954b9 feat(pi): pi-automaton.sh startup context printer for Pi Dev
Pi Dev doesn't auto-load automaton's system prompt (unlike opencode),
so the agent has no awareness of tasks, phases, or status.py commands.
This bridges that gap:

- scripts/pi-automaton.sh: NEW — detects scope (framework/project),
  reads active tasks, default model, config snippet, AGENTS.md rules,
  and prints a formatted context block the user pastes as their first
  message to the Pi Dev agent. Does NOT launch pi.
- scripts/register-guards.sh: EXTENDED — after Pi Dev guard install,
  offers to symlink pi-automaton.sh to ~/bin/pi-automaton
  (interactive prompt only when stdin is a terminal).
- README.md: added 'Using Pi Dev with automaton' FAQ subsection
  documenting the context printer workflow.
2026-06-26 20:47:08 -04:00
Lap Tran 0437cbae6c docs: architecture section, FAQ, cross-references, vault-memory update
README.md:
  - Added §1.5 'Architecture: Framework vs Project' with directory tree
    and lifecycle flow diagram
  - Added §2.5 FAQ covering coexistence, hooks, loops, multi-project
  - Updated install section already done in prior commit

AGENTS.md:
  - Added 'Script Cross-References' table mapping install.sh →
    onboard-project.sh → status.py --create-task

scripts/install-hooks.sh:
  - Updated header to reference onboard-project.sh as caller
  - Added 'Next step' line pointing to --create-task

vault-memory CONTEXT.md:
  - Updated test count (518→611)
  - Added new scripts (detect_models.py, onboard-project.sh)
  - Documented install/onboard flow and split architecture
2026-06-26 16:51:39 -04:00
Lap Tran b880f2535a fix(install): make install.sh idempotent + add onboard-project.sh
- install.sh: no longer exits early when ~/.automaton exists.
  Skips the clone but runs all setup (VRAM detection, guards,
  self-improvement loop, virtualenv). Both curl|bash and
  git-clone + ./install.sh now work correctly.
- onboard-project.sh: new script that bootstraps automaton in a
  project — creates .automaton/ skeleton, detects models, writes
  config.md, inits git, installs hooks, adds .gitignore entries.
- README.md: fix install flow docs (curl|bash + clone-then-run),
  add onboard-project.sh as Option A for project setup
2026-06-26 13:54:28 -04:00
Lap Tran f980ccfe27 feat(dashboard): model badges on kanban cards + design doc update
- task.py: Task dataclass gains models: dict[str, str], loaded from
  .state.models in discover_tasks()
- app.py: models dict included in all task API responses
- dashboard.js: model badges rendered between artifacts and subtask
  progress on kanban cards; ROLE_LABELS map for readable tooltips
- styles.css: .task-card-models and .model-badge styles
- design/loops/technical.md: document {model} substitution token,
  per-role model field, and model-divergence brake gate (gate #7)
2026-06-26 13:26:00 -04:00
Lap Tran 35e449b03e feat(model-divergence): full enforcement — manifest, transition, claim, audit, loop gates, detect script
Completes all 3 model-divergence enforcement subtasks:

- scripts/detect_models.py: probes opencode.json + localhost endpoints,
  builds models.json with --json/--write/--force
- scripts/status.py: CONFLICT_MATRIX, --model flag, --transition --model,
  --claim --model, --audit Category 6, model-divergence brake gate in
  --check-gate, helpers for manifest loading and conflict checking
- scripts/loop-runner.py: _role_model() helper + {model} passed via extras
  dict to _invoke_harness for implement, verify, orchestrate roles
- tests/test_model_divergence.py: 33 tests covering all enforcement layers
- Single-LLM mode: record model advisory, no conflict check
- Multi-LLM mode (2+ models): conflict matrix enforced at transition, claim,
  and loop brake gate
- Project-level models.json preferred over global ~/.automaton/models.json
2026-06-26 13:23:17 -04:00
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00
683 changed files with 2739 additions and 146 deletions
+1
View File
@@ -3,3 +3,4 @@ __pycache__/
*.egg-info/ *.egg-info/
.venv/ .venv/
venv/ venv/
logs/
+15
View File
@@ -136,6 +136,21 @@ Modes:
2. Run `python3 -m pytest tests/test_prompt_paths.py` to ensure task paths are canonical. 2. Run `python3 -m pytest tests/test_prompt_paths.py` to ensure task paths are canonical.
3. Update `CHANGELOG.md` under `[unreleased]`. 3. Update `CHANGELOG.md` under `[unreleased]`.
## Script Cross-References
The framework provides three lifecycle scripts that should be referenced from each other:
| Script | Purpose | Called when | Next step |
|---|---|---|---|
| `scripts/install.sh` | Install framework on a fresh machine | `curl \| bash` or `git clone + bash` | → `scripts/onboard-project.sh` |
| `scripts/onboard-project.sh` | Bootstrap automaton in a project | After framework install, per project | → `status.py --create-task` |
| `scripts/install-hooks.sh` | Install git hooks per project | After onboarding, or manually | see `onboard-project.sh` |
- `install.sh` outputs "Next: onboard-project.sh" at the end.
- `onboard-project.sh` outputs "Next: status.py --create-task" at the end.
- `install-hooks.sh` is called by `onboard-project.sh` automatically.
- `update.sh` does NOT call `onboard-project.sh` — it only updates the framework.
## Adding a New Script ## Adding a New Script
1. Place the script in `scripts/`. 1. Place the script in `scripts/`.
+21
View File
@@ -2,6 +2,27 @@
## [unreleased] ## [unreleased]
### Fixed — dashboard scroll-reset on auto-refresh
- **`automaton/dashboard/html/dashboard.js`** (`renderBoard`): auto-refresh rebuilt the board via `board.innerHTML = html` every tick (default 2s), destroying each `.column-body`'s `scrollTop` and snapping it back to 0 — so users couldn't scroll the Done group down to review older tasks. Now snapshots each column-body's `scrollTop` (plus the board's `scrollLeft` and the active view's `scrollTop`) before the rebuild and restores them after, matched by index (PHASE_GROUPS order is stable).
- **New tests**: `tests/test_dashboard_ui.py` — Playwright browser smoke test (board renders tasks; column scroll survives an auto-refresh tick). Skipped via `importorskip` when playwright/chromium is absent so CI without a browser stays green. Verified the test fails without the fix (scrollTop resets to 0) and passes with it.
### Fixed — failing plist-isolation test (host bleed false positive)
- **`tests/test_cleanup_done.py`** (`TestInstallCleanupScheduleIsolation.test_plist_written_to_override_dir_not_host`): asserted `not host.exists()`, but the host `~/Library/LaunchAgents/com.automaton.cleanup.plist` legitimately exists from a real `--install-cleanup-schedule` run, causing a false failure. Now snapshots the host plist's `st_mtime_ns` (or absence) before the test run and asserts it's unchanged after — a pre-existing real install no longer fails the test; only an actual write during the run would.
### Changed — bind ornith as the Implement model
- **`config.md`** (Model Configuration): set `Model: omlx/Ornith-1.0-35B-4bit-mlx` and `Override context window: 32768` (matches the opencode.json limit for the local LLM). Interactive autopilot already used ornith via opencode's default model; this makes it explicit so auto-detection can't pick another model. Loop ticks still use the single `harness.command` for all roles — per-role model binding (`{model}` substitution in loop-runner.py) is **not** implemented yet (see model-divergence gap below).
### Changed — README: document the self-improvement loop's scope for new projects
- **`README.md`** (Project Setup): added "The Self-Improvement Loop is framework-scoped" note — the default SI loop targets `~/.automaton/` (the framework), not your project, by design. Documents the leave-running / pause / create-a-project-loop paths.
### Known gap — model-divergence was marked complete but unimplemented
- The `model-divergence-enforcement` parent task and its 3 subtasks (`mde-manifest-detection`, `mde-interactive-enforcement`, `mde-loop-enforcement`) are `.state = complete` but contain only `SPEC.md`/`DECOMPOSITION.md` — no `IMPLEMENTATION.md`, no `VERDICT.md`. The promised code (`status.py` model_divergence audit category, `--transition --model`, `.state.models`, `loop.json` per-role `model` + `{model}` substitution in loop-runner.py, dashboard badges) was never written. Consequence: loop roles (implement/verify/orchestrate) all run the same model, so the D12 conflict-of-interest rule (Verify ≠ Implement session/model) is unenforced. Interactive autopilot is unaffected.
### Added — framework agent features design docs ### Added — framework agent features design docs
- **New `design/framework/`** directory: design index, functional design, technical design, and backlog for three framework-level agent features: - **New `design/framework/`** directory: design index, functional design, technical design, and backlog for three framework-level agent features:
+184 -20
View File
@@ -6,22 +6,28 @@ A contract-based operating system for LLM agents, designed to enforce discipline
Before you can use the framework in any project, you must install the core logic into your local environment. Before you can use the framework in any project, you must install the core logic into your local environment.
**Run these commands in your terminal:** Choose one of the following methods:
### Option A: One-liner (curl pipe, recommended)
```bash ```bash
# Clone the framework into the global config directory curl -fsSL https://raw.githubusercontent.com/<your-org>/automaton/main/scripts/install.sh | bash -s -- <your-git-url>
git clone <your-git-url> ~/.automaton
# Enter the directory
cd ~/.automaton
# Make the installation script executable and run it
# You must provide the git URL as the first argument
chmod +x install.sh
./install.sh <your-git-url>
``` ```
The git URL is required because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling. This clones the framework to `~/.automaton/`, runs VRAM detection, installs pre-edit guards, creates the self-improvement loop, and sets up the Python virtualenv.
### Option B: Clone first
```bash
git clone <your-git-url> ~/.automaton
bash ~/.automaton/scripts/install.sh
```
The script detects that `~/.automaton` already exists, skips the clone, and runs all setup steps (VRAM detection, guards, loop, virtualenv).
### Both methods do the same thing
The git URL is required on fresh install because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling.
*Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory. A self-improvement loop is created and scheduled by default (see Loop Engineering below).* *Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory. A self-improvement loop is created and scheduled by default (see Loop Engineering below).*
@@ -45,11 +51,95 @@ If a project was set up under the old model (with copies of framework files), it
--- ---
## 1.5 Architecture: Framework vs Project
Automaton uses a **split architecture** — one shared framework, many project `./.automaton/` directories:
```
~/.automaton/ ← Framework (installed once per machine)
├── scripts/ ← shared tooling: status.py, loop-runner.py
├── prompts/ ← shared LLM prompts
├── plugins/ ← shared harness plugins
├── templates/ ← shared task & loop templates
├── .automaton/tasks/ ← framework housekeeping tasks (self-improvement)
└── .automaton/loops/ ← framework loops (self-improvement loop)
~/projects/my-app/
└── .automaton/ ← Project (onboarded once per project)
├── tasks/ ← YOUR project's tasks
├── models.json ← YOUR project's model config
├── config.md ← YOUR project's VRAM config
├── project-name.md ← YOUR project's display name
└── loops/ ← YOUR project's loops
```
**Key rules:**
- The agent is **scope-aware**: if you're inside `~/.automaton/`, it operates in **framework mode** (reads framework tasks). If you're inside a project dir, it operates in **project mode** (reads project tasks). They never interfere.
- All framework scripts (`status.py`, etc.) live in `~/.automaton/scripts/` and are shared — never copied into projects.
- Framework prompts live in `~/.automaton/prompts/` — projects reference them by path at runtime.
- Git hooks are **per-project**. Each project installs its own via `bash ~/.automaton/scripts/install-hooks.sh <project-path>`.
- The self-improvement loop targets **only** the framework itself. Your project won't get framework-level tasks in its board.
- You can work on both at the same time in different terminals — independent `.automaton/` directories, shared tooling.
### Lifecycle overview
```
┌─────────────────────────────────────────────────────┐
│ 1. Install Framework (once per machine) │
│ curl .../install.sh | bash -s -- <git-url> │
│ → clones to ~/.automaton/ │
│ → VRAM detection, guards, venv, self-improvement │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 2. Onboard Project (once per project) │
│ bash ~/.automaton/scripts/onboard-project.sh <dir> │
│ → creates .automaton/ skeleton │
│ → probes models, writes config.md │
│ → git init + hooks │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 3. Create Task (per feature) │
│ python3 ~/.automaton/scripts/status.py │
│ --create-task my-feature --project . │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 4. Work Through Phases (per task) │
│ status.py --transition research --task my-feature │
│ → agent writes SPEC.md │
│ status.py --transition implement --task my-feature │
│ → agent writes code + IMPLEMENTATION.md │
│ ... → complete │
└─────────────────────────────────────────────────────┘
```
---
## 2. Project Setup (Per project) ## 2. Project Setup (Per project)
Once the framework is installed globally, you must "onboard" every individual project you work on. Once the framework is installed globally, you must "onboard" every individual project you work on.
### Option A: The Agent-Driven Way (Recommended) ### Option A: The Onboarding Script (Recommended)
```bash
bash ~/.automaton/scripts/onboard-project.sh /path/to/project
```
This will:
1. Create `.automaton/` skeleton if missing.
2. Run `detect_models.py --write` to probe local models (falls back to a minimal `models.json`).
3. Generate `config.md` with VRAM recommendations.
4. Write `project-name.md` from the directory name.
5. Initialize git if not already a repo.
6. Install git hooks (pre-commit + pre-push).
7. Add automaton entries to `.gitignore`.
8. Run `status.py --audit` to verify the setup.
### Option B: The Agent-Driven Way
If you want the agent to handle the configuration for you, navigate to your project root and run: If you want the agent to handle the configuration for you, navigate to your project root and run:
> *"Onboard this project into automaton." > *"Onboard this project into automaton."
@@ -60,17 +150,91 @@ The agent will automatically:
3. Generate your `.agent.md` and `.rules.md` files. 3. Generate your `.agent.md` and `.rules.md` files.
4. Initiate the "Exploration Ritual" to understand your codebase. 4. Initiate the "Exploration Ritual" to understand your codebase.
### Option B: The Manual Way ### Option C: The Manual Way
If you prefer to set it up manually, create a `.automaton/` directory in your project root and add: If you prefer to set it up manually:
- `.agent.md`: Project-specific configuration (Mode, rules, etc.).
- `.rules.md`: Project-specific constraints and past failure modes.
Then install the git pre-commit hook:
```bash ```bash
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit # Create the automaton directory
mkdir -p .automaton/tasks .automaton/loops .automaton/design
# Install git hooks
bash ~/.automaton/scripts/install-hooks.sh .
# Configure your model(s)
python3 ~/.automaton/scripts/detect_models.py --write --project .
``` ```
This hook blocks commits when no task is in `implement` or `doc_review` phase, preventing accidental code changes without a tracked task. Then add `.agent.md` and `.rules.md` for the agent.
### The Self-Improvement Loop is framework-scoped
The self-improvement loop installed by default targets the **framework itself** (`work_source.project: ~/.automaton/`), not your project. It ticks hourly against `status.py --audit` on `~/.automaton/` to keep the framework healthy. This is by design — it improves the framework you depend on while you work on your project.
- **Leave it running** if you want the framework maintained in the background (recommended).
- **Disable it** if you want zero background activity: `python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/`
- **Want a loop on your project too?** Create a separate one targeted at the project root:
```bash
python3 ~/.automaton/scripts/status.py --create-loop my-project-loop \
--from-template self-improvement --project /path/to/project
```
---
## 2.5 FAQ
### Can I work on the framework and a project at the same time?
Yes. They have separate `.automaton/` directories. Open two terminals:
```
Terminal 1: cd ~/.automaton → framework mode
Terminal 2: cd ~/projects/my-app → project mode
```
The agent detects scope from your current directory. Each can have its own tasks, loops, and config. They share the same `~/.automaton/scripts/` binaries.
### Why doesn't `install.sh` need a Git URL when run from the repo?
Because the framework is already cloned. `install.sh` skips cloning when `~/.automaton/` exists and runs all the setup steps (VRAM detection, pip deps, self-improvement loop, guards). The Git URL is only required for a fresh install via `curl | bash`.
### Do I need to run `install.sh` again after pulling updates?
No. `git pull` inside `~/.automaton/` updates the code. The self-improvement loop and guards persist across updates. If you want to re-register guards (e.g. after switching harnesses), run `bash ~/.automaton/scripts/register-guards.sh`.
### How do git hooks work per project?
Each project installs its own hooks via:
```bash
bash ~/.automaton/scripts/install-hooks.sh /path/to/project
```
The pre-commit hook blocks commits when no task is in `implement` or `doc_review` phase. The pre-push hook catches `--no-verify` bypasses. They're independent per repo.
### Can I have multiple projects onboarded at once?
Yes. Each project gets its own `.automaton/` directory. Run `onboard-project.sh` once per project. The shared scripts in `~/.automaton/scripts/` enforce the state machine on whichever project you point `--project` at.
### What about loops on my project?
The self-improvement loop runs only on the framework. To add a loop to your project:
```bash
python3 ~/.automaton/scripts/status.py --create-loop my-loop \
--from-template self-improvement --project /path/to/project
python3 ~/.automaton/scripts/status.py --install-schedule my-loop \
--interval 3600 --project /path/to/project
```
### Using Pi Dev with automaton
Pi Dev has the `automaton-guard-pi` plugin (installed by `register-guards.sh`) which blocks edits outside allowed phases. However, Pi Dev does **not** auto-load automaton's system prompt (unlike opencode). For the agent to understand tasks and phases, provide context manually.
**Before starting a Pi Dev session**, run the context printer:
```bash
bash ~/.automaton/scripts/pi-automaton.sh
```
Or if symlinked to `~/bin/`:
```bash
pi-automaton
```
Copy the output and paste it as your first message to the Pi Dev agent. This tells the agent about:
- The automaton workflow framework and phase lifecycle
- Active tasks in the current project
- The default model and VRAM configuration
- Project-specific rules from `AGENTS.md`
- Which commands to use for transitions and task creation
The `automaton-guard-pi` plugin still blocks edits outside `implement`/`doc_review` even without this context — the context printer just makes the agent *aware* of why it's being blocked and how to use the framework correctly.
--- ---
+42 -22
View File
@@ -115,6 +115,7 @@ class Task:
name: str name: str
folder_path: Path folder_path: Path
state: TaskState state: TaskState
phase_raw: str = ""
artifacts: dict[str, ArtifactStatus] = field(default_factory=dict) artifacts: dict[str, ArtifactStatus] = field(default_factory=dict)
sub_tasks: list[SubTask] = field(default_factory=list) sub_tasks: list[SubTask] = field(default_factory=list)
parent_spec: Optional[str] = None parent_spec: Optional[str] = None
@@ -125,10 +126,14 @@ class Task:
design_content: Optional[str] = None design_content: Optional[str] = None
spec_content: Optional[str] = None spec_content: Optional[str] = None
decomposition_content: Optional[str] = None decomposition_content: Optional[str] = None
code_review_content: Optional[str] = None
test_plan_content: Optional[str] = None
implementation_content: Optional[str] = None
parent_spec_content: Optional[str] = None parent_spec_content: Optional[str] = None
vram_config_content: Optional[str] = None vram_config_content: Optional[str] = None
waves: list[WaveGroup] = field(default_factory=list) waves: list[WaveGroup] = field(default_factory=list)
is_corrupted: bool = False is_corrupted: bool = False
models: dict[str, str] = field(default_factory=dict)
@property @property
def display_name(self) -> str: def display_name(self) -> str:
@@ -297,7 +302,7 @@ class Task:
@property @property
def is_approval_gated(self) -> bool: def is_approval_gated(self) -> bool:
return self.state in (TaskState.RESEARCH, TaskState.DECOMPOSITION, TaskState.DESIGN, TaskState.TEST_DESIGN, TaskState.CODE_REVIEW) return self.phase_raw.endswith(":awaiting_approval")
@property @property
def blocker(self) -> str: def blocker(self) -> str:
@@ -310,7 +315,7 @@ class Task:
if not artifact.content: if not artifact.content:
return f"Empty required artifact: {required}" return f"Empty required artifact: {required}"
if self.is_approval_gated: if self.is_approval_gated:
return "Awaiting user approval — use `status.py --approve` to approve" return "Awaiting user approval — click Approve Phase in the dashboard"
if self.state == TaskState.BUG_FIND: if self.state == TaskState.BUG_FIND:
if "BUG_REPORT.md" not in self.artifacts or not self.artifacts["BUG_REPORT.md"].content: if "BUG_REPORT.md" not in self.artifacts or not self.artifacts["BUG_REPORT.md"].content:
return "Agent must generate BUG_REPORT.md" return "Agent must generate BUG_REPORT.md"
@@ -484,7 +489,7 @@ def _state_string_to_task_state(phase: str) -> TaskState:
return mapping.get(base, TaskState.BACKLOG) return mapping.get(base, TaskState.BACKLOG)
def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, ArtifactStatus]]: def determine_task_state(folder_path: Path) -> tuple[TaskState, str, dict[str, ArtifactStatus]]:
artifacts = {} artifacts = {}
for filename in ARTIFACTS: for filename in ARTIFACTS:
@@ -511,7 +516,7 @@ def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, Artifa
try: try:
phase = state_file.read_text(encoding="utf-8").strip() phase = state_file.read_text(encoding="utf-8").strip()
if phase: if phase:
return _state_string_to_task_state(phase), artifacts return _state_string_to_task_state(phase), phase, artifacts
except (OSError, IOError): except (OSError, IOError):
pass # Fall through to artifact heuristic pass # Fall through to artifact heuristic
@@ -521,40 +526,40 @@ def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, Artifa
if "VERDICT.md" in artifacts: if "VERDICT.md" in artifacts:
verdict_content = artifacts["VERDICT.md"].content verdict_content = artifacts["VERDICT.md"].content
if not verdict_content: if not verdict_content:
return TaskState.BLOCKED, artifacts return TaskState.BLOCKED, "blocked", artifacts
verdict_status = parse_verdict_status(verdict_content) verdict_status = parse_verdict_status(verdict_content)
if verdict_status == VERDICT_FAIL or verdict_status == VERDICT_NEEDS_REVIEW: if verdict_status == VERDICT_FAIL or verdict_status == VERDICT_NEEDS_REVIEW:
return TaskState.BLOCKED, artifacts return TaskState.BLOCKED, "blocked", artifacts
if verdict_status == VERDICT_PASS: if verdict_status == VERDICT_PASS:
return TaskState.DONE, artifacts return TaskState.DONE, "done", artifacts
# Verdict exists but status is unparseable — pending referee review # Verdict exists but status is unparseable — pending referee review
if verdict_status is None: if verdict_status is None:
return TaskState.REFEREE, artifacts return TaskState.REFEREE, "referee", artifacts
# State machine aligned with orchestrate.md # State machine aligned with orchestrate.md
# Check from most advanced to least advanced # Check from most advanced to least advanced
if "DOC_REVIEW.md" in artifacts: if "DOC_REVIEW.md" in artifacts:
return TaskState.DOC_REVIEW, artifacts return TaskState.DOC_REVIEW, "doc_review", artifacts
if "ADVERSARIAL_BUG_REPORT.md" in artifacts and "BUG_REPORT.md" in artifacts: if "ADVERSARIAL_BUG_REPORT.md" in artifacts and "BUG_REPORT.md" in artifacts:
return TaskState.ADV_BUG_FIND, artifacts return TaskState.ADV_BUG_FIND, "adv_bug_find", artifacts
if "BUG_REPORT.md" in artifacts: if "BUG_REPORT.md" in artifacts:
return TaskState.BUG_FIND, artifacts return TaskState.BUG_FIND, "bug_find", artifacts
if "ADVERSARIAL_BUG_REPORT.md" in artifacts: if "ADVERSARIAL_BUG_REPORT.md" in artifacts:
return TaskState.BUG_FIND, artifacts return TaskState.BUG_FIND, "bug_find", artifacts
if "CODE_REVIEW.md" in artifacts: if "CODE_REVIEW.md" in artifacts:
return TaskState.CODE_REVIEW, artifacts return TaskState.CODE_REVIEW, "code_review", artifacts
if "IMPLEMENTATION.md" in artifacts: if "IMPLEMENTATION.md" in artifacts:
return TaskState.IMPLEMENT, artifacts return TaskState.IMPLEMENT, "implement", artifacts
if "TEST_PLAN.md" in artifacts: if "TEST_PLAN.md" in artifacts:
return TaskState.TEST_DESIGN, artifacts return TaskState.TEST_DESIGN, "test_design", artifacts
if "DESIGN.md" in artifacts: if "DESIGN.md" in artifacts:
return TaskState.DESIGN, artifacts return TaskState.DESIGN, "design", artifacts
if "DECOMPOSITION.md" in artifacts: if "DECOMPOSITION.md" in artifacts:
return TaskState.DECOMPOSITION, artifacts return TaskState.DECOMPOSITION, "decomposition", artifacts
if "SPEC.md" in artifacts: if "SPEC.md" in artifacts:
return TaskState.RESEARCH, artifacts return TaskState.RESEARCH, "research", artifacts
return TaskState.BACKLOG, artifacts return TaskState.BACKLOG, "backlog", artifacts
def parse_sub_tasks(parent_folder: Path) -> list[SubTask]: def parse_sub_tasks(parent_folder: Path) -> list[SubTask]:
@@ -569,7 +574,7 @@ def parse_sub_tasks(parent_folder: Path) -> list[SubTask]:
name = subtask_folder.name name = subtask_folder.name
if not _VALID_TASK_NAME_CHARS.issuperset(set(name)): if not _VALID_TASK_NAME_CHARS.issuperset(set(name)):
continue continue
state, artifacts = determine_task_state(subtask_folder) state, _, artifacts = determine_task_state(subtask_folder)
verdict_status = None verdict_status = None
has_verdict = False has_verdict = False
if "VERDICT.md" in artifacts and artifacts["VERDICT.md"].content: if "VERDICT.md" in artifacts and artifacts["VERDICT.md"].content:
@@ -653,12 +658,12 @@ def discover_tasks(tasks_dir: Path) -> list[Task]:
if not _VALID_TASK_NAME_CHARS.issuperset(set(folder_path.name)): if not _VALID_TASK_NAME_CHARS.issuperset(set(folder_path.name)):
continue continue
state, artifacts = determine_task_state(folder_path) state, phase_raw, artifacts = determine_task_state(folder_path)
sub_tasks = parse_sub_tasks(folder_path) sub_tasks = parse_sub_tasks(folder_path)
task = Task( task = Task(
name=folder_path.name, folder_path=folder_path, state=state, name=folder_path.name, folder_path=folder_path, state=state,
artifacts=artifacts, sub_tasks=sub_tasks, phase_raw=phase_raw, artifacts=artifacts, sub_tasks=sub_tasks,
) )
# Load specific artifact contents for detail panel # Load specific artifact contents for detail panel
@@ -677,9 +682,24 @@ def discover_tasks(tasks_dir: Path) -> list[Task]:
if "DECOMPOSITION.md" in artifacts and artifacts["DECOMPOSITION.md"].content: if "DECOMPOSITION.md" in artifacts and artifacts["DECOMPOSITION.md"].content:
task.decomposition_content = artifacts["DECOMPOSITION.md"].content task.decomposition_content = artifacts["DECOMPOSITION.md"].content
task.waves = parse_waves(task.decomposition_content) task.waves = parse_waves(task.decomposition_content)
if "CODE_REVIEW.md" in artifacts and artifacts["CODE_REVIEW.md"].content:
task.code_review_content = artifacts["CODE_REVIEW.md"].content
if "TEST_PLAN.md" in artifacts and artifacts["TEST_PLAN.md"].content:
task.test_plan_content = artifacts["TEST_PLAN.md"].content
if "IMPLEMENTATION.md" in artifacts and artifacts["IMPLEMENTATION.md"].content:
task.implementation_content = artifacts["IMPLEMENTATION.md"].content
task.parent_spec_content = parse_parent_spec(folder_path) task.parent_spec_content = parse_parent_spec(folder_path)
task.vram_config_content = parse_vram_config(folder_path) task.vram_config_content = parse_vram_config(folder_path)
# Load .state.models for model divergence badges
state_models_path = folder_path / ".state.models"
if state_models_path.exists():
try:
import json as _json
task.models = _json.loads(state_models_path.read_text(encoding="utf-8"))
except (OSError, IOError, _json.JSONDecodeError):
task.models = {}
tasks.append(task) tasks.append(task)
# Sort by state (most advanced first) # Sort by state (most advanced first)
+143 -4
View File
@@ -110,6 +110,13 @@ async function fetchProjectName() {
function renderBoard() { function renderBoard() {
const board = document.getElementById('board'); const board = document.getElementById('board');
// Preserve scroll positions across re-renders so auto-refresh doesn't
// snap columns back to the top while the user is reviewing older tasks.
const prevBodies = Array.from(board.querySelectorAll('.column-body'));
const savedScrolls = prevBodies.map(el => el.scrollTop);
const savedBoardScrollLeft = board.scrollLeft;
const view = document.querySelector('.view.active');
const savedViewScrollTop = view ? view.scrollTop : 0;
const filtered = getFilteredTasks(); const filtered = getFilteredTasks();
const groups = {}; const groups = {};
PHASE_GROUPS.forEach(group => { groups[group.id] = []; }); PHASE_GROUPS.forEach(group => { groups[group.id] = []; });
@@ -140,9 +147,25 @@ function renderBoard() {
</div>`; </div>`;
}).join(''); }).join('');
board.innerHTML = html; board.innerHTML = html;
// Restore scroll positions (matched by index — PHASE_GROUPS order is stable).
const newBodies = board.querySelectorAll('.column-body');
newBodies.forEach((el, i) => { if (savedScrolls[i] != null) el.scrollTop = savedScrolls[i]; });
board.scrollLeft = savedBoardScrollLeft;
if (view) view.scrollTop = savedViewScrollTop;
attachCardListeners(); attachCardListeners();
} }
const ROLE_LABELS = {
'implement': 'Implement',
'code_review': 'Code Review',
'bug_find': 'Bug Find',
'adversarial_bug_find': 'Adv Bug Find',
'doc_review': 'Doc Review',
'referee': 'Referee',
'loop-implement': 'Loop Impl',
'loop-verify': 'Loop Verify',
};
const ARTIFACT_LABELS = { const ARTIFACT_LABELS = {
'research': 'SPEC.md', 'decomposition': 'DECOMPOSITION.md', 'research': 'SPEC.md', 'decomposition': 'DECOMPOSITION.md',
'design': 'DESIGN.md', 'test_design': 'TEST_PLAN.md', 'design': 'DESIGN.md', 'test_design': 'TEST_PLAN.md',
@@ -170,11 +193,20 @@ function renderTaskCard(task) {
const label = ARTIFACT_LABELS[col.id] || col.label; const label = ARTIFACT_LABELS[col.id] || col.label;
return `<span class="artifact-badge" title="${col.label}">${label}</span>`; return `<span class="artifact-badge" title="${col.label}">${label}</span>`;
}).join('')}</div>`; }).join('')}</div>`;
const modelKeys = Object.keys(task.models || {});
const modelsHtml = modelKeys.length > 0
? `<div class="task-card-models">${modelKeys.map(role => {
const m = task.models[role];
const roleLabel = ROLE_LABELS[role] || role;
return `<span class="model-badge" title="${roleLabel}: ${m}">${m}</span>`;
}).join('')}</div>`
: '';
return `<div class="task-card" data-task="${task.name}" data-status="${statusClass}"> return `<div class="task-card" data-task="${task.name}" data-status="${statusClass}">
<div class="task-card-header"><span class="task-card-name">${escapeHtml(task.display_name)}</span><span class="task-card-status ${statusClass}">${statusIcon}</span></div> <div class="task-card-header"><span class="task-card-name">${escapeHtml(task.display_name)}</span><span class="task-card-status ${statusClass}">${statusIcon}</span></div>
<div class="task-card-sublabel">${subLabel}</div> <div class="task-card-sublabel">${subLabel}</div>
${task.status_reason ? `<div class="task-card-reason">${escapeHtml(task.status_reason)}</div>` : ''} ${task.status_reason ? `<div class="task-card-reason">${escapeHtml(task.status_reason)}</div>` : ''}
${artifactsHtml} ${artifactsHtml}
${modelsHtml}
${progressHtml ? `<div class="task-card-footer"><span class="subtask-progress">${progressHtml}</span></div>` : ''} ${progressHtml ? `<div class="task-card-footer"><span class="subtask-progress">${progressHtml}</span></div>` : ''}
${subtasksHtml} ${subtasksHtml}
</div>`; </div>`;
@@ -198,13 +230,69 @@ function renderDetail(task) {
return `<span class="detail-artifact"><span class="${cls}">${icon}</span>${col.label}</span>`; return `<span class="detail-artifact"><span class="${cls}">${icon}</span>${col.label}</span>`;
}).join(''); }).join('');
const phaseGroupHtml = phaseGroup ? `<span class="detail-phase-badge" style="background: ${phaseGroupColor}20; color: ${phaseGroupColor}">${PHASE_GROUPS.find(g => g.id === phaseGroup).label}</span>` : ''; const phaseGroupHtml = phaseGroup ? `<span class="detail-phase-badge" style="background: ${phaseGroupColor}20; color: ${phaseGroupColor}">${PHASE_GROUPS.find(g => g.id === phaseGroup).label}</span>` : '';
const approvePhaseBtn = task.is_approval_gated
? `<button class="approve-phase-btn" data-task="${escapeHtml(task.name)}">🔓 Approve Phase</button>` // Build approval section for approval-gated phases
const APPROVAL_ARTIFACT_MAP = {
'research': { contentKey: 'spec_content', label: 'SPEC.md', phaseName: 'Research' },
'decomposition': { contentKey: 'decomposition_content', label: 'DECOMPOSITION.md', phaseName: 'Decomposition' },
'design': { contentKey: 'design_content', label: 'DESIGN.md', phaseName: 'Design' },
'test_design': { contentKey: 'test_plan_content', label: 'TEST_PLAN.md', phaseName: 'Test Design' },
'code_review': { contentKey: 'code_review_content', label: 'CODE_REVIEW.md', phaseName: 'Code Review' },
};
const awaitingPhase = task.phase_raw ? task.phase_raw.replace(':awaiting_approval', '') : null;
const approvalInfo = awaitingPhase ? APPROVAL_ARTIFACT_MAP[awaitingPhase] : null;
const approvalContent = approvalInfo && approvalInfo.contentKey ? task[approvalInfo.contentKey] : null;
const approvalHtml = task.is_approval_gated
? `<div class="approval-section">
<div class="approval-header">
<span class="approval-icon">🔒</span>
<span class="approval-title">${approvalInfo ? approvalInfo.phaseName : 'Phase'} — Needs Approval</span>
<button class="approve-phase-btn" data-task="${escapeHtml(task.name)}">Approve Phase</button>
</div>
${task.blocker ? `<p class="approval-blocker">${escapeHtml(task.blocker)}</p>` : ''}
${approvalContent ? `<details class="approval-artifact" open>
<summary>${approvalInfo.label} — review content before approving</summary>
<pre class="detail-content-text">${escapeHtml(approvalContent)}</pre>
</details>` : `<p class="approval-missing">${approvalInfo ? approvalInfo.label : 'Artifact'} not yet written — an agent must create it before this phase can complete.</p>`}
</div>`
: ''; : '';
// Build transition section for phases that can advance
const TRANSITION_MAP = {
'research:approved': { target: 'decomposition', label: 'Advance to Decomposition' },
'decomposition:approved': { target: 'design', label: 'Advance to Design' },
'design:approved': { target: 'implement', label: 'Advance to Implementation' },
'test_design:approved': { target: 'implement', label: 'Advance to Implementation' },
'implement': { target: 'code_review', label: 'Advance to Code Review' },
'code_review:approved': { target: 'bug_find', label: 'Advance to Bug Finding' },
'bug_find': { target: 'adv_bug_find', label: 'Advance to Adversarial Bug Finding' },
'adv_bug_find': { target: 'doc_review', label: 'Advance to Document Review' },
'doc_review': { target: 'referee', label: 'Advance to Referee' },
};
const nextTransition = TRANSITION_MAP[task.phase_raw];
const transitionHtml = nextTransition
? `<div class="detail-section detail-actions"><h4>🚀 Phase Actions</h4><button class="transition-btn" data-task="${escapeHtml(task.name)}" data-target="${nextTransition.target}">${nextTransition.label}</button></div>`
: '';
// Build artifact editor for missing required artifacts
const requiredArtifact = task.required_artifact_name;
const hasRequiredArtifact = task.artifacts[requiredArtifact];
const artifactEditorHtml = requiredArtifact && !hasRequiredArtifact && !task.is_approval_gated
? `<div class="detail-section detail-artifact-editor">
<h4>📝 Write ${requiredArtifact}</h4>
<p class="artifact-editor-hint">This artifact is required before the task can advance. Write it below and save.</p>
<textarea class="artifact-editor" data-task="${escapeHtml(task.name)}" data-filename="${requiredArtifact}" rows="12" placeholder="# ${requiredArtifact.replace('.md','')}\n\nWrite content here..."></textarea>
<button class="save-artifact-btn" data-task="${escapeHtml(task.name)}" data-filename="${requiredArtifact}">Save ${requiredArtifact}</button>
</div>`
: '';
content.innerHTML = ` content.innerHTML = `
<div class="detail-section"><h4>Status</h4><span class="detail-status-badge ${statusClass}">${statusText}</span>${phaseGroupHtml}${task.is_edit_phase ? '<span class="detail-edit-badge">✏️ Edit Allowed</span>' : ''}${task.is_approval_gated ? '<span class="detail-approval-badge">🔒 Requires Approval</span>' : ''}${approvePhaseBtn}</div> <div class="detail-section"><h4>Status</h4><span class="detail-status-badge ${statusClass}">${statusText}</span>${phaseGroupHtml}${task.is_edit_phase ? '<span class="detail-edit-badge">✏️ Edit Allowed</span>' : ''}</div>
${statusReason ? `<div class="detail-section detail-status-reason">${escapeHtml(statusReason)}</div>` : ''} ${statusReason ? `<div class="detail-section detail-status-reason">${escapeHtml(statusReason)}</div>` : ''}
${task.blocker ? `<div class="detail-section detail-blocker"><h4>⚠️ What's Blocking</h4><p class="detail-blocker-text">${escapeHtml(task.blocker)}</p></div>` : ''} ${approvalHtml}
${task.blocker && !task.is_approval_gated ? `<div class="detail-section detail-blocker"><h4>⚠️ What's Blocking</h4><p class="detail-blocker-text">${escapeHtml(task.blocker)}</p></div>` : ''}
${artifactEditorHtml}
${transitionHtml}
${task.phase_guidance ? `<div class="detail-section detail-guidance"><h4>▶ What's Next</h4><pre class="detail-guidance-text">${escapeHtml(task.phase_guidance)}</pre></div>` : ''} ${task.phase_guidance ? `<div class="detail-section detail-guidance"><h4>▶ What's Next</h4><pre class="detail-guidance-text">${escapeHtml(task.phase_guidance)}</pre></div>` : ''}
${task.state === 'blocked' && task.blocked_action_items && task.blocked_action_items.length > 0 ? `<div class="detail-section detail-action-items"><h4>📋 Action Items</h4><ul class="detail-action-list">${task.blocked_action_items.map(item => `<li>${escapeHtml(item)}</li>`).join('')}</ul></div>` : ''} ${task.state === 'blocked' && task.blocked_action_items && task.blocked_action_items.length > 0 ? `<div class="detail-section detail-action-items"><h4>📋 Action Items</h4><ul class="detail-action-list">${task.blocked_action_items.map(item => `<li>${escapeHtml(item)}</li>`).join('')}</ul></div>` : ''}
<div class="detail-section"><h4>Artifacts</h4><div class="detail-artifacts">${artifactsHtml}</div></div> <div class="detail-section"><h4>Artifacts</h4><div class="detail-artifacts">${artifactsHtml}</div></div>
@@ -504,6 +592,44 @@ async function approvePhase(taskName) {
} }
} }
async function transitionTask(taskName, target) {
try {
const res = await fetch(`/api/transition/${taskName}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ target }),
});
const data = await res.json();
if (data.success) {
closeDetail();
await refreshData();
} else {
console.error('Transition failed:', data);
}
} catch (err) {
console.error('Transition failed:', err);
}
}
async function saveArtifact(taskName, filename, content) {
try {
const res = await fetch(`/api/write-artifact/${taskName}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ filename, content }),
});
const data = await res.json();
if (data.success) {
closeDetail();
await refreshData();
} else {
console.error('Save artifact failed:', data);
}
} catch (err) {
console.error('Save artifact failed:', err);
}
}
function getFilteredTasks() { function getFilteredTasks() {
let filtered = [...state.tasks]; let filtered = [...state.tasks];
if (state.filterPhase !== 'all') filtered = filtered.filter(t => t.state === state.filterPhase); if (state.filterPhase !== 'all') filtered = filtered.filter(t => t.state === state.filterPhase);
@@ -593,6 +719,19 @@ function setupUI() {
const taskName = approveBtn.dataset.task; const taskName = approveBtn.dataset.task;
if (taskName) approvePhase(taskName); if (taskName) approvePhase(taskName);
} }
const transitionBtn = e.target.closest('.transition-btn');
if (transitionBtn) {
const taskName = transitionBtn.dataset.task;
const target = transitionBtn.dataset.target;
if (taskName && target) transitionTask(taskName, target);
}
const saveBtn = e.target.closest('.save-artifact-btn');
if (saveBtn) {
const taskName = saveBtn.dataset.task;
const filename = saveBtn.dataset.filename;
const editor = document.querySelector(`.artifact-editor[data-task="${taskName}"][data-filename="${filename}"]`);
if (taskName && filename && editor) saveArtifact(taskName, filename, editor.value);
}
}); });
} }
+40
View File
@@ -307,8 +307,48 @@ kbd {
::-webkit-scrollbar-thumb:hover { background: var(--border-active); } ::-webkit-scrollbar-thumb:hover { background: var(--border-active); }
.approve-phase-btn { display: inline-block; margin-left: 8px; padding: 4px 12px; border: 1px solid var(--warning); border-radius: var(--radius-sm); cursor: pointer; font-size: 11px; font-weight: 500; background: var(--warning-bg); color: var(--warning); font-family: inherit; transition: all 0.15s; } .approve-phase-btn { display: inline-block; margin-left: 8px; padding: 4px 12px; border: 1px solid var(--warning); border-radius: var(--radius-sm); cursor: pointer; font-size: 11px; font-weight: 500; background: var(--warning-bg); color: var(--warning); font-family: inherit; transition: all 0.15s; }
.approve-phase-btn:hover { filter: brightness(1.1); } .approve-phase-btn:hover { filter: brightness(1.1); }
/* Approval section — prominent card for approval-gated tasks */
.approval-section {
margin: 12px 0;
padding: 14px 16px;
border: 2px solid var(--warning);
border-radius: var(--radius-md);
background: var(--warning-bg);
}
.approval-header {
display: flex;
align-items: center;
gap: 8px;
margin-bottom: 8px;
}
.approval-icon { font-size: 18px; }
.approval-title { font-size: 14px; font-weight: 600; flex: 1; }
.approval-blocker { font-size: 12px; color: var(--text-secondary); margin: 0 0 8px 0; padding: 6px 10px; background: var(--bg-card); border-radius: var(--radius-sm); }
.approval-missing { font-size: 12px; color: var(--text-secondary); margin: 8px 0 0 0; font-style: italic; }
.approval-artifact { margin-top: 8px; font-size: 12px; }
.approval-artifact summary { cursor: pointer; font-weight: 500; padding: 4px 0; color: var(--text-primary); }
.approval-artifact summary:hover { color: var(--primary); }
.task-card-artifacts { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 6px; } .task-card-artifacts { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 6px; }
.artifact-badge { font-size: 10px; padding: 2px 8px; background: var(--bg-primary); border: 1px solid var(--border-color); border-radius: 4px; color: var(--text-secondary); font-family: 'SF Mono', 'Fira Code', monospace; font-weight: 500; } .artifact-badge { font-size: 10px; padding: 2px 8px; background: var(--bg-primary); border: 1px solid var(--border-color); border-radius: 4px; color: var(--text-secondary); font-family: 'SF Mono', 'Fira Code', monospace; font-weight: 500; }
.task-card-models { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 6px; }
.model-badge { font-size: 10px; padding: 1px 6px; background: #e3f2fd; border: 1px solid #90caf9; border-radius: 4px; color: #1565c0; font-family: 'SF Mono', 'Fira Code', monospace; font-weight: 500; }
/* Transition button — advance to next phase */
.transition-btn { display: inline-block; padding: 6px 16px; border: 1px solid var(--primary); border-radius: var(--radius-sm); cursor: pointer; font-size: 12px; font-weight: 500; background: rgba(74, 144, 226, 0.1); color: var(--primary); font-family: inherit; transition: all 0.15s; }
.transition-btn:hover { filter: brightness(1.15); background: rgba(74, 144, 226, 0.2); }
.detail-actions { margin: 8px 0; padding: 10px 12px; background: var(--bg-card); border-radius: var(--radius-md); border: 1px solid var(--border-color); }
.detail-actions h4 { margin: 0 0 8px 0; font-size: 12px; color: var(--text-secondary); font-weight: 500; }
/* Artifact editor — inline textarea for writing missing artifacts */
.detail-artifact-editor { margin: 8px 0; padding: 12px; background: var(--bg-card); border: 1px solid var(--primary); border-radius: var(--radius-md); }
.detail-artifact-editor h4 { margin: 0 0 4px 0; font-size: 12px; font-weight: 600; }
.artifact-editor-hint { font-size: 11px; color: var(--text-secondary); margin: 0 0 8px 0; }
.artifact-editor { width: 100%; padding: 10px; border: 1px solid var(--border-color); border-radius: var(--radius-sm); background: var(--bg-primary); color: var(--text-primary); font-family: 'SF Mono', 'Fira Code', monospace; font-size: 12px; line-height: 1.5; resize: vertical; box-sizing: border-box; }
.artifact-editor:focus { outline: none; border-color: var(--primary); }
.save-artifact-btn { display: inline-block; margin-top: 8px; padding: 6px 16px; border: 1px solid var(--success); border-radius: var(--radius-sm); cursor: pointer; font-size: 12px; font-weight: 500; background: rgba(74, 208, 120, 0.1); color: var(--success); font-family: inherit; transition: all 0.15s; }
.save-artifact-btn:hover { filter: brightness(1.15); background: rgba(74, 208, 120, 0.2); }
@media (max-width: 768px) { @media (max-width: 768px) {
.header { flex-wrap: wrap; gap: 8px; } .header { flex-wrap: wrap; gap: 8px; }
+119 -2
View File
@@ -51,6 +51,21 @@ TASK_STATE_ARTIFACT = {
TASK_STATES = list(TASK_STATE_ARTIFACT.keys()) TASK_STATES = list(TASK_STATE_ARTIFACT.keys())
def next_transition(phase_raw: str) -> str | None:
transitions = {
"research:approved": "decomposition",
"decomposition:approved": "design",
"design:approved": "implement",
"test_design:approved": "implement",
"implement": "code_review",
"code_review:approved": "bug_find",
"bug_find": "adv_bug_find",
"adv_bug_find": "doc_review",
"doc_review": "referee",
"referee": "complete",
}
return transitions.get(phase_raw)
MAX_POST_BODY = 65536 # 64KB MAX_POST_BODY = 65536 # 64KB
MAX_REVIEW_COMMENT_LENGTH = 4096 MAX_REVIEW_COMMENT_LENGTH = 4096
CACHE_TTL = 1.0 # seconds CACHE_TTL = 1.0 # seconds
@@ -120,6 +135,12 @@ class DashboardHandler(SimpleHTTPRequestHandler):
elif self.path.startswith("/api/approve/"): elif self.path.startswith("/api/approve/"):
task_name = unquote(self.path.split("/api/approve/")[1]) task_name = unquote(self.path.split("/api/approve/")[1])
self._handle_phase_approval(task_name) self._handle_phase_approval(task_name)
elif self.path.startswith("/api/transition/"):
task_name = unquote(self.path.split("/api/transition/")[1])
self._handle_transition(task_name)
elif self.path.startswith("/api/write-artifact/"):
task_name = unquote(self.path.split("/api/write-artifact/")[1])
self._handle_write_artifact(task_name)
else: else:
self._send_error(404, "Not found") self._send_error(404, "Not found")
@@ -196,6 +217,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"name": t.name, "name": t.name,
"display_name": t.display_name, "display_name": t.display_name,
"state": t.state.value, "state": t.state.value,
"phase_raw": t.phase_raw,
"status_reason": t.status_reason, "status_reason": t.status_reason,
"artifacts": {s: TASK_STATE_ARTIFACT[s] in t.artifacts for s in TASK_STATES}, "artifacts": {s: TASK_STATE_ARTIFACT[s] in t.artifacts for s in TASK_STATES},
"sub_tasks": [ "sub_tasks": [
@@ -204,8 +226,14 @@ class DashboardHandler(SimpleHTTPRequestHandler):
], ],
"verdict_content": t.verdict_content, "verdict_content": t.verdict_content,
"bug_report_content": t.bug_report_content, "bug_report_content": t.bug_report_content,
"adversarial_bug_report_content": t.adversarial_bug_report_content,
"doc_review_content": t.doc_review_content,
"design_content": t.design_content,
"spec_content": t.spec_content, "spec_content": t.spec_content,
"decomposition_content": t.decomposition_content, "decomposition_content": t.decomposition_content,
"code_review_content": t.code_review_content,
"test_plan_content": t.test_plan_content,
"implementation_content": t.implementation_content,
"parent_spec_content": t.parent_spec_content, "parent_spec_content": t.parent_spec_content,
"vram_config_content": t.vram_config_content, "vram_config_content": t.vram_config_content,
"blocked_action_items": t.blocked_action_items, "blocked_action_items": t.blocked_action_items,
@@ -217,7 +245,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"is_approval_gated": t.is_approval_gated, "is_approval_gated": t.is_approval_gated,
"blocker": t.blocker, "blocker": t.blocker,
"waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in t.waves], "waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in t.waves],
"review": self._get_review_status(t.name), "models": t.models,
} }
for t in tasks for t in tasks
] ]
@@ -302,6 +330,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"name": task.name, "name": task.name,
"display_name": task.display_name, "display_name": task.display_name,
"state": task.state.value, "state": task.state.value,
"phase_raw": task.phase_raw,
"status_reason": task.status_reason, "status_reason": task.status_reason,
"artifacts": {s: TASK_STATE_ARTIFACT[s] in task.artifacts for s in TASK_STATES}, "artifacts": {s: TASK_STATE_ARTIFACT[s] in task.artifacts for s in TASK_STATES},
"sub_tasks": [ "sub_tasks": [
@@ -310,12 +339,22 @@ class DashboardHandler(SimpleHTTPRequestHandler):
], ],
"verdict_content": task.verdict_content, "verdict_content": task.verdict_content,
"bug_report_content": task.bug_report_content, "bug_report_content": task.bug_report_content,
"adversarial_bug_report_content": task.adversarial_bug_report_content,
"doc_review_content": task.doc_review_content,
"design_content": task.design_content,
"spec_content": task.spec_content, "spec_content": task.spec_content,
"decomposition_content": task.decomposition_content, "decomposition_content": task.decomposition_content,
"parent_spec_content": task.parent_spec_content, "parent_spec_content": task.parent_spec_content,
"vram_config_content": task.vram_config_content, "vram_config_content": task.vram_config_content,
"blocked_action_items": task.blocked_action_items,
"unblock_instructions": task.unblock_instructions,
"phase_guidance": task.phase_guidance,
"required_artifact_name": task.required_artifact_name,
"next_phase_name": task.next_phase_name,
"is_edit_phase": task.is_edit_phase,
"is_approval_gated": task.is_approval_gated,
"blocker": task.blocker,
"waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in task.waves], "waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in task.waves],
"review": self._get_review_status(task.name),
} }
self._send_json(task_data) self._send_json(task_data)
@@ -416,6 +455,84 @@ class DashboardHandler(SimpleHTTPRequestHandler):
else: else:
self._send_error(400, result.stderr.strip() or result.stdout.strip()) self._send_error(400, result.stderr.strip() or result.stdout.strip())
def _handle_transition(self, task_name: str):
if not self._validate_task_name(task_name):
self._send_error(400, "Invalid task name")
return
project_root = self.project_root
if not project_root:
self._send_error(503, "Not in automaton project")
return
status_py = Path.home() / ".automaton" / "scripts" / "status.py"
if not status_py.exists():
self._send_error(500, "status.py not found")
return
import subprocess
content_length = int(self.headers.get('Content-Length', 0))
target = None
if content_length > 0:
try:
body = json.loads(self.rfile.read(content_length))
target = body.get("target")
except (json.JSONDecodeError, UnicodeDecodeError):
pass
if not target:
tasks = _get_cached_tasks(project_root)
task = next((t for t in tasks if t.name == task_name), None)
if not task:
self._send_error(404, "Task not found")
return
target = next_transition(task.phase_raw)
if not target:
self._send_json({"success": False, "message": "Cannot determine next transition from current phase"})
return
result = subprocess.run(
[sys.executable, str(status_py), "--transition", target, "--task", task_name,
"--project", str(project_root)],
capture_output=True, text=True, timeout=30,
)
_invalidate_task_cache()
if result.returncode == 0:
self._send_json({"success": True, "message": f"Transitioned to {target}: {result.stdout.strip()}"})
else:
self._send_json({"success": False, "message": result.stderr.strip() or result.stdout.strip()})
def _handle_write_artifact(self, task_name: str):
if not self._validate_task_name(task_name):
self._send_error(400, "Invalid task name")
return
project_root = self.project_root
if not project_root:
self._send_error(503, "Not in automaton project")
return
content_length = int(self.headers.get('Content-Length', 0))
if content_length > MAX_POST_BODY:
self._send_error(413, "Payload too large")
return
if content_length == 0:
self._send_error(400, "Empty request body")
return
try:
body = json.loads(self.rfile.read(content_length))
filename = body.get("filename", "")
content = body.get("content", "")
if not filename or not filename.endswith(".md"):
self._send_error(400, "Invalid filename — must be a .md file")
return
tasks = _get_cached_tasks(project_root)
task = next((t for t in tasks if t.name == task_name), None)
if not task:
self._send_error(404, "Task not found")
return
artifact_path = Path(task.folder_path) / filename
artifact_path.write_text(content, encoding="utf-8")
_invalidate_task_cache()
self._send_json({"success": True, "message": f"Written {filename}"})
except (json.JSONDecodeError, UnicodeDecodeError):
self._send_error(400, "Invalid JSON")
except (OSError, IOError) as e:
self._send_error(500, f"Failed to write file: {e}")
def _serve_review_summary(self): def _serve_review_summary(self):
project_root = self.project_root project_root = self.project_root
if not project_root: if not project_root:
+13 -2
View File
@@ -33,8 +33,8 @@ To disable auto-detection and use manual values:
Settings for the LLM model being used. Settings for the LLM model being used.
- **Model**: auto # Use auto-detection from API config, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet) - **Model**: omlx/Ornith-1.0-35B-4bit-mlx # Local LLM (opencode provider); used as the Implement role
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k) - **Override context window**: 32768 # Matches opencode.json limit.context for ornith
### Auto-detection ### Auto-detection
@@ -51,6 +51,17 @@ To disable auto-detection and use manual values:
- **Override context window**: 128k - **Override context window**: 128k
``` ```
## Available Models
Models available for model-divergence enforcement. This file is managed by `scripts/detect_models.py`. In single-LLM mode (0-1 models), no hard blocks are enforced. In multi-LLM mode (2+ models), the conflict matrix enforces role-model separation.
- **Default**: omlx/Ornith-1.0-35B-4bit-mlx # Used when no role-specific binding is set
- **Advised**: true # Recommend a second model in single-LLM mode
No additional models are configured in the manifest. To add models:
1. Run `python3 ~/.automaton/scripts/detect_models.py --write` to auto-detect from opencode.json and localhost endpoints.
2. Or manually create `~/.automaton/models.json` (see `design/framework/technical.md` §2 for schema).
## System Requirements ## System Requirements
Requirements for the environment the framework runs in. Requirements for the environment the framework runs in.
+5
View File
@@ -103,6 +103,7 @@ Gate checks, in order:
4. **Task phase** — if `current_task` is set, that task's `.state` must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to `human_intervention` by some other path), halt as `human_intervention`. 4. **Task phase** — if `current_task` is set, that task's `.state` must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to `human_intervention` by some other path), halt as `human_intervention`.
5. **Worktree drift** — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via `git diff --name-only main...HEAD` restricted to `file_scope`), halt as `drift_detected`. 5. **Worktree drift** — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via `git diff --name-only main...HEAD` restricted to `file_scope`), halt as `drift_detected`.
6. **Score plateau** — last N entries in `score_history` are flat or monotonically decreasing (where N = `score_plateau_window`). Trip → halt as `verifier_failed`. 6. **Score plateau** — last N entries in `score_history` are flat or monotonically decreasing (where N = `score_plateau_window`). Trip → halt as `verifier_failed`.
7. **Model divergence** — in multi-LLM mode (2+ models in `models.json`), checks that the loop's implement and verify roles use different models. If they share the same model, halt as `human_intervention` (this prevents same-model verification / rubber-stamping within a loop tick). Single-LLM mode is exempt. Model is resolved from `roles[<role>].model` if set, otherwise the manifest default.
All halts atomically set `status=halted`, `halt_reason=<reason>`, write to `.state.log`, and call `--pause-loop`'s schedule-disable step (see §6). All halts atomically set `status=halted`, `halt_reason=<reason>`, write to `.state.log`, and call `--pause-loop`'s schedule-disable step (see §6).
@@ -258,11 +259,13 @@ To bound `outputs/` directory growth (O5 from `add-loop-runner/BUG_REPORT.md`),
v1.1's default `harness.command` is `opencode run` -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the `command` list (single argv element per token, no shell expansion): v1.1's default `harness.command` is `opencode run` -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the `command` list (single argv element per token, no shell expansion):
- `{model}` -- the model assigned to the role (from `roles[<role>].model` or manifest default). Passed via `extras["model"]`. If the role has no model assignment, the token is left unsubstituted.
- `{prompt}` -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path. - `{prompt}` -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path.
- `{prompt_content}` -- the resolved prompt file's text content as a single argv element. Safe under `subprocess.run` list mode; no shell quoting needed. Used by the default command since `opencode run` takes the message as a positional argument and has no `--prompt-file` flag. - `{prompt_content}` -- the resolved prompt file's text content as a single argv element. Safe under `subprocess.run` list mode; no shell quoting needed. Used by the default command since `opencode run` takes the message as a positional argument and has no `--prompt-file` flag.
- `{cwd}` -- the working directory the harness should run in (the loop's project root or worktree). - `{cwd}` -- the working directory the harness should run in (the loop's project root or worktree).
- `{output}`, `{artifact}` -- role-specific extras (the implement output path handed to verify). - `{output}`, `{artifact}` -- role-specific extras (the implement output path handed to verify).
- `{verdict}`, `{current_task}`, `{current_phase}`, etc. -- other runtime extras; see `_resolve_prompt` below. - `{verdict}`, `{current_task}`, `{current_phase}`, etc. -- other runtime extras; see `_resolve_prompt` below.
- `{model}` -- the model assigned to the role being invoked (from `roles[<role>].model` in `loop.json`, or the manifest default). The runner passes it via the `extras["model"]` key. If the role has no explicit model, `{model}` is left as-is (no substitution). This allows per-role model pinning without hardcoding the model name in `harness.command`.
The default command does NOT hardcode a `--model` flag; the spawned `opencode run` inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set `harness.command` in their `loop.json`: The default command does NOT hardcode a `--model` flag; the spawned `opencode run` inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set `harness.command` in their `loop.json`:
@@ -336,6 +339,8 @@ This means the harness receives a fully-resolved prompt file with all context ba
} }
``` ```
Each role in `roles` accepts an optional `"model"` field to pin a specific model for that role (e.g. `"implement": {"prompt": "loop-implement.md", "tier": 16000, "model": "model-a"}`). When set, the runner passes `model=<value>` in the harness extras for that role, enabling `{model}` substitution in `harness.command`. This is how multi-LLM loops prevent same-model verification — see `CONFLICT_MATRIX` in `status.py`.
Installs default-on at `install.sh` time: `status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` then `status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"`. Both commands use `|| true` so the framework works even if loop creation fails. `update.sh` bootstraps the loop idempotently for existing users (checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`). Disabling: `status.py --pause-loop self-improvement --project ~/.automaton/`. Installs default-on at `install.sh` time: `status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` then `status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"`. Both commands use `|| true` so the framework works even if loop creation fails. `update.sh` bootstraps the loop idempotently for existing users (checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`). Disabling: `status.py --pause-loop self-improvement --project ~/.automaton/`.
## 10. Tests (`tests/test_loops.py`) ## 10. Tests (`tests/test_loops.py`)
+141
View File
@@ -0,0 +1,141 @@
import type { ExtensionAPI, ExtensionContext, BeforeAgentStartEventResult } from "@earendil-works/pi-coding-agent";
import { existsSync, readFileSync, readdirSync, statSync } from "fs";
import { basename, join } from "path";
import { homedir } from "os";
const AUTOMATON_HOME = join(homedir(), ".automaton");
const STALE_MINUTES = 30;
interface TaskInfo {
name: string;
phase: string;
mtime: Date;
}
function getTasks(autoDir: string): TaskInfo[] {
const tasksDir = join(autoDir, ".automaton", "tasks");
if (!existsSync(tasksDir)) return [];
try {
const entries = readdirSync(tasksDir);
const tasks: TaskInfo[] = [];
for (const entry of entries) {
const stateFile = join(tasksDir, entry, ".state");
if (existsSync(stateFile)) {
try {
const phase = readFileSync(stateFile, "utf-8").trim();
if (!phase) continue;
const stat = statSync(stateFile);
tasks.push({ name: entry, phase, mtime: stat.mtime });
} catch {
// skip unreadable
}
}
}
tasks.sort((a, b) => b.mtime.getTime() - a.mtime.getTime());
return tasks;
} catch {
return [];
}
}
function getLoopInfo(autoDir: string): string[] {
const loopsDir = join(autoDir, ".automaton", "loops");
if (!existsSync(loopsDir)) return [];
try {
const entries = readdirSync(loopsDir);
const lines: string[] = [];
for (const entry of entries) {
const stateFile = join(loopsDir, entry, ".state.loop");
if (existsSync(stateFile)) {
try {
const content = readFileSync(stateFile, "utf-8").trim();
const state = JSON.parse(content);
const taskRef = state.current_task ? `, active task: ${state.current_task}` : "";
lines.push(`Loop "${entry}": ${state.status || "unknown"}${taskRef}`);
} catch {
// skip unparseable
}
}
}
return lines;
} catch {
return [];
}
}
export default function (pi: ExtensionAPI) {
pi.on("before_agent_start", async (_event, ctx): Promise<BeforeAgentStartEventResult | undefined> => {
const cwd = ctx.cwd;
const inFramework = cwd === AUTOMATON_HOME || cwd.startsWith(AUTOMATON_HOME + "/");
let baseDir: string | null = null;
let scopeLabel: string;
if (inFramework) {
baseDir = AUTOMATON_HOME;
scopeLabel = "framework (~/.automaton/)";
} else {
const projectAuto = join(cwd, ".automaton");
if (existsSync(projectAuto)) {
baseDir = cwd;
const projectName = basename(cwd) || "project";
scopeLabel = `project (${projectName}/.automaton/)`;
} else {
return; // not in automaton context
}
}
const tasks = getTasks(baseDir);
const editTasks = tasks.filter((t) => t.phase === "implement" || t.phase === "doc_review");
const currentTask = editTasks.length > 0 ? editTasks[0] : tasks.length > 0 ? tasks[0] : null;
const lines: string[] = [];
lines.push(`Automaton scope: ${scopeLabel}`);
lines.push("");
lines.push("How Automaton works:");
lines.push("- Automaton enforces a task phase state machine. File edits are ONLY allowed when a task is in 'implement' or 'doc_review' phase.");
lines.push("- When the user asks you to build, fix, or change something, you MUST first create a task (use automaton_create_task) and transition it to 'implement' before you can edit any files.");
lines.push("- Without a task in 'implement', the automaton-guard-pi extension will BLOCK all file edits.");
lines.push(`- Available tools: automaton_create_task (create+transition), automaton_transition (change phase), automaton_status (check state).`);
lines.push("");
if (currentTask) {
lines.push(`State: Active task "${currentTask.name}" is in "${currentTask.phase}" phase.`);
if (currentTask.phase === "implement" || currentTask.phase === "doc_review") {
lines.push(`You CAN edit files under this task.`);
const ageMinutes = Math.round((Date.now() - currentTask.mtime.getTime()) / 60000);
if (ageMinutes > STALE_MINUTES) {
lines.push(
`Note: task has been in ${currentTask.phase} for ${ageMinutes} min and is considered stale. ` +
`Run automaton_transition (or touch via status.py) if still active.`,
);
}
} else {
lines.push(`File edits are BLOCKED. Call automaton_transition to move it to 'implement' before editing.`);
}
} else {
lines.push(`State: No active tasks. When the user makes a work request, call automaton_create_task first.`);
}
const byPhase = new Map<string, number>();
for (const t of tasks) {
byPhase.set(t.phase, (byPhase.get(t.phase) || 0) + 1);
}
const summary = Array.from(byPhase.entries())
.sort((a, b) => b[1] - a[1])
.map(([p, c]) => `${p} (${c})`)
.join(", ");
lines.push(`All tasks: ${summary || "none"}`);
const loopInfo = getLoopInfo(baseDir);
for (const l of loopInfo) {
lines.push(l);
}
return {
systemPrompt: `${_event.systemPrompt}\n\n## Automaton Context\n\n${lines.join("\n")}`,
};
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-automaton-context",
"version": "1.0.0",
"description": "Auto-injects Automaton framework context into Pi Dev system prompt — no more manual copy-paste of pi-automaton.sh output",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+171
View File
@@ -0,0 +1,171 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "@sinclair/typebox";
import { execSync } from "child_process";
import { existsSync, readFileSync, readdirSync } from "fs";
import { join } from "path";
import { homedir } from "os";
const STATUS_SCRIPT = join(homedir(), ".automaton", "scripts", "status.py");
function resolveProjectDir(cwd: string): string | null {
const inFramework = cwd === join(homedir(), ".automaton") || cwd.startsWith(join(homedir(), ".automaton") + "/");
if (inFramework) return homedir() + "/.automaton";
if (existsSync(join(cwd, ".automaton"))) return cwd;
return null;
}
function getTasks(projectDir: string) {
const tasksDir = join(projectDir, ".automaton", "tasks");
if (!existsSync(tasksDir)) return [];
const tasks: { name: string; phase: string }[] = [];
for (const entry of readdirSync(tasksDir)) {
const stateFile = join(tasksDir, entry, ".state");
if (existsSync(stateFile)) {
try {
const phase = readFileSync(stateFile, "utf-8").trim();
if (phase) tasks.push({ name: entry, phase });
} catch {}
}
}
tasks.sort((a, b) => a.name.localeCompare(b.name));
return tasks;
}
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "automaton_create_task",
label: "Create Automaton Task",
description:
"Create a new automaton task and transition it to the specified phase. " +
"Use this when the user asks you to do work that requires file edits — you need a task in 'implement' or 'doc_review' phase before you can edit files.",
promptSnippet: "Create automaton tasks for work management",
promptGuidelines: [
"When the user asks you to build, implement, fix, or change something, first call automaton_create_task to create a task and transition it to 'implement' phase",
"Only after the task is in 'implement' can you edit files — the guard will block edits otherwise",
"Use a descriptive task name based on what the user wants (e.g., 'add-login-page', 'fix-api-timeout')",
"Default phase is 'implement' — omit phase for most cases",
],
parameters: Type.Object({
name: Type.String({ minLength: 1, description: "Task name (kebab-case, e.g. 'add-login-page')" }),
phase: Type.Optional(Type.String({ default: "implement", description: "Phase to transition to after creation" })),
}),
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { name, phase = "implement" } = params;
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project. No .automaton/ directory found." }],
details: {},
};
}
try {
const createOut = execSync(
`python3 ${STATUS_SCRIPT} --create-task ${JSON.stringify(name)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
const transitionOut = execSync(
`python3 ${STATUS_SCRIPT} --transition ${phase} --task ${JSON.stringify(name)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
return {
content: [{ type: "text", text: `Created task "${name}" and transitioned to "${phase}".` }],
details: { task: name, phase },
};
} catch (e: any) {
return {
content: [{
type: "text",
text: `Failed to create task: ${e.stderr || e.message || String(e)}`,
}],
details: { error: e.stderr || e.message, task: name },
};
}
},
});
pi.registerTool({
name: "automaton_transition",
label: "Transition Automaton Task",
description:
"Transition an existing automaton task to a new phase. " +
"Valid phases: research, decomposition, design, implement, test_design, testing, doc_review, complete. " +
"Use this to move a task forward (e.g. from research to implement).",
promptSnippet: "Transition automaton tasks between phases",
parameters: Type.Object({
task: Type.String({ minLength: 1, description: "Task name" }),
phase: Type.String({ minLength: 1, description: "Target phase" }),
}),
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { task, phase } = params;
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project." }],
details: {},
};
}
try {
execSync(
`python3 ${STATUS_SCRIPT} --transition ${phase} --task ${JSON.stringify(task)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
return {
content: [{ type: "text", text: `Task "${task}" transitioned to "${phase}".` }],
details: { task, phase },
};
} catch (e: any) {
return {
content: [{
type: "text",
text: `Failed to transition task: ${e.stderr || e.message || String(e)}`,
}],
details: { error: e.stderr || e.message, task, phase },
};
}
},
});
pi.registerTool({
name: "automaton_status",
label: "Automaton Status",
description:
"Show current automaton project status — all tasks, their phases, and any running loops. " +
"Call this to check what state things are in before deciding what to do.",
promptSnippet: "Check automaton project status",
parameters: Type.Object({}),
async execute(_toolCallId, _params, _signal, _onUpdate, _ctx) {
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project." }],
details: {},
};
}
const tasks = getTasks(projectDir);
const editTasks = tasks.filter((t) => t.phase === "implement" || t.phase === "doc_review");
const currentTask = editTasks.length > 0 ? editTasks[0] : tasks.length > 0 ? tasks[0] : null;
const lines: string[] = [];
lines.push(`Project: ${projectDir.split("/").pop()}`);
if (currentTask) {
lines.push(`Current: "${currentTask.name}" (${currentTask.phase})`);
} else {
lines.push("No tasks. Create one with automaton_create_task.");
}
lines.push("");
for (const t of tasks) {
const marker = currentTask && t.name === currentTask.name ? ">" : " ";
lines.push(`${marker} ${t.name.padEnd(35)} ${t.phase}`);
}
return {
content: [{ type: "text", text: lines.join("\n") }],
details: { tasks: tasks.length, current: currentTask?.name || null },
};
},
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-automaton-tools",
"version": "1.0.0",
"description": "Registers automaton_create_task and automaton_transition tools so the agent can manage task lifecycle automatically",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+91
View File
@@ -0,0 +1,91 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "@sinclair/typebox";
const BIRD_BIN = "/opt/homebrew/bin/bird";
const TweetParams = Type.Object({
url: Type.String({ minLength: 1, description: "Tweet URL or numeric ID" }),
mode: Type.Optional(
Type.Union([
Type.Literal("read"),
Type.Literal("thread"),
Type.Literal("replies"),
]),
),
});
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "read_tweet",
label: "Read Tweet",
description:
"Read an X/Twitter tweet or thread using the local bird CLI. " +
"Returns tweet author, text, media, metrics, and timestamps. " +
"Use this instead of webfetch for x.com/twitter.com URLs.",
promptSnippet: "Read X/Twitter tweets with bird CLI",
promptGuidelines: [
"Use read_tweet when the user shares an x.com or twitter.com URL and asks what it says",
"Use mode='thread' for full conversation threads",
"Use mode='replies' to fetch replies to a tweet",
"If fetch fails with auth errors, ask the user to sign in to x.com in their browser and retry",
],
parameters: TweetParams,
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { url, mode = "read" } = params;
try {
const { execSync } = await import("child_process");
const cmd = `${BIRD_BIN} ${mode} ${JSON.stringify(url)} --json`;
const stdout = execSync(cmd, { encoding: "utf-8", timeout: 15000 });
const parsed = JSON.parse(stdout);
const author = parsed.author?.name || parsed.author?.screen_name || "unknown";
const text = parsed.text || parsed.content || "";
const createdAt = parsed.created_at || "";
const retweetCount = parsed.metrics?.retweet_count ?? parsed.metrics?.retweets ?? 0;
const likeCount = parsed.metrics?.like_count ?? parsed.metrics?.likes ?? 0;
const replyCount = parsed.metrics?.reply_count ?? parsed.metrics?.replies ?? 0;
const mediaCount = parsed.media?.length ?? 0;
const threadCount = parsed.thread?.tweets?.length ?? 0;
const repliesCount = parsed.replies?.length ?? 0;
const lines: string[] = [];
lines.push(`Author: ${author}`);
if (createdAt) lines.push(`Posted: ${createdAt}`);
lines.push("");
lines.push(text);
if (retweetCount || likeCount || replyCount) {
lines.push("");
lines.push(`Retweets: ${retweetCount} Likes: ${likeCount} Replies: ${replyCount}`);
}
if (mediaCount) lines.push(`Media: ${mediaCount} attachment(s)`);
if (threadCount) lines.push(`Thread: ${threadCount} tweets total`);
if (repliesCount) lines.push(`Replies fetched: ${repliesCount}`);
return {
content: [
{ type: "text", text: lines.join("\n") },
{ type: "text", text: `\n--- raw ---\n${JSON.stringify(parsed, null, 2)}` },
],
details: {
author,
text: text.slice(0, 500),
url,
mode,
tweetCount: threadCount || repliesCount || 1,
},
};
} catch (e: any) {
const errMsg = e.stderr || e.message || String(e);
const hint = errMsg.includes("auth")
? " Sign in to x.com in your browser and retry."
: "";
return {
content: [{ type: "text", text: `Failed to fetch tweet: ${errMsg}${hint}` }],
details: { error: errMsg, url, mode },
};
}
},
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-read-tweet",
"version": "1.0.0",
"description": "Read X/Twitter tweets using the local `bird` CLI — registered as a read_tweet tool for Pi Dev",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+1 -1
View File
@@ -1,2 +1,2 @@
#!/usr/bin/env bash #!/usr/bin/env bash
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/Users/laptran/.automaton" python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/private/var/folders/f5/yv0dzbnx47x3yp8sc_2519gh0000gn/T/pytest-of-laptran/pytest-130/test_uninstall_via_disabled_re0"
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env python3
"""Probe opencode.json and localhost endpoints to produce a candidate models.json.
Usage:
python3 scripts/detect_models.py [--json] [--write]
Without --json, prints a human-readable report.
With --json, emits the candidate models.json to stdout as the last JSON line.
With --write, writes the candidate to ~/.automaton/models.json (idempotent,
never overwrites an existing file unless --force is also given).
Probing strategy (stdlib only):
1. Parse opencode.json (or opencode.jsonc) for configured provider+model pairs.
2. Probe localhost endpoints to find locally-running LLM servers:
- http://localhost:8080/v1/models (llama.cpp / generic OpenAI-compatible)
- http://localhost:11434/api/tags (Ollama)
- http://localhost:1234/v1/models (LM Studio)
- http://localhost:8000/v1/models (vLLM)
3. Merge results into a candidate models.json.
"""
from __future__ import annotations
import json
import os
import re
import sys
from pathlib import Path
from typing import Optional
AUTOMATON_DIR = Path.home() / ".automaton"
# ---------------------------------------------------------------------------
# opencode.json parsing
# ---------------------------------------------------------------------------
def _find_opencode_json() -> Optional[Path]:
"""Locate the opencode config file (opencode.json or opencode.jsonc)."""
candidates = [
Path.cwd() / "opencode.json",
Path.cwd() / "opencode.jsonc",
AUTOMATON_DIR / "opencode.json",
AUTOMATON_DIR / "opencode.jsonc",
Path.home() / ".opencode.json",
Path.home() / ".config" / "opencode" / "opencode.json",
Path.home() / ".config" / "opencode" / "opencode.jsonc",
]
for p in candidates:
if p.exists():
return p
return None
def _parse_opencode_models(config_path: Path) -> list[dict]:
"""Extract model entries from an opencode.json config.
Expected structure (common patterns):
{
"providers": {
"opencode": { "model": "glm-4.6", ... },
...
}
}
or a flatter:
{
"model": "glm-4.6",
...
}
"""
try:
content = config_path.read_text(encoding="utf-8")
except OSError:
return []
# Strip JSONC comments (// line comments only, sufficient for our use)
content = re.sub(r"//.*", "", content)
try:
data = json.loads(content)
except json.JSONDecodeError:
return []
if not isinstance(data, dict):
return []
models: list[dict] = []
seen: set[str] = set()
# Check top-level "model" field (single-model config)
single = data.get("model")
if isinstance(single, str) and single not in seen:
seen.add(single)
models.append({"name": single, "provider": "opencode", "context_window": None, "location": "remote"})
# Check providers dict
providers = data.get("providers") or {}
for prov_name, prov_cfg in providers.items():
if isinstance(prov_cfg, dict):
model_name = prov_cfg.get("model")
if isinstance(model_name, str) and model_name not in seen:
seen.add(model_name)
models.append({"name": model_name, "provider": prov_name, "context_window": None, "location": "remote"})
# Check "models" list (explicit model roster)
model_list = data.get("models")
if isinstance(model_list, list):
for entry in model_list:
if isinstance(entry, dict):
name = entry.get("name") or entry.get("model")
if isinstance(name, str) and name not in seen:
seen.add(name)
models.append({
"name": name,
"provider": entry.get("provider", "opencode"),
"context_window": entry.get("context_window"),
"location": entry.get("location", "remote"),
})
return models
# ---------------------------------------------------------------------------
# Localhost probing
# ---------------------------------------------------------------------------
def _fetch_json(url: str, timeout: int = 5) -> Optional[dict]:
"""Fetch a JSON response from a URL using urllib (stdlib)."""
import urllib.request
import urllib.error
try:
req = urllib.request.Request(url, method="GET")
with urllib.request.urlopen(req, timeout=timeout) as resp:
body = resp.read().decode("utf-8")
return json.loads(body)
except (OSError, urllib.error.URLError, json.JSONDecodeError, ValueError):
return None
def _probe_ollama() -> list[dict]:
"""Probe Ollama: GET http://localhost:11434/api/tags → models[].name"""
data = _fetch_json("http://localhost:11434/api/tags")
if not data:
return []
models_list = data.get("models") or []
return [
{"name": m.get("name"), "provider": "ollama", "context_window": None, "location": "http://localhost:11434"}
for m in models_list
if isinstance(m, dict) and isinstance(m.get("name"), str)
]
def _probe_openai_compatible(url: str, provider: str) -> list[dict]:
"""Probe an OpenAI-compatible /v1/models endpoint."""
data = _fetch_json(url)
if not data:
return []
model_list = data.get("data") or []
return [
{"name": m.get("id"), "provider": provider, "context_window": None, "location": url}
for m in model_list
if isinstance(m, dict) and isinstance(m.get("id"), str)
]
_ENDPOINTS = [
("http://localhost:8080/v1/models", "llama.cpp"),
("http://localhost:11434/api/tags", "ollama"), # handled separately above
("http://localhost:1234/v1/models", "lm-studio"),
("http://localhost:8000/v1/models", "vllm"),
]
def _probe_localhost() -> list[dict]:
"""Probe all known localhost endpoints and merge results."""
seen_names: set[str] = set()
models: list[dict] = []
for url, provider in _ENDPOINTS:
if provider == "ollama":
result = _probe_ollama()
else:
result = _probe_openai_compatible(url, provider)
for m in result:
n = m.get("name")
if isinstance(n, str) and n not in seen_names:
seen_names.add(n)
models.append(m)
return models
# ---------------------------------------------------------------------------
# Merge & write
# ---------------------------------------------------------------------------
def build_candidate_models(probe_local: bool = True) -> dict:
"""Build a candidate models.json dict.
1. Parse models from opencode.json
2. Optionally probe localhost endpoints
3. Merge: opencode config models come first; local probes fill in gaps.
4. Build result with default, advised, models[].
"""
opencode_path = _find_opencode_json()
config_models: list[dict] = []
if opencode_path:
config_models = _parse_opencode_models(opencode_path)
local_models: list[dict] = []
if probe_local:
local_models = _probe_localhost()
# Merge: key by name, config models take priority (unordered)
merged: dict[str, dict] = {}
for m in config_models:
n = m["name"]
if n not in merged:
merged[n] = m
for m in local_models:
n = m.get("name")
if n and n not in merged:
merged[n] = m
models_list = list(merged.values())
# Determine default: first config model, or first local model, or empty
default_name: Optional[str] = None
if config_models:
default_name = config_models[0].get("name")
elif local_models:
default_name = local_models[0].get("name")
# Determine advised: if only 0-1 models, set advised=true; else false
advised = len(models_list) <= 1
result: dict = {
"schema_version": 1,
"default": default_name,
"advised": advised,
"models": models_list,
}
return result
def write_models_file(candidate: dict, force: bool = False) -> bool:
"""Write candidate models.json to AUTOMATON_DIR.
Never overwrites an existing file unless force=True.
Returns True if written, False if skipped.
"""
target = AUTOMATON_DIR / "models.json"
if target.exists() and not force:
return False
target.write_text(json.dumps(candidate, indent=2) + "\n")
return True
def format_report(candidate: dict) -> str:
"""Human-readable report of the candidate models."""
lines = []
lines.append("=== Model Detection Report ===")
lines.append("")
source = "No opencode.json found" if not _find_opencode_json() else f"Config: {_find_opencode_json()}"
lines.append(f"Source: {source}")
lines.append("")
models = candidate.get("models", [])
if not models:
lines.append("No models detected.")
else:
lines.append(f"Detected {len(models)} model(s):")
for m in models:
loc = m.get("location", "unknown")
prov = m.get("provider", "?")
ctx = m.get("context_window")
ctx_str = f", context: {ctx}" if ctx else ""
lines.append(f" - {m['name']} ({prov}, {loc}{ctx_str})")
lines.append("")
lines.append(f"Default: {candidate.get('default', 'none')}")
lines.append(f"Advised: {candidate.get('advised', False)}")
lines.append(f"Mode: {'multi-LLM' if len(models) >= 2 else 'single-LLM'}")
lines.append("")
target = AUTOMATON_DIR / "models.json"
if target.exists():
lines.append(f"models.json already exists at {target} (use --force to overwrite)")
else:
lines.append(f"Ready to write to {target} (use --write to create)")
return "\n".join(lines)
def main() -> int:
import argparse
parser = argparse.ArgumentParser(description="Detect available LLM models and write models.json")
parser.add_argument("--json", action="store_true", help="Output candidate JSON on last line")
parser.add_argument("--write", action="store_true", help="Write candidate models.json to ~/.automaton/ (idempotent)")
parser.add_argument("--force", action="store_true", help="Overwrite existing models.json")
parser.add_argument("--no-probe", action="store_true", help="Skip localhost endpoint probing")
args = parser.parse_args()
candidate = build_candidate_models(probe_local=not args.no_probe)
if args.write:
written = write_models_file(candidate, force=args.force)
if written:
print(f"Written models.json to {AUTOMATON_DIR / 'models.json'}")
else:
print(f"Skipped: {AUTOMATON_DIR / 'models.json'} already exists (use --force to overwrite)")
if args.json:
print(json.dumps(candidate))
else:
print(format_report(candidate))
return 0
if __name__ == "__main__":
sys.exit(main())
+5
View File
@@ -3,9 +3,14 @@
# #
# Usage: bash ~/.automaton/scripts/install-hooks.sh [project-path] # Usage: bash ~/.automaton/scripts/install-hooks.sh [project-path]
# #
# Called automatically by onboard-project.sh. Can also be run manually
# after framework updates to refresh hooks.
#
# Installs pre-commit and pre-push hooks. The pre-commit hook blocks # Installs pre-commit and pre-push hooks. The pre-commit hook blocks
# commits when no task is in implement/doc_review. The pre-push hook # commits when no task is in implement/doc_review. The pre-push hook
# blocks pushes in the same condition, catching --no-verify bypasses. # blocks pushes in the same condition, catching --no-verify bypasses.
#
# Next step: python3 ~/.automaton/scripts/status.py --create-task --project .
set -euo pipefail set -euo pipefail
+34 -28
View File
@@ -3,27 +3,39 @@ set -e
FRAMEWORK_DIR="$HOME/.automaton" FRAMEWORK_DIR="$HOME/.automaton"
if [ -d "$FRAMEWORK_DIR" ]; then
echo "automaton already installed at $FRAMEWORK_DIR"
echo "Run './update.sh' to update."
exit 0
fi
GIT_URL="${1:-}" GIT_URL="${1:-}"
if [ -z "$GIT_URL" ]; then
echo "ERROR: Git URL required." if [ ! -d "$FRAMEWORK_DIR" ]; then
echo "Usage: ./install.sh <git-url>" # Fresh install — need a Git URL to clone
echo "Example: ./install.sh https://github.com/user/automaton.git" if [ -z "$GIT_URL" ]; then
echo "ERROR: Git URL required for fresh install."
echo ""
echo "Usage:"
echo " curl -fsSL https://raw.githubusercontent.com/<user>/automaton/main/scripts/install.sh | bash -s -- <git-url>"
echo ""
echo " --- or ---"
echo ""
echo " git clone <git-url> ~/.automaton"
echo " bash ~/.automaton/scripts/install.sh"
echo ""
echo "The framework is cloned to ~/.automaton. Choose your URL carefully"
echo "as it cannot be changed later without reinstalling."
exit 1
fi
echo "Cloning automaton from $GIT_URL to $FRAMEWORK_DIR..."
git clone "$GIT_URL" "$FRAMEWORK_DIR"
echo ""
else
echo "automaton already installed at $FRAMEWORK_DIR — running setup steps..."
if [ -n "$GIT_URL" ]; then
echo "Note: Git URL argument ignored because ~/.automaton already exists."
echo "To update, run: cd ~/.automaton && ./update.sh"
fi
echo "" echo ""
echo "The framework is cloned to ~/.automaton. Choose your URL carefully"
echo "as it cannot be changed later without reinstalling."
exit 1
fi fi
echo "Cloning automaton from $GIT_URL to $FRAMEWORK_DIR..." # --- Everything below is idempotent and runs on both fresh and existing installs ---
git clone "$GIT_URL" "$FRAMEWORK_DIR"
echo ""
echo "=== VRAM / Context Detection ===" echo "=== VRAM / Context Detection ==="
echo "Detecting your system's VRAM to recommend task decomposition settings..." echo "Detecting your system's VRAM to recommend task decomposition settings..."
echo "" echo ""
@@ -31,7 +43,7 @@ echo ""
# Run VRAM detection script if it exists # Run VRAM detection script if it exists
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
# Run in project-dir context so it can read framework overhead # Run in project-dir context so it can read framework overhead
detection_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>&1) detection_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>&1 || true)
# Extract JSON output (the block after "=== JSON Output ===") # Extract JSON output (the block after "=== JSON Output ===")
json_output=$(echo "$detection_output" | sed -n '/=== JSON Output ===/,$p' | tail -n +2) json_output=$(echo "$detection_output" | sed -n '/=== JSON Output ===/,$p' | tail -n +2)
@@ -40,12 +52,9 @@ if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
echo "$detection_output" echo "$detection_output"
# Extract key values from JSON using Python # Extract key values from JSON using Python
recommended_k=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["recommended_k"])') recommended_k=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["recommended_k"])' 2>/dev/null || echo "?")
max_peak_kb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["max_peak_context_kb"])') max_peak_kb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["max_peak_context_kb"])' 2>/dev/null || echo "?")
headroom=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["headroom"])') headroom=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["headroom"])' 2>/dev/null || echo "?")
gpu_vram=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["gpu_vram_gb"])')
ram_gb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["ram_gb"])')
model_context=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["model_context_kb"])')
echo "" echo ""
echo "=== Recommended VRAM Configuration ===" echo "=== Recommended VRAM Configuration ==="
@@ -72,11 +81,8 @@ echo ""
echo "Installation complete." echo "Installation complete."
echo "" echo ""
echo "Next steps:" echo "Next steps:"
echo " 1. cd into a project and run the onboarding prompt" echo " 1. Onboard a project: bash ~/.automaton/scripts/onboard-project.sh /path/to/project"
echo " 2. In each project that uses git, install the automaton hooks:" echo " 2. Or tell your agent: 'Onboard this project into automaton'"
echo " bash ~/.automaton/scripts/install-hooks.sh /path/to/project"
echo ""
echo "These hooks block commits and pushes when no task is in an edit-allowed phase."
echo "" echo ""
# Register pre-edit guards for detected harnesses # Register pre-edit guards for detected harnesses
+54 -15
View File
@@ -375,6 +375,9 @@ def _invoke_harness(
"""Build the harness command from loop.json and invoke it. Returns stdout. """Build the harness command from loop.json and invoke it. Returns stdout.
extras: substitution tokens specific to this role ({artifact}, {verdict}, etc). extras: substitution tokens specific to this role ({artifact}, {verdict}, etc).
If extras contains a "model" key, the ``{model}`` token in the harness
command is substituted. The caller is responsible for passing the model
via extras (extracted from loop.json role config or manifest default).
""" """
resolved_prompt = prompt_path resolved_prompt = prompt_path
if loop_path is not None: if loop_path is not None:
@@ -524,6 +527,24 @@ def _role_prompt(cfg: dict, role: str) ->Optional[str]:
return role_cfg.get("prompt") return role_cfg.get("prompt")
def _role_model(cfg: dict, role: str) -> Optional[str]:
"""Get the model configured for a role in loop.json, or the manifest default."""
roles = cfg.get("roles") or {}
role_cfg = roles.get(role) or {}
model = role_cfg.get("model")
if model:
return model
models_file = AUTOMATON_DIR / "models.json"
if models_file.exists():
try:
import json as _mj
manifest = _mj.loads(models_file.read_text())
return manifest.get("default")
except (OSError, _mj.JSONDecodeError):
pass
return None
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Work sources (task add-goal-mode) # Work sources (task add-goal-mode)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@@ -809,27 +830,39 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
out_dir = _outputs_dir(loop_path) out_dir = _outputs_dir(loop_path)
tick_num = state.get('iteration_count', 0) + 1 tick_num = state.get('iteration_count', 0) + 1
impl_output = str(out_dir / f"tick{tick_num}-implement.json") impl_output = str(out_dir / f"tick{tick_num}-implement.json")
impl_model = _role_model(cfg, "implement")
implement_extras = {
"output": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint,
}
if impl_model:
implement_extras["model"] = impl_model
implement_stdout = _invoke_harness( implement_stdout = _invoke_harness(
harness_cfg, "implement", implement_prompt, cwd, harness_cfg, "implement", implement_prompt, cwd,
extras={"output": impl_output, extras=implement_extras,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
loop_path=loop_path, tick_num=tick_num) loop_path=loop_path, tick_num=tick_num)
(Path(impl_output)).write_text(implement_stdout) (Path(impl_output)).write_text(implement_stdout)
# Step 6: spawn Verify # Step 6: spawn Verify
verify_prompt = _role_prompt(cfg, "verify") or "" verify_prompt = _role_prompt(cfg, "verify") or ""
verify_output = str(out_dir / f"tick{tick_num}-verify.json") verify_output = str(out_dir / f"tick{tick_num}-verify.json")
verify_model = _role_model(cfg, "verify")
verify_extras = {
"output": verify_output,
"artifact": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint,
}
if verify_model:
verify_extras["model"] = verify_model
verify_stdout = _invoke_harness( verify_stdout = _invoke_harness(
harness_cfg, "verify", verify_prompt, cwd, harness_cfg, "verify", verify_prompt, cwd,
extras={"output": verify_output, extras=verify_extras,
"artifact": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
loop_path=loop_path, tick_num=tick_num) loop_path=loop_path, tick_num=tick_num)
(Path(verify_output)).write_text(verify_stdout) (Path(verify_output)).write_text(verify_stdout)
@@ -854,12 +887,18 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
# Step 9: spawn Orchestrate # Step 9: spawn Orchestrate
orch_prompt = _role_prompt(cfg, "orchestrate") or "" orch_prompt = _role_prompt(cfg, "orchestrate") or ""
orch_output = str(out_dir / f"tick{tick_num}-orchestrate.json") orch_output = str(out_dir / f"tick{tick_num}-orchestrate.json")
orch_model = _role_model(cfg, "orchestrate")
orch_extras = {
"output": orch_output,
"verdict": json.dumps(verdict),
"current_task": current_task,
"current_phase": state.get("current_phase", ""),
}
if orch_model:
orch_extras["model"] = orch_model
orch_stdout = _invoke_harness( orch_stdout = _invoke_harness(
harness_cfg, "orchestrate", orch_prompt, cwd, harness_cfg, "orchestrate", orch_prompt, cwd,
extras={"output": orch_output, extras=orch_extras,
"verdict": json.dumps(verdict),
"current_task": current_task,
"current_phase": state.get("current_phase", "")},
loop_path=loop_path, tick_num=tick_num) loop_path=loop_path, tick_num=tick_num)
(Path(orch_output)).write_text(orch_stdout) (Path(orch_output)).write_text(orch_stdout)
+184
View File
@@ -0,0 +1,184 @@
#!/usr/bin/env bash
# onboard-project.sh — Bootstrap automaton in a new or existing project.
#
# Usage:
# bash ~/.automaton/scripts/onboard-project.sh /path/to/project
#
# Creates .automaton/ skeleton, detects models, creates config, inits git,
# installs hooks, and verifies everything works.
set -euo pipefail
FRAMEWORK_DIR="$HOME/.automaton"
PROJECT_DIR="${1:-}"
# Colors
RED='\033[0;31m'
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
BLUE='\033[0;34m'
NC='\033[0m' # No Color
info() { echo -e "${BLUE}INFO:${NC} $1"; }
ok() { echo -e "${GREEN}OK:${NC} $1"; }
warn() { echo -e "${YELLOW}WARN:${NC} $1"; }
error() { echo -e "${RED}ERROR:${NC} $1"; }
# --- Argument checks ---
if [ -z "$PROJECT_DIR" ]; then
error "Usage: bash ~/.automaton/scripts/onboard-project.sh /path/to/project"
exit 1
fi
PROJECT_DIR="$(cd "$PROJECT_DIR" 2>/dev/null && pwd)" || true
if [ -z "$PROJECT_DIR" ] || [ ! -d "$PROJECT_DIR" ]; then
echo ""
error "'$1' does not exist."
echo " Create it first: mkdir -p '$1'"
echo " Then re-run this script."
exit 1
fi
if [ ! -d "$FRAMEWORK_DIR/scripts" ]; then
error "Framework not found at $FRAMEWORK_DIR."
echo " Install the framework first:"
echo " curl -fsSL https://raw.githubusercontent.com/<user>/automaton/main/scripts/install.sh | bash -s -- <git-url>"
echo " Or: git clone <git-url> ~/.automaton && bash ~/.automaton/scripts/install.sh"
exit 1
fi
echo ""
echo "========================================"
echo " Automaton Project Onboarding"
echo " Project: $PROJECT_DIR"
echo "========================================"
echo ""
# --- Step 1: Create .automaton/ skeleton ---
AUTO_DIR="$PROJECT_DIR/.automaton"
if [ -d "$AUTO_DIR" ]; then
warn "$AUTO_DIR already exists — skipping skeleton creation"
else
info "Creating .automaton/ skeleton..."
mkdir -p "$AUTO_DIR/tasks" "$AUTO_DIR/loops" "$AUTO_DIR/design"
ok "Created $AUTO_DIR/"
fi
# --- Step 2: Models ---
MODELS_FILE="$AUTO_DIR/models.json"
if [ -f "$MODELS_FILE" ]; then
warn "$MODELS_FILE already exists — skipping model detection"
else
info "Probing local models..."
if python3 "$FRAMEWORK_DIR/scripts/detect_models.py" --write --project "$PROJECT_DIR" 2>/dev/null; then
ok "Detected models written to $MODELS_FILE"
else
info "Auto-detection failed. Creating minimal models.json..."
cat > "$MODELS_FILE" <<- 'EOF'
{
"models": [
{"name": "default-model", "provider": "local", "context": 32768}
],
"default": "default-model"
}
EOF
warn "Edit $MODELS_FILE to set your actual model(s)."
fi
fi
# --- Step 3: Config ---
CONFIG_FILE="$AUTO_DIR/config.md"
if [ -f "$CONFIG_FILE" ]; then
warn "$CONFIG_FILE already exists — skipping"
else
info "Creating config.md..."
PROJECT_NAME="$(basename "$PROJECT_DIR")"
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
vram_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>/dev/null || true)
json_part=$(echo "$vram_output" | sed -n '/=== JSON Output ===/,$p' | tail -n +2)
recommended=$(echo "$json_part" | python3 -c 'import json,sys; print(json.load(sys.stdin).get("recommended_k","16"))' 2>/dev/null || echo "16")
else
recommended="16"
fi
cat > "$CONFIG_FILE" <<- EOF
# $PROJECT_NAME — Automaton Configuration
## VRAM Configuration
- **Auto-detect**: Yes
- **Target context**: ${recommended}k tokens
- **Headroom**: 25%
- **Max peak context per sub-task**: $((recommended * 3 / 4))k tokens
## Model Configuration
# Uses models.json for model divergence enforcement.
# Default model is read from models.json's "default" key.
EOF
ok "Created $CONFIG_FILE"
fi
# --- Step 4: Project name ---
NAME_FILE="$AUTO_DIR/project-name.md"
if [ -f "$NAME_FILE" ]; then
warn "$NAME_FILE already exists — skipping"
else
PROJECT_NAME="$(basename "$PROJECT_DIR")"
echo "$PROJECT_NAME" > "$NAME_FILE"
ok "Created $NAME_FILE ($PROJECT_NAME)"
fi
# --- Step 5: Git ---
GIT_DIR="$PROJECT_DIR/.git"
if [ -d "$GIT_DIR" ]; then
ok "Git repository already initialized"
else
info "Initializing git repository..."
cd "$PROJECT_DIR" && git init
ok "Git initialized"
fi
# --- Step 6: Git hooks ---
if [ -d "$GIT_DIR" ]; then
info "Installing git hooks..."
bash "$FRAMEWORK_DIR/scripts/install-hooks.sh" "$PROJECT_DIR"
fi
# --- Step 7: .gitignore ---
GITIGNORE="$PROJECT_DIR/.gitignore"
if [ -f "$GITIGNORE" ]; then
if ! grep -q ".automaton/tasks/" "$GITIGNORE" 2>/dev/null; then
echo "" >> "$GITIGNORE"
echo "# Automaton" >> "$GITIGNORE"
echo ".automaton/tasks/" >> "$GITIGNORE"
echo ".automaton/loops/*/worktree/" >> "$GITIGNORE"
warn "Added automaton entries to .gitignore"
fi
else
cat > "$GITIGNORE" <<- 'EOF'
# Automaton
.automaton/tasks/
.automaton/loops/*/worktree/
.automaton/loops/*/outputs/
EOF
ok "Created .gitignore with automaton entries"
fi
# --- Step 8: Verify ---
info "Verifying setup..."
cd "$FRAMEWORK_DIR"
python3 "$FRAMEWORK_DIR/scripts/status.py" --project "$PROJECT_DIR" --audit 2>&1 | head -5 || true
echo ""
echo "========================================"
echo -e "${GREEN} Onboarding complete!${NC}"
echo "========================================"
echo ""
echo " Project: $PROJECT_DIR"
echo " Config: $CONFIG_FILE"
echo " Models: $MODELS_FILE"
echo ""
echo " Next steps:"
echo " 1. cd $PROJECT_DIR"
echo " 2. Create a task:"
echo " python3 ~/.automaton/scripts/status.py --create-task my-first-task --project ."
echo " 3. Start working with your agent."
echo ""
+174
View File
@@ -0,0 +1,174 @@
#!/usr/bin/env bash
# pi-automaton.sh — Pi Dev automaton context printer.
#
# Prints automaton project/framework context for the user to paste as their
# first message to a Pi Dev agent. Does NOT launch pi.
#
# Usage:
# cd /path/to/project
# bash ~/.automaton/scripts/pi-automaton.sh
#
# Copy the output and paste it as your first message in Pi Dev.
set -euo pipefail
FRAMEWORK_DIR="$HOME/.automaton"
STATUS_PY="$FRAMEWORK_DIR/scripts/status.py"
# Colors
BOLD='\033[1m'
DIM='\033[2m'
NC='\033[0m'
info() { echo -e " $1"; }
dim() { echo -e " ${DIM}$1${NC}"; }
dim_nl(){ echo -e "${DIM}$1${NC}"; }
# --- Scope detection ---
CWD="$(pwd)"
if [ "$CWD" = "$FRAMEWORK_DIR" ] || [ "${CWD##"$FRAMEWORK_DIR"}" != "$CWD" ]; then
SCOPE="framework"
SCOPE_LABEL="framework mode (automaton itself)"
PROJECT_DIR="$FRAMEWORK_DIR"
elif [ -d "$CWD/.automaton" ]; then
SCOPE="project"
SCOPE_LABEL="project mode ($(basename "$CWD"))"
PROJECT_DIR="$CWD"
else
echo ""
echo "No automaton project detected in $CWD"
echo ""
echo "To onboard this project:"
echo " bash $FRAMEWORK_DIR/scripts/onboard-project.sh ."
echo ""
exit 1
fi
TASKS_DIR="$PROJECT_DIR/.automaton/tasks"
MODELS_FILE="$PROJECT_DIR/.automaton/models.json"
CONFIG_FILE="$PROJECT_DIR/.automaton/config.md"
AGENTS_FILE="$PROJECT_DIR/.automaton/AGENTS.md"
# --- Collect data ---
# Active tasks via status.py --audit --json
TASKS_JSON=""
if [ -f "$STATUS_PY" ] && [ -d "$TASKS_DIR" ]; then
TASKS_JSON=$(python3 "$STATUS_PY" --audit --json --project "$PROJECT_DIR" 2>/dev/null || true)
fi
# Default model
DEFAULT_MODEL=""
if [ -f "$MODELS_FILE" ]; then
DEFAULT_MODEL=$(python3 -c "
import json
with open('$MODELS_FILE') as f:
m = json.load(f)
print(m.get('default', ''))
" 2>/dev/null || echo "")
fi
# VRAM context snippet from config.md
CONFIG_SNIPPET=""
if [ -f "$CONFIG_FILE" ]; then
CONFIG_SNIPPET=$(grep -i 'context\|headroom\|target' "$CONFIG_FILE" 2>/dev/null | head -3 | sed 's/^/ /')
fi
# Project rules from AGENTS.md
RULES_TEXT=""
if [ -f "$AGENTS_FILE" ]; then
RULES_TEXT=$(grep -v -E '^#|^$' "$AGENTS_FILE" 2>/dev/null | head -10)
fi
# Can-edit status
CAN_EDIT_OUTPUT=""
if [ -f "$STATUS_PY" ]; then
CAN_EDIT_OUTPUT=$(python3 "$STATUS_PY" --can-edit --json --project "$PROJECT_DIR" 2>/dev/null || true)
fi
# --- Build task list ---
TASK_LINES=""
TASK_COUNT=0
if [ -n "$TASKS_JSON" ]; then
while IFS=$'\t' read -r name state edit_phase; do
if [ -n "$name" ]; then
FLAG=""
if [ "$edit_phase" = "true" ]; then
FLAG=" (edit allowed)"
elif [ "$state" != "backlog" ] && [ "$state" != "done" ] && [ "$state" != "blocked" ]; then
FLAG=" (read-only)"
fi
TASK_LINES+=" * $name\t\t$state$FLAG\n"
TASK_COUNT=$((TASK_COUNT + 1))
fi
done < <(echo "$TASKS_JSON" | python3 -c "
import json, sys
data = json.load(sys.stdin)
tasks = []
for t in data.get('tasks', []):
tasks.append((t['name'], t['state']))
# Map edit-eligible phases
EDIT_PHASES = {'implement', 'doc_review'}
for name, state in tasks:
ep = 'true' if state in EDIT_PHASES else 'false'
print(f'{name}\t{state}\t{ep}')
" 2>/dev/null || true)
fi
# --- Render ---
LINE="══════════════════════════════════════════════════"
SEP="──────────────────────────────────────────────────"
echo ""
echo -e "${BOLD}${LINE}${NC}"
echo -e "${BOLD} Automaton Context — $SCOPE_LABEL${NC}"
echo -e "${BOLD}${LINE}${NC}"
echo ""
echo -e "${BOLD}This project uses the automaton workflow framework.${NC}"
info "Tasks are tracked in $(basename "$PROJECT_DIR")/.automaton/tasks/"
info "and flow through phases:"
info "backlog → research → implement → code_review → bug_find → ... → complete"
echo ""
if [ -n "$TASK_LINES" ]; then
echo -e "${BOLD}Active tasks:${NC}"
echo -e "$TASK_LINES"
echo ""
fi
if [ -n "$DEFAULT_MODEL" ]; then
echo -e "${BOLD}Default model:${NC} $DEFAULT_MODEL"
echo ""
fi
if [ -n "$CONFIG_SNIPPET" ]; then
echo -e "${BOLD}Configuration:${NC}"
echo "$CONFIG_SNIPPET"
echo ""
fi
echo -e "${BOLD}To work on a task:${NC}"
info "python3 ~/.automaton/scripts/status.py --transition <phase> --task <name>"
echo ""
echo -e "${BOLD}To create a new task:${NC}"
info "python3 ~/.automaton/scripts/status.py --create-task <name>"
echo ""
if [ -n "$RULES_TEXT" ]; then
echo -e "${BOLD}Project rules (from AGENTS.md):${NC}"
echo "$RULES_TEXT" | head -5
echo ""
fi
echo -e "${BOLD}Important:${NC}"
info "Only modify files when a task is in ${BOLD}implement${NC} or ${BOLD}doc_review${NC} phase"
info "All phase transitions go through status.py"
info "The automaton-guard-pi plugin blocks edits outside allowed phases"
echo ""
echo -e "${DIM}$SEP${NC}"
dim_nl "Copy this entire block and paste it as your first message"
dim_nl "to the Pi Dev agent to provide automaton context."
echo -e "${DIM}$SEP${NC}"
echo ""
+44 -11
View File
@@ -54,17 +54,50 @@ else
echo "OpenCode: not detected (no ~/.config/opencode/opencode.json or .jsonc)" echo "OpenCode: not detected (no ~/.config/opencode/opencode.json or .jsonc)"
fi fi
# Pi Dev guard # Pi Dev extensions
PI_SOURCE="$FRAMEWORK_DIR/plugins/automaton-guard-pi" PI_EXTENSIONS=(
"$FRAMEWORK_DIR/plugins/automaton-guard-pi"
"$FRAMEWORK_DIR/plugins/pi-read-tweet"
"$FRAMEWORK_DIR/plugins/pi-automaton-context"
"$FRAMEWORK_DIR/plugins/pi-automaton-tools"
)
if command -v pi &>/dev/null; then if command -v pi &>/dev/null; then
INSTALLED=$(pi list 2>/dev/null | grep -c "automaton-guard-pi" || true) INSTALLED_PI=true
if [ "$INSTALLED" -gt 0 ]; then for ext in "${PI_EXTENSIONS[@]}"; do
echo "Pi Dev: already installed" ext_name=$(basename "$ext")
else FOUND=$(pi list 2>/dev/null | grep -c "$ext_name" || true)
echo "Pi Dev: installing guard extension..." if [ "$FOUND" -gt 0 ]; then
pi install "$PI_SOURCE" 2>&1 | sed 's/^/ /' echo "Pi Dev: $ext_name already installed"
INSTALLED_PI=true else
echo "Pi Dev: installed" echo "Pi Dev: installing $ext_name..."
pi install "$ext" 2>&1 | sed 's/^/ /'
echo "Pi Dev: $ext_name installed"
fi
done
# Offer pi-automaton startup wrapper
PI_WRAPPER="$FRAMEWORK_DIR/scripts/pi-automaton.sh"
if [ -f "$PI_WRAPPER" ]; then
echo ""
echo "Pi Dev context wrapper available at:"
echo " $PI_WRAPPER"
echo ""
echo "Before starting a Pi Dev session, run this script to print"
echo "automaton project context that you can paste as your first message."
echo ""
echo " bash ~/.automaton/scripts/pi-automaton.sh"
echo ""
# Only prompt interactively if stdin is a terminal
BIN_DIR="$HOME/bin"
if [ -t 0 ] && [ ! -f "$BIN_DIR/pi-automaton" ]; then
echo -n "Symlink to ~/bin/pi-automaton for easier access? [Y/n] "
read -r REPLY
if [ -z "$REPLY" ] || [ "$REPLY" = "y" ] || [ "$REPLY" = "Y" ]; then
mkdir -p "$BIN_DIR"
ln -sf "$PI_WRAPPER" "$BIN_DIR/pi-automaton"
echo " Created $BIN_DIR/pi-automaton → $PI_WRAPPER"
echo " (ensure ~/bin is in your PATH)"
fi
fi
fi fi
else else
echo "Pi Dev: not detected (pi not in PATH)" echo "Pi Dev: not detected (pi not in PATH)"
@@ -77,7 +110,7 @@ fi
if ! $INSTALLED_OPENCODE && ! $INSTALLED_PI; then if ! $INSTALLED_OPENCODE && ! $INSTALLED_PI; then
echo "No harness detected. To install a guard manually:" echo "No harness detected. To install a guard manually:"
echo " OpenCode: add '\"plugin\": [\"$OPENCODE_SOURCE\"]' to ~/.config/opencode/opencode.json" echo " OpenCode: add '\"plugin\": [\"$OPENCODE_SOURCE\"]' to ~/.config/opencode/opencode.json"
echo " Pi Dev: pi install $PI_SOURCE" echo " Pi Dev: pi install $FRAMEWORK_DIR/plugins/automaton-guard-pi"
echo "" echo ""
echo "Without a pre-edit guard, git hooks (pre-commit + pre-push)" echo "Without a pre-edit guard, git hooks (pre-commit + pre-push)"
echo "provide enforcement at commit/push time instead." echo "provide enforcement at commit/push time instead."
+261 -2
View File
@@ -153,7 +153,19 @@ FORBIDDEN_ARTIFACTS = {
} }
NON_ARTIFACT_FILES = {".state", ".state.tmp", ".state.lock", ".state.approvals", NON_ARTIFACT_FILES = {".state", ".state.tmp", ".state.lock", ".state.approvals",
".state.implementer", ".state.lastedit", "VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"} ".state.implementer", ".state.lastedit", ".state.models",
"VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"}
# Model-divergence enforcement
CONFLICT_MATRIX = {
"code_review": {"implement"},
"bug_find": {"implement"},
"adversarial_bug_find": {"implement", "bug_find"},
"referee": {"implement", "bug_find", "adversarial_bug_find"},
"loop-verify": {"loop-implement"},
}
MODELS_JSON_FILE = "models.json"
PHASE_PRIORITY = { PHASE_PRIORITY = {
"referee": 12, "doc_review": 11, "adversarial_bug_find": 10, "referee": 12, "doc_review": 11, "adversarial_bug_find": 10,
@@ -438,6 +450,124 @@ def _lock_timeout_seconds(project: Optional[str] = None) -> int:
return val return val
# ---------------------------------------------------------------------------
# Model-divergence enforcement helpers
# ---------------------------------------------------------------------------
def _load_models_manifest(project: Optional[str] = None) -> Optional[dict]:
"""Load the models.json manifest for the given project.
Searches:
1. project/.automaton/models.json
2. ~/.automaton/models.json (fallback)
Returns None if no models.json exists (single-LLM mode, backward compatible).
"""
project_dir = _find_project_dir(project)
candidates = [
project_dir / ".automaton" / MODELS_JSON_FILE,
AUTOMATON_DIR / MODELS_JSON_FILE,
]
for path in candidates:
if path.exists():
try:
return json.loads(path.read_text())
except (OSError, json.JSONDecodeError):
return None
return None
def _get_model_mode(manifest: Optional[dict]) -> str:
"""Determine the model mode: 'single' or 'multi-llm'.
- Missing manifest → single-LLM (backward compatible)
- 0-1 models → single-LLM
- 2+ models → multi-LLM
"""
if manifest is None:
return "single"
models = manifest.get("models") or []
if len(models) >= 2:
return "multi-llm"
return "single"
def _check_conflict(state_models: dict, role: str, model: str, matrix: Optional[dict] = None) -> Optional[str]:
"""Check if the given model conflicts with already-filled roles.
state_models: dict of {role: model_name} from .state.models
role: the role being entered (e.g. 'code_review')
model: the model name being assigned
matrix: conflict matrix (defaults to CONFLICT_MATRIX)
Returns the name of the conflicting role, or None if no conflict.
"""
if matrix is None:
matrix = CONFLICT_MATRIX
if role not in matrix:
return None
conflicting_roles = matrix[role]
for filled_role, filled_model in state_models.items():
if filled_model == model and filled_role in conflicting_roles:
return filled_role
return None
def _read_state_models(task_path: Path) -> dict:
"""Read .state.models from the task directory. Returns {} if missing."""
f = task_path / ".state.models"
if not f.exists():
return {}
try:
data = json.loads(f.read_text())
if isinstance(data, dict):
return data
except (OSError, json.JSONDecodeError):
pass
return {}
def _write_state_models(task_path: Path, state_models: dict) -> None:
"""Write .state.models atomically."""
tmp = task_path / ".state.models.tmp"
tmp.write_text(json.dumps(state_models, indent=2, sort_keys=True) + "\n")
tmp.replace(task_path / ".state.models")
def _model_divergence_violations(project: Optional[str] = None) -> list[dict]:
"""Scan all tasks for model-divergence violations.
Returns list of violation dicts:
{"task": str, "message": str, "severity": "high", "resolved": False}
"""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode == "single":
return []
violations = []
tasks = _all_task_dirs(project)
for name, path in tasks:
sm = _read_state_models(path)
if not sm:
continue
for role, model in sm.items():
if model is None:
continue
# D8: doc_review, code_review, bug_find have no cross-conflicts
# with each other; only conflicts documented in CONFLICT_MATRIX apply.
conflict = _check_conflict(sm, role, str(model))
if conflict:
violations.append({
"task": name,
"severity": "high",
"message": f"model-divergence: role '{role}' uses model '{model}' "
f"which conflicts with role '{conflict}' (same model)",
"resolved": False,
})
return violations
# --- Command implementations --- # --- Command implementations ---
def cmd_show_task(args): def cmd_show_task(args):
@@ -593,6 +723,60 @@ def cmd_transition(args):
return 1 return 1
if current == "human_intervention" and target == "complete": if current == "human_intervention" and target == "complete":
_auto_update_verdict_on_complete(task_path) _auto_update_verdict_on_complete(task_path)
# Model-divergence enforcement (Subtask 2)
# When entering a phase that maps to a role, record the model
target_base = _base_phase(target)
ROLE_PHASES = {"implement", "code_review", "bug_find", "adversarial_bug_find", "doc_review", "referee"}
if target_base in ROLE_PHASES and current != target:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
state_models = _read_state_models(task_path)
model_arg = getattr(args, "model", None)
if model_arg:
# --model explicitly provided — record advisory in single mode, check in multi
state_models[target_base] = model_arg
if mode == "multi-llm":
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 1
conflict = _check_conflict(state_models, target_base, model_arg)
if conflict:
# Remove the entry we just added
del state_models[target_base]
print(f"ERROR: Model '{model_arg}' assigned to role '{target_base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model> to specify a different model.")
return 1
elif mode == "multi-llm":
# Auto-assign: try default, then next-available non-conflicting
default = (manifest or {}).get("default")
assigned = False
if default and default in model_names:
conflict = _check_conflict(state_models, target_base, default)
if not conflict:
state_models[target_base] = default
assigned = True
if not assigned:
for m_name in model_names:
if m_name == default:
continue
conflict = _check_conflict(state_models, target_base, m_name)
if not conflict:
state_models[target_base] = m_name
assigned = True
break
if not assigned:
print(f"ERROR: Cannot auto-assign a model for role '{target_base}'. "
f"All available models conflict with already-filled roles. "
f"Use --model <name> to override.")
return 1
if model_arg or mode == "multi-llm":
_write_state_models(task_path, state_models)
if current == "implement" and target == "code_review": if current == "implement" and target == "code_review":
lock_file = task_path / ".state.lock" lock_file = task_path / ".state.lock"
if lock_file.exists(): if lock_file.exists():
@@ -980,6 +1164,14 @@ def _audit_collect(args):
"halt_reason": lhalt, "current_task": ltask, "halt_reason": lhalt, "current_task": ltask,
"violation": is_violation, "message": msg}) "violation": is_violation, "message": msg})
# Model-divergence violations (Category 6)
for mv in _model_divergence_violations(args.project):
violations.append({
"category": 6, "severity": mv["severity"],
"task": mv["task"], "message": mv["message"],
"resolved": False,
})
return {"violations": violations, return {"violations": violations,
"loops": loops, "loops": loops,
"total_tasks": len(tasks), "total_tasks": len(tasks),
@@ -1118,7 +1310,16 @@ def cmd_audit(args):
if stuck_found == 0: if stuck_found == 0:
print(f"[PASS] No stuck tasks (threshold: {stuck_threshold} min)") print(f"[PASS] No stuck tasks (threshold: {stuck_threshold} min)")
print("\n=== Category 6: Loops ===") print("\n=== Category 6: Model-Divergence Violations ===")
md_violations = _model_divergence_violations(args.project)
if md_violations:
for v in md_violations:
print(f"[FAIL] {v['task']}: {v['message']}")
violations += 1
else:
print("[PASS] No model-divergence violations found")
print("\n=== Category 7: Loops ===")
violations += _audit_loops_block(args) violations += _audit_loops_block(args)
print(f"\n=== Summary ===") print(f"\n=== Summary ===")
@@ -1435,6 +1636,26 @@ def cmd_claim(args):
if implementer == args.agent: if implementer == args.agent:
print(f"ERROR: Agent '{args.agent}' implemented this task and cannot claim the code_review phase. Reviewer must be different from implementer.") print(f"ERROR: Agent '{args.agent}' implemented this task and cannot claim the code_review phase. Reviewer must be different from implementer.")
return 1 return 1
# Model-divergence check on claim (Subtask 2, multi-LLM only)
model_arg = getattr(args, "model", None)
if model_arg:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
if mode == "multi-llm":
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 2
state_models = _read_state_models(task_path)
conflict = _check_conflict(state_models, base, model_arg)
if conflict:
print(f"ERROR: Model '{model_arg}' for role '{base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model>.")
return 1
lock_file = task_path / ".state.lock" lock_file = task_path / ".state.lock"
timeout_sec = _lock_timeout_seconds(args.project) timeout_sec = _lock_timeout_seconds(args.project)
if lock_file.exists(): if lock_file.exists():
@@ -2600,6 +2821,42 @@ def _gate_worktree_drift(state: dict, cfg: dict, project: Optional[str]) -> Opti
return None return None
def _gate_model_divergence(state: dict, cfg: dict, project: Optional[str]) -> Optional[dict]:
"""Model-divergence brake: in multi-LLM mode, verify and implement
roles must use different models. This prevents same-model verification
(rubber-stamping) within a loop tick."""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode != "multi-llm":
return None
roles = cfg.get("roles") or {}
impl_model = None
verify_model = None
impl_cfg = roles.get("implement") or {}
verify_cfg = roles.get("verify") or {}
impl_model = impl_cfg.get("model")
verify_model = verify_cfg.get("model")
# Fall back to manifest default if role has no explicit model
if not impl_model or not verify_model:
default = (manifest or {}).get("default")
if not impl_model:
impl_model = default
if not verify_model:
verify_model = default
if impl_model and verify_model and impl_model == verify_model:
return {
"ok": False,
"reason": "halted:model_conflict",
"halt_reason": "human_intervention",
"remaining_iterations": None,
"remaining_budget_usd": None,
"task_phase": None,
"task_in_halt_loop": True,
"out_of_scope_files": [],
}
return None
def _gate_score_plateau(state: dict, cfg: dict) -> Optional[dict]: def _gate_score_plateau(state: dict, cfg: dict) -> Optional[dict]:
window = int(cfg.get("brakes", {}).get("score_plateau_window", 0)) window = int(cfg.get("brakes", {}).get("score_plateau_window", 0))
if window <= 0: if window <= 0:
@@ -2648,6 +2905,7 @@ def cmd_check_gate(args) -> int:
_gate_task_phase(state, cfg, args.project), _gate_task_phase(state, cfg, args.project),
_gate_worktree_drift(state, cfg, args.project), _gate_worktree_drift(state, cfg, args.project),
_gate_score_plateau(state, cfg), _gate_score_plateau(state, cfg),
_gate_model_divergence(state, cfg, args.project),
] ]
failure = next((g for g in gates if g is not None), None) failure = next((g for g in gates if g is not None), None)
if failure is None: if failure is None:
@@ -2864,6 +3122,7 @@ def main():
parser.add_argument("--days", type=int, help="Cleanup age threshold in days (default 7, used with --cleanup-done / --install-cleanup-schedule)") parser.add_argument("--days", type=int, help="Cleanup age threshold in days (default 7, used with --cleanup-done / --install-cleanup-schedule)")
parser.add_argument("--dry-run", action="store_true", help="With --cleanup-done, list candidates without moving them") parser.add_argument("--dry-run", action="store_true", help="With --cleanup-done, list candidates without moving them")
parser.add_argument("--version", action="store_true", help="Print framework version and exit") parser.add_argument("--version", action="store_true", help="Print framework version and exit")
parser.add_argument("--model", metavar="NAME", help="Model name for model-divergence enforcement (used with --transition, --claim)")
args = parser.parse_args() args = parser.parse_args()

Some files were not shown because too many files have changed in this diff Show More