Compare commits

..
7 Commits
Author SHA1 Message Date
Lap Tran 287cfeff7c feat(pi): 3 new Pi Dev extensions — read-tweet, auto-context, task-lifecycle tools
CI / build (push) Has been cancelled
2026-06-26 22:40:49 -04:00
Lap Tran c2355954b9 feat(pi): pi-automaton.sh startup context printer for Pi Dev
Pi Dev doesn't auto-load automaton's system prompt (unlike opencode),
so the agent has no awareness of tasks, phases, or status.py commands.
This bridges that gap:

- scripts/pi-automaton.sh: NEW — detects scope (framework/project),
  reads active tasks, default model, config snippet, AGENTS.md rules,
  and prints a formatted context block the user pastes as their first
  message to the Pi Dev agent. Does NOT launch pi.
- scripts/register-guards.sh: EXTENDED — after Pi Dev guard install,
  offers to symlink pi-automaton.sh to ~/bin/pi-automaton
  (interactive prompt only when stdin is a terminal).
- README.md: added 'Using Pi Dev with automaton' FAQ subsection
  documenting the context printer workflow.
2026-06-26 20:47:08 -04:00
Lap Tran 0437cbae6c docs: architecture section, FAQ, cross-references, vault-memory update
README.md:
  - Added §1.5 'Architecture: Framework vs Project' with directory tree
    and lifecycle flow diagram
  - Added §2.5 FAQ covering coexistence, hooks, loops, multi-project
  - Updated install section already done in prior commit

AGENTS.md:
  - Added 'Script Cross-References' table mapping install.sh →
    onboard-project.sh → status.py --create-task

scripts/install-hooks.sh:
  - Updated header to reference onboard-project.sh as caller
  - Added 'Next step' line pointing to --create-task

vault-memory CONTEXT.md:
  - Updated test count (518→611)
  - Added new scripts (detect_models.py, onboard-project.sh)
  - Documented install/onboard flow and split architecture
2026-06-26 16:51:39 -04:00
Lap Tran b880f2535a fix(install): make install.sh idempotent + add onboard-project.sh
- install.sh: no longer exits early when ~/.automaton exists.
  Skips the clone but runs all setup (VRAM detection, guards,
  self-improvement loop, virtualenv). Both curl|bash and
  git-clone + ./install.sh now work correctly.
- onboard-project.sh: new script that bootstraps automaton in a
  project — creates .automaton/ skeleton, detects models, writes
  config.md, inits git, installs hooks, adds .gitignore entries.
- README.md: fix install flow docs (curl|bash + clone-then-run),
  add onboard-project.sh as Option A for project setup
2026-06-26 13:54:28 -04:00
Lap Tran f980ccfe27 feat(dashboard): model badges on kanban cards + design doc update
- task.py: Task dataclass gains models: dict[str, str], loaded from
  .state.models in discover_tasks()
- app.py: models dict included in all task API responses
- dashboard.js: model badges rendered between artifacts and subtask
  progress on kanban cards; ROLE_LABELS map for readable tooltips
- styles.css: .task-card-models and .model-badge styles
- design/loops/technical.md: document {model} substitution token,
  per-role model field, and model-divergence brake gate (gate #7)
2026-06-26 13:26:00 -04:00
Lap Tran 35e449b03e feat(model-divergence): full enforcement — manifest, transition, claim, audit, loop gates, detect script
Completes all 3 model-divergence enforcement subtasks:

- scripts/detect_models.py: probes opencode.json + localhost endpoints,
  builds models.json with --json/--write/--force
- scripts/status.py: CONFLICT_MATRIX, --model flag, --transition --model,
  --claim --model, --audit Category 6, model-divergence brake gate in
  --check-gate, helpers for manifest loading and conflict checking
- scripts/loop-runner.py: _role_model() helper + {model} passed via extras
  dict to _invoke_harness for implement, verify, orchestrate roles
- tests/test_model_divergence.py: 33 tests covering all enforcement layers
- Single-LLM mode: record model advisory, no conflict check
- Multi-LLM mode (2+ models): conflict matrix enforced at transition, claim,
  and loop brake gate
- Project-level models.json preferred over global ~/.automaton/models.json
2026-06-26 13:23:17 -04:00
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00
683 changed files with 2738 additions and 145 deletions
+1
View File
@@ -3,3 +3,4 @@ __pycache__/
*.egg-info/
.venv/
venv/
logs/
+15
View File
@@ -136,6 +136,21 @@ Modes:
2. Run `python3 -m pytest tests/test_prompt_paths.py` to ensure task paths are canonical.
3. Update `CHANGELOG.md` under `[unreleased]`.
## Script Cross-References
The framework provides three lifecycle scripts that should be referenced from each other:
| Script | Purpose | Called when | Next step |
|---|---|---|---|
| `scripts/install.sh` | Install framework on a fresh machine | `curl \| bash` or `git clone + bash` | → `scripts/onboard-project.sh` |
| `scripts/onboard-project.sh` | Bootstrap automaton in a project | After framework install, per project | → `status.py --create-task` |
| `scripts/install-hooks.sh` | Install git hooks per project | After onboarding, or manually | see `onboard-project.sh` |
- `install.sh` outputs "Next: onboard-project.sh" at the end.
- `onboard-project.sh` outputs "Next: status.py --create-task" at the end.
- `install-hooks.sh` is called by `onboard-project.sh` automatically.
- `update.sh` does NOT call `onboard-project.sh` — it only updates the framework.
## Adding a New Script
1. Place the script in `scripts/`.
+21
View File
@@ -2,6 +2,27 @@
## [unreleased]
### Fixed — dashboard scroll-reset on auto-refresh
- **`automaton/dashboard/html/dashboard.js`** (`renderBoard`): auto-refresh rebuilt the board via `board.innerHTML = html` every tick (default 2s), destroying each `.column-body`'s `scrollTop` and snapping it back to 0 — so users couldn't scroll the Done group down to review older tasks. Now snapshots each column-body's `scrollTop` (plus the board's `scrollLeft` and the active view's `scrollTop`) before the rebuild and restores them after, matched by index (PHASE_GROUPS order is stable).
- **New tests**: `tests/test_dashboard_ui.py` — Playwright browser smoke test (board renders tasks; column scroll survives an auto-refresh tick). Skipped via `importorskip` when playwright/chromium is absent so CI without a browser stays green. Verified the test fails without the fix (scrollTop resets to 0) and passes with it.
### Fixed — failing plist-isolation test (host bleed false positive)
- **`tests/test_cleanup_done.py`** (`TestInstallCleanupScheduleIsolation.test_plist_written_to_override_dir_not_host`): asserted `not host.exists()`, but the host `~/Library/LaunchAgents/com.automaton.cleanup.plist` legitimately exists from a real `--install-cleanup-schedule` run, causing a false failure. Now snapshots the host plist's `st_mtime_ns` (or absence) before the test run and asserts it's unchanged after — a pre-existing real install no longer fails the test; only an actual write during the run would.
### Changed — bind ornith as the Implement model
- **`config.md`** (Model Configuration): set `Model: omlx/Ornith-1.0-35B-4bit-mlx` and `Override context window: 32768` (matches the opencode.json limit for the local LLM). Interactive autopilot already used ornith via opencode's default model; this makes it explicit so auto-detection can't pick another model. Loop ticks still use the single `harness.command` for all roles — per-role model binding (`{model}` substitution in loop-runner.py) is **not** implemented yet (see model-divergence gap below).
### Changed — README: document the self-improvement loop's scope for new projects
- **`README.md`** (Project Setup): added "The Self-Improvement Loop is framework-scoped" note — the default SI loop targets `~/.automaton/` (the framework), not your project, by design. Documents the leave-running / pause / create-a-project-loop paths.
### Known gap — model-divergence was marked complete but unimplemented
- The `model-divergence-enforcement` parent task and its 3 subtasks (`mde-manifest-detection`, `mde-interactive-enforcement`, `mde-loop-enforcement`) are `.state = complete` but contain only `SPEC.md`/`DECOMPOSITION.md` — no `IMPLEMENTATION.md`, no `VERDICT.md`. The promised code (`status.py` model_divergence audit category, `--transition --model`, `.state.models`, `loop.json` per-role `model` + `{model}` substitution in loop-runner.py, dashboard badges) was never written. Consequence: loop roles (implement/verify/orchestrate) all run the same model, so the D12 conflict-of-interest rule (Verify ≠ Implement session/model) is unenforced. Interactive autopilot is unaffected.
### Added — framework agent features design docs
- **New `design/framework/`** directory: design index, functional design, technical design, and backlog for three framework-level agent features:
+184 -20
View File
@@ -6,22 +6,28 @@ A contract-based operating system for LLM agents, designed to enforce discipline
Before you can use the framework in any project, you must install the core logic into your local environment.
**Run these commands in your terminal:**
Choose one of the following methods:
### Option A: One-liner (curl pipe, recommended)
```bash
# Clone the framework into the global config directory
git clone <your-git-url> ~/.automaton
# Enter the directory
cd ~/.automaton
# Make the installation script executable and run it
# You must provide the git URL as the first argument
chmod +x install.sh
./install.sh <your-git-url>
curl -fsSL https://raw.githubusercontent.com/<your-org>/automaton/main/scripts/install.sh | bash -s -- <your-git-url>
```
The git URL is required because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling.
This clones the framework to `~/.automaton/`, runs VRAM detection, installs pre-edit guards, creates the self-improvement loop, and sets up the Python virtualenv.
### Option B: Clone first
```bash
git clone <your-git-url> ~/.automaton
bash ~/.automaton/scripts/install.sh
```
The script detects that `~/.automaton` already exists, skips the clone, and runs all setup steps (VRAM detection, guards, loop, virtualenv).
### Both methods do the same thing
The git URL is required on fresh install because the framework uses it for self-updates and the self-improvement loop. Choose carefully -- it cannot be changed later without reinstalling.
*Note: This creates the "brain" of the framework (prompts, state machines, and rules) in your home directory. A self-improvement loop is created and scheduled by default (see Loop Engineering below).*
@@ -45,11 +51,95 @@ If a project was set up under the old model (with copies of framework files), it
---
## 1.5 Architecture: Framework vs Project
Automaton uses a **split architecture** — one shared framework, many project `./.automaton/` directories:
```
~/.automaton/ ← Framework (installed once per machine)
├── scripts/ ← shared tooling: status.py, loop-runner.py
├── prompts/ ← shared LLM prompts
├── plugins/ ← shared harness plugins
├── templates/ ← shared task & loop templates
├── .automaton/tasks/ ← framework housekeeping tasks (self-improvement)
└── .automaton/loops/ ← framework loops (self-improvement loop)
~/projects/my-app/
└── .automaton/ ← Project (onboarded once per project)
├── tasks/ ← YOUR project's tasks
├── models.json ← YOUR project's model config
├── config.md ← YOUR project's VRAM config
├── project-name.md ← YOUR project's display name
└── loops/ ← YOUR project's loops
```
**Key rules:**
- The agent is **scope-aware**: if you're inside `~/.automaton/`, it operates in **framework mode** (reads framework tasks). If you're inside a project dir, it operates in **project mode** (reads project tasks). They never interfere.
- All framework scripts (`status.py`, etc.) live in `~/.automaton/scripts/` and are shared — never copied into projects.
- Framework prompts live in `~/.automaton/prompts/` — projects reference them by path at runtime.
- Git hooks are **per-project**. Each project installs its own via `bash ~/.automaton/scripts/install-hooks.sh <project-path>`.
- The self-improvement loop targets **only** the framework itself. Your project won't get framework-level tasks in its board.
- You can work on both at the same time in different terminals — independent `.automaton/` directories, shared tooling.
### Lifecycle overview
```
┌─────────────────────────────────────────────────────┐
│ 1. Install Framework (once per machine) │
│ curl .../install.sh | bash -s -- <git-url> │
│ → clones to ~/.automaton/ │
│ → VRAM detection, guards, venv, self-improvement │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 2. Onboard Project (once per project) │
│ bash ~/.automaton/scripts/onboard-project.sh <dir> │
│ → creates .automaton/ skeleton │
│ → probes models, writes config.md │
│ → git init + hooks │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 3. Create Task (per feature) │
│ python3 ~/.automaton/scripts/status.py │
│ --create-task my-feature --project . │
└────────────────────────┬────────────────────────────┘
│
┌────────────────────────▼────────────────────────────┐
│ 4. Work Through Phases (per task) │
│ status.py --transition research --task my-feature │
│ → agent writes SPEC.md │
│ status.py --transition implement --task my-feature │
│ → agent writes code + IMPLEMENTATION.md │
│ ... → complete │
└─────────────────────────────────────────────────────┘
```
---
## 2. Project Setup (Per project)
Once the framework is installed globally, you must "onboard" every individual project you work on.
### Option A: The Agent-Driven Way (Recommended)
### Option A: The Onboarding Script (Recommended)
```bash
bash ~/.automaton/scripts/onboard-project.sh /path/to/project
```
This will:
1. Create `.automaton/` skeleton if missing.
2. Run `detect_models.py --write` to probe local models (falls back to a minimal `models.json`).
3. Generate `config.md` with VRAM recommendations.
4. Write `project-name.md` from the directory name.
5. Initialize git if not already a repo.
6. Install git hooks (pre-commit + pre-push).
7. Add automaton entries to `.gitignore`.
8. Run `status.py --audit` to verify the setup.
### Option B: The Agent-Driven Way
If you want the agent to handle the configuration for you, navigate to your project root and run:
> *"Onboard this project into automaton."
@@ -60,17 +150,91 @@ The agent will automatically:
3. Generate your `.agent.md` and `.rules.md` files.
4. Initiate the "Exploration Ritual" to understand your codebase.
### Option B: The Manual Way
If you prefer to set it up manually, create a `.automaton/` directory in your project root and add:
- `.agent.md`: Project-specific configuration (Mode, rules, etc.).
- `.rules.md`: Project-specific constraints and past failure modes.
### Option C: The Manual Way
If you prefer to set it up manually:
Then install the git pre-commit hook:
```bash
ln -sf ~/.automaton/scripts/git-hooks/pre-commit .git/hooks/pre-commit
# Create the automaton directory
mkdir -p .automaton/tasks .automaton/loops .automaton/design
# Install git hooks
bash ~/.automaton/scripts/install-hooks.sh .
# Configure your model(s)
python3 ~/.automaton/scripts/detect_models.py --write --project .
```
This hook blocks commits when no task is in `implement` or `doc_review` phase, preventing accidental code changes without a tracked task.
Then add `.agent.md` and `.rules.md` for the agent.
### The Self-Improvement Loop is framework-scoped
The self-improvement loop installed by default targets the **framework itself** (`work_source.project: ~/.automaton/`), not your project. It ticks hourly against `status.py --audit` on `~/.automaton/` to keep the framework healthy. This is by design — it improves the framework you depend on while you work on your project.
- **Leave it running** if you want the framework maintained in the background (recommended).
- **Disable it** if you want zero background activity: `python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/`
- **Want a loop on your project too?** Create a separate one targeted at the project root:
```bash
python3 ~/.automaton/scripts/status.py --create-loop my-project-loop \
--from-template self-improvement --project /path/to/project
```
---
## 2.5 FAQ
### Can I work on the framework and a project at the same time?
Yes. They have separate `.automaton/` directories. Open two terminals:
```
Terminal 1: cd ~/.automaton → framework mode
Terminal 2: cd ~/projects/my-app → project mode
```
The agent detects scope from your current directory. Each can have its own tasks, loops, and config. They share the same `~/.automaton/scripts/` binaries.
### Why doesn't `install.sh` need a Git URL when run from the repo?
Because the framework is already cloned. `install.sh` skips cloning when `~/.automaton/` exists and runs all the setup steps (VRAM detection, pip deps, self-improvement loop, guards). The Git URL is only required for a fresh install via `curl | bash`.
### Do I need to run `install.sh` again after pulling updates?
No. `git pull` inside `~/.automaton/` updates the code. The self-improvement loop and guards persist across updates. If you want to re-register guards (e.g. after switching harnesses), run `bash ~/.automaton/scripts/register-guards.sh`.
### How do git hooks work per project?
Each project installs its own hooks via:
```bash
bash ~/.automaton/scripts/install-hooks.sh /path/to/project
```
The pre-commit hook blocks commits when no task is in `implement` or `doc_review` phase. The pre-push hook catches `--no-verify` bypasses. They're independent per repo.
### Can I have multiple projects onboarded at once?
Yes. Each project gets its own `.automaton/` directory. Run `onboard-project.sh` once per project. The shared scripts in `~/.automaton/scripts/` enforce the state machine on whichever project you point `--project` at.
### What about loops on my project?
The self-improvement loop runs only on the framework. To add a loop to your project:
```bash
python3 ~/.automaton/scripts/status.py --create-loop my-loop \
--from-template self-improvement --project /path/to/project
python3 ~/.automaton/scripts/status.py --install-schedule my-loop \
--interval 3600 --project /path/to/project
```
### Using Pi Dev with automaton
Pi Dev has the `automaton-guard-pi` plugin (installed by `register-guards.sh`) which blocks edits outside allowed phases. However, Pi Dev does **not** auto-load automaton's system prompt (unlike opencode). For the agent to understand tasks and phases, provide context manually.
**Before starting a Pi Dev session**, run the context printer:
```bash
bash ~/.automaton/scripts/pi-automaton.sh
```
Or if symlinked to `~/bin/`:
```bash
pi-automaton
```
Copy the output and paste it as your first message to the Pi Dev agent. This tells the agent about:
- The automaton workflow framework and phase lifecycle
- Active tasks in the current project
- The default model and VRAM configuration
- Project-specific rules from `AGENTS.md`
- Which commands to use for transitions and task creation
The `automaton-guard-pi` plugin still blocks edits outside `implement`/`doc_review` even without this context — the context printer just makes the agent *aware* of why it's being blocked and how to use the framework correctly.
---
+42 -22
View File
@@ -115,6 +115,7 @@ class Task:
name: str
folder_path: Path
state: TaskState
phase_raw: str = ""
artifacts: dict[str, ArtifactStatus] = field(default_factory=dict)
sub_tasks: list[SubTask] = field(default_factory=list)
parent_spec: Optional[str] = None
@@ -125,10 +126,14 @@ class Task:
design_content: Optional[str] = None
spec_content: Optional[str] = None
decomposition_content: Optional[str] = None
code_review_content: Optional[str] = None
test_plan_content: Optional[str] = None
implementation_content: Optional[str] = None
parent_spec_content: Optional[str] = None
vram_config_content: Optional[str] = None
waves: list[WaveGroup] = field(default_factory=list)
is_corrupted: bool = False
models: dict[str, str] = field(default_factory=dict)
@property
def display_name(self) -> str:
@@ -297,7 +302,7 @@ class Task:
@property
def is_approval_gated(self) -> bool:
return self.state in (TaskState.RESEARCH, TaskState.DECOMPOSITION, TaskState.DESIGN, TaskState.TEST_DESIGN, TaskState.CODE_REVIEW)
return self.phase_raw.endswith(":awaiting_approval")
@property
def blocker(self) -> str:
@@ -310,7 +315,7 @@ class Task:
if not artifact.content:
return f"Empty required artifact: {required}"
if self.is_approval_gated:
return "Awaiting user approval — use `status.py --approve` to approve"
return "Awaiting user approval — click Approve Phase in the dashboard"
if self.state == TaskState.BUG_FIND:
if "BUG_REPORT.md" not in self.artifacts or not self.artifacts["BUG_REPORT.md"].content:
return "Agent must generate BUG_REPORT.md"
@@ -484,7 +489,7 @@ def _state_string_to_task_state(phase: str) -> TaskState:
return mapping.get(base, TaskState.BACKLOG)
def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, ArtifactStatus]]:
def determine_task_state(folder_path: Path) -> tuple[TaskState, str, dict[str, ArtifactStatus]]:
artifacts = {}
for filename in ARTIFACTS:
@@ -511,7 +516,7 @@ def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, Artifa
try:
phase = state_file.read_text(encoding="utf-8").strip()
if phase:
return _state_string_to_task_state(phase), artifacts
return _state_string_to_task_state(phase), phase, artifacts
except (OSError, IOError):
pass # Fall through to artifact heuristic
@@ -521,40 +526,40 @@ def determine_task_state(folder_path: Path) -> tuple[TaskState, dict[str, Artifa
if "VERDICT.md" in artifacts:
verdict_content = artifacts["VERDICT.md"].content
if not verdict_content:
return TaskState.BLOCKED, artifacts
return TaskState.BLOCKED, "blocked", artifacts
verdict_status = parse_verdict_status(verdict_content)
if verdict_status == VERDICT_FAIL or verdict_status == VERDICT_NEEDS_REVIEW:
return TaskState.BLOCKED, artifacts
return TaskState.BLOCKED, "blocked", artifacts
if verdict_status == VERDICT_PASS:
return TaskState.DONE, artifacts
return TaskState.DONE, "done", artifacts
# Verdict exists but status is unparseable — pending referee review
if verdict_status is None:
return TaskState.REFEREE, artifacts
return TaskState.REFEREE, "referee", artifacts
# State machine aligned with orchestrate.md
# Check from most advanced to least advanced
if "DOC_REVIEW.md" in artifacts:
return TaskState.DOC_REVIEW, artifacts
return TaskState.DOC_REVIEW, "doc_review", artifacts
if "ADVERSARIAL_BUG_REPORT.md" in artifacts and "BUG_REPORT.md" in artifacts:
return TaskState.ADV_BUG_FIND, artifacts
return TaskState.ADV_BUG_FIND, "adv_bug_find", artifacts
if "BUG_REPORT.md" in artifacts:
return TaskState.BUG_FIND, artifacts
return TaskState.BUG_FIND, "bug_find", artifacts
if "ADVERSARIAL_BUG_REPORT.md" in artifacts:
return TaskState.BUG_FIND, artifacts
return TaskState.BUG_FIND, "bug_find", artifacts
if "CODE_REVIEW.md" in artifacts:
return TaskState.CODE_REVIEW, artifacts
return TaskState.CODE_REVIEW, "code_review", artifacts
if "IMPLEMENTATION.md" in artifacts:
return TaskState.IMPLEMENT, artifacts
return TaskState.IMPLEMENT, "implement", artifacts
if "TEST_PLAN.md" in artifacts:
return TaskState.TEST_DESIGN, artifacts
return TaskState.TEST_DESIGN, "test_design", artifacts
if "DESIGN.md" in artifacts:
return TaskState.DESIGN, artifacts
return TaskState.DESIGN, "design", artifacts
if "DECOMPOSITION.md" in artifacts:
return TaskState.DECOMPOSITION, artifacts
return TaskState.DECOMPOSITION, "decomposition", artifacts
if "SPEC.md" in artifacts:
return TaskState.RESEARCH, artifacts
return TaskState.RESEARCH, "research", artifacts
return TaskState.BACKLOG, artifacts
return TaskState.BACKLOG, "backlog", artifacts
def parse_sub_tasks(parent_folder: Path) -> list[SubTask]:
@@ -569,7 +574,7 @@ def parse_sub_tasks(parent_folder: Path) -> list[SubTask]:
name = subtask_folder.name
if not _VALID_TASK_NAME_CHARS.issuperset(set(name)):
continue
state, artifacts = determine_task_state(subtask_folder)
state, _, artifacts = determine_task_state(subtask_folder)
verdict_status = None
has_verdict = False
if "VERDICT.md" in artifacts and artifacts["VERDICT.md"].content:
@@ -653,12 +658,12 @@ def discover_tasks(tasks_dir: Path) -> list[Task]:
if not _VALID_TASK_NAME_CHARS.issuperset(set(folder_path.name)):
continue
state, artifacts = determine_task_state(folder_path)
state, phase_raw, artifacts = determine_task_state(folder_path)
sub_tasks = parse_sub_tasks(folder_path)
task = Task(
name=folder_path.name, folder_path=folder_path, state=state,
artifacts=artifacts, sub_tasks=sub_tasks,
phase_raw=phase_raw, artifacts=artifacts, sub_tasks=sub_tasks,
)
# Load specific artifact contents for detail panel
@@ -677,9 +682,24 @@ def discover_tasks(tasks_dir: Path) -> list[Task]:
if "DECOMPOSITION.md" in artifacts and artifacts["DECOMPOSITION.md"].content:
task.decomposition_content = artifacts["DECOMPOSITION.md"].content
task.waves = parse_waves(task.decomposition_content)
if "CODE_REVIEW.md" in artifacts and artifacts["CODE_REVIEW.md"].content:
task.code_review_content = artifacts["CODE_REVIEW.md"].content
if "TEST_PLAN.md" in artifacts and artifacts["TEST_PLAN.md"].content:
task.test_plan_content = artifacts["TEST_PLAN.md"].content
if "IMPLEMENTATION.md" in artifacts and artifacts["IMPLEMENTATION.md"].content:
task.implementation_content = artifacts["IMPLEMENTATION.md"].content
task.parent_spec_content = parse_parent_spec(folder_path)
task.vram_config_content = parse_vram_config(folder_path)
# Load .state.models for model divergence badges
state_models_path = folder_path / ".state.models"
if state_models_path.exists():
try:
import json as _json
task.models = _json.loads(state_models_path.read_text(encoding="utf-8"))
except (OSError, IOError, _json.JSONDecodeError):
task.models = {}
tasks.append(task)
# Sort by state (most advanced first)
+143 -4
View File
@@ -110,6 +110,13 @@ async function fetchProjectName() {
function renderBoard() {
const board = document.getElementById('board');
// Preserve scroll positions across re-renders so auto-refresh doesn't
// snap columns back to the top while the user is reviewing older tasks.
const prevBodies = Array.from(board.querySelectorAll('.column-body'));
const savedScrolls = prevBodies.map(el => el.scrollTop);
const savedBoardScrollLeft = board.scrollLeft;
const view = document.querySelector('.view.active');
const savedViewScrollTop = view ? view.scrollTop : 0;
const filtered = getFilteredTasks();
const groups = {};
PHASE_GROUPS.forEach(group => { groups[group.id] = []; });
@@ -140,9 +147,25 @@ function renderBoard() {
</div>`;
}).join('');
board.innerHTML = html;
// Restore scroll positions (matched by index — PHASE_GROUPS order is stable).
const newBodies = board.querySelectorAll('.column-body');
newBodies.forEach((el, i) => { if (savedScrolls[i] != null) el.scrollTop = savedScrolls[i]; });
board.scrollLeft = savedBoardScrollLeft;
if (view) view.scrollTop = savedViewScrollTop;
attachCardListeners();
}
const ROLE_LABELS = {
'implement': 'Implement',
'code_review': 'Code Review',
'bug_find': 'Bug Find',
'adversarial_bug_find': 'Adv Bug Find',
'doc_review': 'Doc Review',
'referee': 'Referee',
'loop-implement': 'Loop Impl',
'loop-verify': 'Loop Verify',
};
const ARTIFACT_LABELS = {
'research': 'SPEC.md', 'decomposition': 'DECOMPOSITION.md',
'design': 'DESIGN.md', 'test_design': 'TEST_PLAN.md',
@@ -170,11 +193,20 @@ function renderTaskCard(task) {
const label = ARTIFACT_LABELS[col.id] || col.label;
return `<span class="artifact-badge" title="${col.label}">${label}</span>`;
}).join('')}</div>`;
const modelKeys = Object.keys(task.models || {});
const modelsHtml = modelKeys.length > 0
? `<div class="task-card-models">${modelKeys.map(role => {
const m = task.models[role];
const roleLabel = ROLE_LABELS[role] || role;
return `<span class="model-badge" title="${roleLabel}: ${m}">${m}</span>`;
}).join('')}</div>`
: '';
return `<div class="task-card" data-task="${task.name}" data-status="${statusClass}">
<div class="task-card-header"><span class="task-card-name">${escapeHtml(task.display_name)}</span><span class="task-card-status ${statusClass}">${statusIcon}</span></div>
<div class="task-card-sublabel">${subLabel}</div>
${task.status_reason ? `<div class="task-card-reason">${escapeHtml(task.status_reason)}</div>` : ''}
${artifactsHtml}
${modelsHtml}
${progressHtml ? `<div class="task-card-footer"><span class="subtask-progress">${progressHtml}</span></div>` : ''}
${subtasksHtml}
</div>`;
@@ -198,13 +230,69 @@ function renderDetail(task) {
return `<span class="detail-artifact"><span class="${cls}">${icon}</span>${col.label}</span>`;
}).join('');
const phaseGroupHtml = phaseGroup ? `<span class="detail-phase-badge" style="background: ${phaseGroupColor}20; color: ${phaseGroupColor}">${PHASE_GROUPS.find(g => g.id === phaseGroup).label}</span>` : '';
const approvePhaseBtn = task.is_approval_gated
? `<button class="approve-phase-btn" data-task="${escapeHtml(task.name)}">🔓 Approve Phase</button>`
// Build approval section for approval-gated phases
const APPROVAL_ARTIFACT_MAP = {
'research': { contentKey: 'spec_content', label: 'SPEC.md', phaseName: 'Research' },
'decomposition': { contentKey: 'decomposition_content', label: 'DECOMPOSITION.md', phaseName: 'Decomposition' },
'design': { contentKey: 'design_content', label: 'DESIGN.md', phaseName: 'Design' },
'test_design': { contentKey: 'test_plan_content', label: 'TEST_PLAN.md', phaseName: 'Test Design' },
'code_review': { contentKey: 'code_review_content', label: 'CODE_REVIEW.md', phaseName: 'Code Review' },
};
const awaitingPhase = task.phase_raw ? task.phase_raw.replace(':awaiting_approval', '') : null;
const approvalInfo = awaitingPhase ? APPROVAL_ARTIFACT_MAP[awaitingPhase] : null;
const approvalContent = approvalInfo && approvalInfo.contentKey ? task[approvalInfo.contentKey] : null;
const approvalHtml = task.is_approval_gated
? `<div class="approval-section">
<div class="approval-header">
<span class="approval-icon">🔒</span>
<span class="approval-title">${approvalInfo ? approvalInfo.phaseName : 'Phase'} — Needs Approval</span>
<button class="approve-phase-btn" data-task="${escapeHtml(task.name)}">Approve Phase</button>
</div>
${task.blocker ? `<p class="approval-blocker">${escapeHtml(task.blocker)}</p>` : ''}
${approvalContent ? `<details class="approval-artifact" open>
<summary>${approvalInfo.label} — review content before approving</summary>
<pre class="detail-content-text">${escapeHtml(approvalContent)}</pre>
</details>` : `<p class="approval-missing">${approvalInfo ? approvalInfo.label : 'Artifact'} not yet written — an agent must create it before this phase can complete.</p>`}
</div>`
: '';
// Build transition section for phases that can advance
const TRANSITION_MAP = {
'research:approved': { target: 'decomposition', label: 'Advance to Decomposition' },
'decomposition:approved': { target: 'design', label: 'Advance to Design' },
'design:approved': { target: 'implement', label: 'Advance to Implementation' },
'test_design:approved': { target: 'implement', label: 'Advance to Implementation' },
'implement': { target: 'code_review', label: 'Advance to Code Review' },
'code_review:approved': { target: 'bug_find', label: 'Advance to Bug Finding' },
'bug_find': { target: 'adv_bug_find', label: 'Advance to Adversarial Bug Finding' },
'adv_bug_find': { target: 'doc_review', label: 'Advance to Document Review' },
'doc_review': { target: 'referee', label: 'Advance to Referee' },
};
const nextTransition = TRANSITION_MAP[task.phase_raw];
const transitionHtml = nextTransition
? `<div class="detail-section detail-actions"><h4>🚀 Phase Actions</h4><button class="transition-btn" data-task="${escapeHtml(task.name)}" data-target="${nextTransition.target}">${nextTransition.label}</button></div>`
: '';
// Build artifact editor for missing required artifacts
const requiredArtifact = task.required_artifact_name;
const hasRequiredArtifact = task.artifacts[requiredArtifact];
const artifactEditorHtml = requiredArtifact && !hasRequiredArtifact && !task.is_approval_gated
? `<div class="detail-section detail-artifact-editor">
<h4>📝 Write ${requiredArtifact}</h4>
<p class="artifact-editor-hint">This artifact is required before the task can advance. Write it below and save.</p>
<textarea class="artifact-editor" data-task="${escapeHtml(task.name)}" data-filename="${requiredArtifact}" rows="12" placeholder="# ${requiredArtifact.replace('.md','')}\n\nWrite content here..."></textarea>
<button class="save-artifact-btn" data-task="${escapeHtml(task.name)}" data-filename="${requiredArtifact}">Save ${requiredArtifact}</button>
</div>`
: '';
content.innerHTML = `
<div class="detail-section"><h4>Status</h4><span class="detail-status-badge ${statusClass}">${statusText}</span>${phaseGroupHtml}${task.is_edit_phase ? '<span class="detail-edit-badge">✏️ Edit Allowed</span>' : ''}${task.is_approval_gated ? '<span class="detail-approval-badge">🔒 Requires Approval</span>' : ''}${approvePhaseBtn}</div>
<div class="detail-section"><h4>Status</h4><span class="detail-status-badge ${statusClass}">${statusText}</span>${phaseGroupHtml}${task.is_edit_phase ? '<span class="detail-edit-badge">✏️ Edit Allowed</span>' : ''}</div>
${statusReason ? `<div class="detail-section detail-status-reason">${escapeHtml(statusReason)}</div>` : ''}
${task.blocker ? `<div class="detail-section detail-blocker"><h4>⚠️ What's Blocking</h4><p class="detail-blocker-text">${escapeHtml(task.blocker)}</p></div>` : ''}
${approvalHtml}
${task.blocker && !task.is_approval_gated ? `<div class="detail-section detail-blocker"><h4>⚠️ What's Blocking</h4><p class="detail-blocker-text">${escapeHtml(task.blocker)}</p></div>` : ''}
${artifactEditorHtml}
${transitionHtml}
${task.phase_guidance ? `<div class="detail-section detail-guidance"><h4>▶ What's Next</h4><pre class="detail-guidance-text">${escapeHtml(task.phase_guidance)}</pre></div>` : ''}
${task.state === 'blocked' && task.blocked_action_items && task.blocked_action_items.length > 0 ? `<div class="detail-section detail-action-items"><h4>📋 Action Items</h4><ul class="detail-action-list">${task.blocked_action_items.map(item => `<li>${escapeHtml(item)}</li>`).join('')}</ul></div>` : ''}
<div class="detail-section"><h4>Artifacts</h4><div class="detail-artifacts">${artifactsHtml}</div></div>
@@ -504,6 +592,44 @@ async function approvePhase(taskName) {
}
}
async function transitionTask(taskName, target) {
try {
const res = await fetch(`/api/transition/${taskName}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ target }),
});
const data = await res.json();
if (data.success) {
closeDetail();
await refreshData();
} else {
console.error('Transition failed:', data);
}
} catch (err) {
console.error('Transition failed:', err);
}
}
async function saveArtifact(taskName, filename, content) {
try {
const res = await fetch(`/api/write-artifact/${taskName}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ filename, content }),
});
const data = await res.json();
if (data.success) {
closeDetail();
await refreshData();
} else {
console.error('Save artifact failed:', data);
}
} catch (err) {
console.error('Save artifact failed:', err);
}
}
function getFilteredTasks() {
let filtered = [...state.tasks];
if (state.filterPhase !== 'all') filtered = filtered.filter(t => t.state === state.filterPhase);
@@ -593,6 +719,19 @@ function setupUI() {
const taskName = approveBtn.dataset.task;
if (taskName) approvePhase(taskName);
}
const transitionBtn = e.target.closest('.transition-btn');
if (transitionBtn) {
const taskName = transitionBtn.dataset.task;
const target = transitionBtn.dataset.target;
if (taskName && target) transitionTask(taskName, target);
}
const saveBtn = e.target.closest('.save-artifact-btn');
if (saveBtn) {
const taskName = saveBtn.dataset.task;
const filename = saveBtn.dataset.filename;
const editor = document.querySelector(`.artifact-editor[data-task="${taskName}"][data-filename="${filename}"]`);
if (taskName && filename && editor) saveArtifact(taskName, filename, editor.value);
}
});
}
+40
View File
@@ -307,8 +307,48 @@ kbd {
::-webkit-scrollbar-thumb:hover { background: var(--border-active); }
.approve-phase-btn { display: inline-block; margin-left: 8px; padding: 4px 12px; border: 1px solid var(--warning); border-radius: var(--radius-sm); cursor: pointer; font-size: 11px; font-weight: 500; background: var(--warning-bg); color: var(--warning); font-family: inherit; transition: all 0.15s; }
.approve-phase-btn:hover { filter: brightness(1.1); }
/* Approval section — prominent card for approval-gated tasks */
.approval-section {
margin: 12px 0;
padding: 14px 16px;
border: 2px solid var(--warning);
border-radius: var(--radius-md);
background: var(--warning-bg);
}
.approval-header {
display: flex;
align-items: center;
gap: 8px;
margin-bottom: 8px;
}
.approval-icon { font-size: 18px; }
.approval-title { font-size: 14px; font-weight: 600; flex: 1; }
.approval-blocker { font-size: 12px; color: var(--text-secondary); margin: 0 0 8px 0; padding: 6px 10px; background: var(--bg-card); border-radius: var(--radius-sm); }
.approval-missing { font-size: 12px; color: var(--text-secondary); margin: 8px 0 0 0; font-style: italic; }
.approval-artifact { margin-top: 8px; font-size: 12px; }
.approval-artifact summary { cursor: pointer; font-weight: 500; padding: 4px 0; color: var(--text-primary); }
.approval-artifact summary:hover { color: var(--primary); }
.task-card-artifacts { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 6px; }
.artifact-badge { font-size: 10px; padding: 2px 8px; background: var(--bg-primary); border: 1px solid var(--border-color); border-radius: 4px; color: var(--text-secondary); font-family: 'SF Mono', 'Fira Code', monospace; font-weight: 500; }
.task-card-models { display: flex; flex-wrap: wrap; gap: 4px; margin-top: 6px; }
.model-badge { font-size: 10px; padding: 1px 6px; background: #e3f2fd; border: 1px solid #90caf9; border-radius: 4px; color: #1565c0; font-family: 'SF Mono', 'Fira Code', monospace; font-weight: 500; }
/* Transition button — advance to next phase */
.transition-btn { display: inline-block; padding: 6px 16px; border: 1px solid var(--primary); border-radius: var(--radius-sm); cursor: pointer; font-size: 12px; font-weight: 500; background: rgba(74, 144, 226, 0.1); color: var(--primary); font-family: inherit; transition: all 0.15s; }
.transition-btn:hover { filter: brightness(1.15); background: rgba(74, 144, 226, 0.2); }
.detail-actions { margin: 8px 0; padding: 10px 12px; background: var(--bg-card); border-radius: var(--radius-md); border: 1px solid var(--border-color); }
.detail-actions h4 { margin: 0 0 8px 0; font-size: 12px; color: var(--text-secondary); font-weight: 500; }
/* Artifact editor — inline textarea for writing missing artifacts */
.detail-artifact-editor { margin: 8px 0; padding: 12px; background: var(--bg-card); border: 1px solid var(--primary); border-radius: var(--radius-md); }
.detail-artifact-editor h4 { margin: 0 0 4px 0; font-size: 12px; font-weight: 600; }
.artifact-editor-hint { font-size: 11px; color: var(--text-secondary); margin: 0 0 8px 0; }
.artifact-editor { width: 100%; padding: 10px; border: 1px solid var(--border-color); border-radius: var(--radius-sm); background: var(--bg-primary); color: var(--text-primary); font-family: 'SF Mono', 'Fira Code', monospace; font-size: 12px; line-height: 1.5; resize: vertical; box-sizing: border-box; }
.artifact-editor:focus { outline: none; border-color: var(--primary); }
.save-artifact-btn { display: inline-block; margin-top: 8px; padding: 6px 16px; border: 1px solid var(--success); border-radius: var(--radius-sm); cursor: pointer; font-size: 12px; font-weight: 500; background: rgba(74, 208, 120, 0.1); color: var(--success); font-family: inherit; transition: all 0.15s; }
.save-artifact-btn:hover { filter: brightness(1.15); background: rgba(74, 208, 120, 0.2); }
@media (max-width: 768px) {
.header { flex-wrap: wrap; gap: 8px; }
+119 -2
View File
@@ -51,6 +51,21 @@ TASK_STATE_ARTIFACT = {
TASK_STATES = list(TASK_STATE_ARTIFACT.keys())
def next_transition(phase_raw: str) -> str | None:
transitions = {
"research:approved": "decomposition",
"decomposition:approved": "design",
"design:approved": "implement",
"test_design:approved": "implement",
"implement": "code_review",
"code_review:approved": "bug_find",
"bug_find": "adv_bug_find",
"adv_bug_find": "doc_review",
"doc_review": "referee",
"referee": "complete",
}
return transitions.get(phase_raw)
MAX_POST_BODY = 65536 # 64KB
MAX_REVIEW_COMMENT_LENGTH = 4096
CACHE_TTL = 1.0 # seconds
@@ -120,6 +135,12 @@ class DashboardHandler(SimpleHTTPRequestHandler):
elif self.path.startswith("/api/approve/"):
task_name = unquote(self.path.split("/api/approve/")[1])
self._handle_phase_approval(task_name)
elif self.path.startswith("/api/transition/"):
task_name = unquote(self.path.split("/api/transition/")[1])
self._handle_transition(task_name)
elif self.path.startswith("/api/write-artifact/"):
task_name = unquote(self.path.split("/api/write-artifact/")[1])
self._handle_write_artifact(task_name)
else:
self._send_error(404, "Not found")
@@ -196,6 +217,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"name": t.name,
"display_name": t.display_name,
"state": t.state.value,
"phase_raw": t.phase_raw,
"status_reason": t.status_reason,
"artifacts": {s: TASK_STATE_ARTIFACT[s] in t.artifacts for s in TASK_STATES},
"sub_tasks": [
@@ -204,8 +226,14 @@ class DashboardHandler(SimpleHTTPRequestHandler):
],
"verdict_content": t.verdict_content,
"bug_report_content": t.bug_report_content,
"adversarial_bug_report_content": t.adversarial_bug_report_content,
"doc_review_content": t.doc_review_content,
"design_content": t.design_content,
"spec_content": t.spec_content,
"decomposition_content": t.decomposition_content,
"code_review_content": t.code_review_content,
"test_plan_content": t.test_plan_content,
"implementation_content": t.implementation_content,
"parent_spec_content": t.parent_spec_content,
"vram_config_content": t.vram_config_content,
"blocked_action_items": t.blocked_action_items,
@@ -217,7 +245,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"is_approval_gated": t.is_approval_gated,
"blocker": t.blocker,
"waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in t.waves],
"review": self._get_review_status(t.name),
"models": t.models,
}
for t in tasks
]
@@ -302,6 +330,7 @@ class DashboardHandler(SimpleHTTPRequestHandler):
"name": task.name,
"display_name": task.display_name,
"state": task.state.value,
"phase_raw": task.phase_raw,
"status_reason": task.status_reason,
"artifacts": {s: TASK_STATE_ARTIFACT[s] in task.artifacts for s in TASK_STATES},
"sub_tasks": [
@@ -310,12 +339,22 @@ class DashboardHandler(SimpleHTTPRequestHandler):
],
"verdict_content": task.verdict_content,
"bug_report_content": task.bug_report_content,
"adversarial_bug_report_content": task.adversarial_bug_report_content,
"doc_review_content": task.doc_review_content,
"design_content": task.design_content,
"spec_content": task.spec_content,
"decomposition_content": task.decomposition_content,
"parent_spec_content": task.parent_spec_content,
"vram_config_content": task.vram_config_content,
"blocked_action_items": task.blocked_action_items,
"unblock_instructions": task.unblock_instructions,
"phase_guidance": task.phase_guidance,
"required_artifact_name": task.required_artifact_name,
"next_phase_name": task.next_phase_name,
"is_edit_phase": task.is_edit_phase,
"is_approval_gated": task.is_approval_gated,
"blocker": task.blocker,
"waves": [{"wave_number": w.wave_number, "label": w.label, "sub_task_names": w.sub_task_names} for w in task.waves],
"review": self._get_review_status(task.name),
}
self._send_json(task_data)
@@ -416,6 +455,84 @@ class DashboardHandler(SimpleHTTPRequestHandler):
else:
self._send_error(400, result.stderr.strip() or result.stdout.strip())
def _handle_transition(self, task_name: str):
if not self._validate_task_name(task_name):
self._send_error(400, "Invalid task name")
return
project_root = self.project_root
if not project_root:
self._send_error(503, "Not in automaton project")
return
status_py = Path.home() / ".automaton" / "scripts" / "status.py"
if not status_py.exists():
self._send_error(500, "status.py not found")
return
import subprocess
content_length = int(self.headers.get('Content-Length', 0))
target = None
if content_length > 0:
try:
body = json.loads(self.rfile.read(content_length))
target = body.get("target")
except (json.JSONDecodeError, UnicodeDecodeError):
pass
if not target:
tasks = _get_cached_tasks(project_root)
task = next((t for t in tasks if t.name == task_name), None)
if not task:
self._send_error(404, "Task not found")
return
target = next_transition(task.phase_raw)
if not target:
self._send_json({"success": False, "message": "Cannot determine next transition from current phase"})
return
result = subprocess.run(
[sys.executable, str(status_py), "--transition", target, "--task", task_name,
"--project", str(project_root)],
capture_output=True, text=True, timeout=30,
)
_invalidate_task_cache()
if result.returncode == 0:
self._send_json({"success": True, "message": f"Transitioned to {target}: {result.stdout.strip()}"})
else:
self._send_json({"success": False, "message": result.stderr.strip() or result.stdout.strip()})
def _handle_write_artifact(self, task_name: str):
if not self._validate_task_name(task_name):
self._send_error(400, "Invalid task name")
return
project_root = self.project_root
if not project_root:
self._send_error(503, "Not in automaton project")
return
content_length = int(self.headers.get('Content-Length', 0))
if content_length > MAX_POST_BODY:
self._send_error(413, "Payload too large")
return
if content_length == 0:
self._send_error(400, "Empty request body")
return
try:
body = json.loads(self.rfile.read(content_length))
filename = body.get("filename", "")
content = body.get("content", "")
if not filename or not filename.endswith(".md"):
self._send_error(400, "Invalid filename — must be a .md file")
return
tasks = _get_cached_tasks(project_root)
task = next((t for t in tasks if t.name == task_name), None)
if not task:
self._send_error(404, "Task not found")
return
artifact_path = Path(task.folder_path) / filename
artifact_path.write_text(content, encoding="utf-8")
_invalidate_task_cache()
self._send_json({"success": True, "message": f"Written {filename}"})
except (json.JSONDecodeError, UnicodeDecodeError):
self._send_error(400, "Invalid JSON")
except (OSError, IOError) as e:
self._send_error(500, f"Failed to write file: {e}")
def _serve_review_summary(self):
project_root = self.project_root
if not project_root:
+13 -2
View File
@@ -33,8 +33,8 @@ To disable auto-detection and use manual values:
Settings for the LLM model being used.
- **Model**: auto # Use auto-detection from API config, or specify explicitly (e.g., gpt-4o, claude-3-5-sonnet)
- **Override context window**: auto # Override auto-detection, or specify (e.g., 128k, 200k)
- **Model**: omlx/Ornith-1.0-35B-4bit-mlx # Local LLM (opencode provider); used as the Implement role
- **Override context window**: 32768 # Matches opencode.json limit.context for ornith
### Auto-detection
@@ -51,6 +51,17 @@ To disable auto-detection and use manual values:
- **Override context window**: 128k
```
## Available Models
Models available for model-divergence enforcement. This file is managed by `scripts/detect_models.py`. In single-LLM mode (0-1 models), no hard blocks are enforced. In multi-LLM mode (2+ models), the conflict matrix enforces role-model separation.
- **Default**: omlx/Ornith-1.0-35B-4bit-mlx # Used when no role-specific binding is set
- **Advised**: true # Recommend a second model in single-LLM mode
No additional models are configured in the manifest. To add models:
1. Run `python3 ~/.automaton/scripts/detect_models.py --write` to auto-detect from opencode.json and localhost endpoints.
2. Or manually create `~/.automaton/models.json` (see `design/framework/technical.md` §2 for schema).
## System Requirements
Requirements for the environment the framework runs in.
+5
View File
@@ -103,6 +103,7 @@ Gate checks, in order:
4. **Task phase** — if `current_task` is set, that task's `.state` must still be one of the phases this loop is allowed to operate on. If the task has transitioned out (e.g. to `human_intervention` by some other path), halt as `human_intervention`.
5. **Worktree drift** — if worktree branch diverges from main in a way that indicates the loop wrote files outside its scope (checked via `git diff --name-only main...HEAD` restricted to `file_scope`), halt as `drift_detected`.
6. **Score plateau** — last N entries in `score_history` are flat or monotonically decreasing (where N = `score_plateau_window`). Trip → halt as `verifier_failed`.
7. **Model divergence** — in multi-LLM mode (2+ models in `models.json`), checks that the loop's implement and verify roles use different models. If they share the same model, halt as `human_intervention` (this prevents same-model verification / rubber-stamping within a loop tick). Single-LLM mode is exempt. Model is resolved from `roles[<role>].model` if set, otherwise the manifest default.
All halts atomically set `status=halted`, `halt_reason=<reason>`, write to `.state.log`, and call `--pause-loop`'s schedule-disable step (see §6).
@@ -258,11 +259,13 @@ To bound `outputs/` directory growth (O5 from `add-loop-runner/BUG_REPORT.md`),
v1.1's default `harness.command` is `opencode run` -- matching the framework's primary harness -- but the shape is generic. The runner substitutes the following tokens into the `command` list (single argv element per token, no shell expansion):
- `{model}` -- the model assigned to the role (from `roles[<role>].model` or manifest default). Passed via `extras["model"]`. If the role has no model assignment, the token is left unsubstituted.
- `{prompt}` -- resolved prompt file path (loop-local override or framework default). Kept for backwards compat and harnesses that prefer a file path.
- `{prompt_content}` -- the resolved prompt file's text content as a single argv element. Safe under `subprocess.run` list mode; no shell quoting needed. Used by the default command since `opencode run` takes the message as a positional argument and has no `--prompt-file` flag.
- `{cwd}` -- the working directory the harness should run in (the loop's project root or worktree).
- `{output}`, `{artifact}` -- role-specific extras (the implement output path handed to verify).
- `{verdict}`, `{current_task}`, `{current_phase}`, etc. -- other runtime extras; see `_resolve_prompt` below.
- `{model}` -- the model assigned to the role being invoked (from `roles[<role>].model` in `loop.json`, or the manifest default). The runner passes it via the `extras["model"]` key. If the role has no explicit model, `{model}` is left as-is (no substitution). This allows per-role model pinning without hardcoding the model name in `harness.command`.
The default command does NOT hardcode a `--model` flag; the spawned `opencode run` inherits the model from the project/user config. Users who want a per-loop model override (e.g. a local LLM for ticks) set `harness.command` in their `loop.json`:
@@ -336,6 +339,8 @@ This means the harness receives a fully-resolved prompt file with all context ba
}
```
Each role in `roles` accepts an optional `"model"` field to pin a specific model for that role (e.g. `"implement": {"prompt": "loop-implement.md", "tier": 16000, "model": "model-a"}`). When set, the runner passes `model=<value>` in the harness extras for that role, enabling `{model}` substitution in `harness.command`. This is how multi-LLM loops prevent same-model verification — see `CONFLICT_MATRIX` in `status.py`.
Installs default-on at `install.sh` time: `status.py --create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` then `status.py --install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"`. Both commands use `|| true` so the framework works even if loop creation fails. `update.sh` bootstraps the loop idempotently for existing users (checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`). Disabling: `status.py --pause-loop self-improvement --project ~/.automaton/`.
## 10. Tests (`tests/test_loops.py`)
+141
View File
@@ -0,0 +1,141 @@
import type { ExtensionAPI, ExtensionContext, BeforeAgentStartEventResult } from "@earendil-works/pi-coding-agent";
import { existsSync, readFileSync, readdirSync, statSync } from "fs";
import { basename, join } from "path";
import { homedir } from "os";
const AUTOMATON_HOME = join(homedir(), ".automaton");
const STALE_MINUTES = 30;
interface TaskInfo {
name: string;
phase: string;
mtime: Date;
}
function getTasks(autoDir: string): TaskInfo[] {
const tasksDir = join(autoDir, ".automaton", "tasks");
if (!existsSync(tasksDir)) return [];
try {
const entries = readdirSync(tasksDir);
const tasks: TaskInfo[] = [];
for (const entry of entries) {
const stateFile = join(tasksDir, entry, ".state");
if (existsSync(stateFile)) {
try {
const phase = readFileSync(stateFile, "utf-8").trim();
if (!phase) continue;
const stat = statSync(stateFile);
tasks.push({ name: entry, phase, mtime: stat.mtime });
} catch {
// skip unreadable
}
}
}
tasks.sort((a, b) => b.mtime.getTime() - a.mtime.getTime());
return tasks;
} catch {
return [];
}
}
function getLoopInfo(autoDir: string): string[] {
const loopsDir = join(autoDir, ".automaton", "loops");
if (!existsSync(loopsDir)) return [];
try {
const entries = readdirSync(loopsDir);
const lines: string[] = [];
for (const entry of entries) {
const stateFile = join(loopsDir, entry, ".state.loop");
if (existsSync(stateFile)) {
try {
const content = readFileSync(stateFile, "utf-8").trim();
const state = JSON.parse(content);
const taskRef = state.current_task ? `, active task: ${state.current_task}` : "";
lines.push(`Loop "${entry}": ${state.status || "unknown"}${taskRef}`);
} catch {
// skip unparseable
}
}
}
return lines;
} catch {
return [];
}
}
export default function (pi: ExtensionAPI) {
pi.on("before_agent_start", async (_event, ctx): Promise<BeforeAgentStartEventResult | undefined> => {
const cwd = ctx.cwd;
const inFramework = cwd === AUTOMATON_HOME || cwd.startsWith(AUTOMATON_HOME + "/");
let baseDir: string | null = null;
let scopeLabel: string;
if (inFramework) {
baseDir = AUTOMATON_HOME;
scopeLabel = "framework (~/.automaton/)";
} else {
const projectAuto = join(cwd, ".automaton");
if (existsSync(projectAuto)) {
baseDir = cwd;
const projectName = basename(cwd) || "project";
scopeLabel = `project (${projectName}/.automaton/)`;
} else {
return; // not in automaton context
}
}
const tasks = getTasks(baseDir);
const editTasks = tasks.filter((t) => t.phase === "implement" || t.phase === "doc_review");
const currentTask = editTasks.length > 0 ? editTasks[0] : tasks.length > 0 ? tasks[0] : null;
const lines: string[] = [];
lines.push(`Automaton scope: ${scopeLabel}`);
lines.push("");
lines.push("How Automaton works:");
lines.push("- Automaton enforces a task phase state machine. File edits are ONLY allowed when a task is in 'implement' or 'doc_review' phase.");
lines.push("- When the user asks you to build, fix, or change something, you MUST first create a task (use automaton_create_task) and transition it to 'implement' before you can edit any files.");
lines.push("- Without a task in 'implement', the automaton-guard-pi extension will BLOCK all file edits.");
lines.push(`- Available tools: automaton_create_task (create+transition), automaton_transition (change phase), automaton_status (check state).`);
lines.push("");
if (currentTask) {
lines.push(`State: Active task "${currentTask.name}" is in "${currentTask.phase}" phase.`);
if (currentTask.phase === "implement" || currentTask.phase === "doc_review") {
lines.push(`You CAN edit files under this task.`);
const ageMinutes = Math.round((Date.now() - currentTask.mtime.getTime()) / 60000);
if (ageMinutes > STALE_MINUTES) {
lines.push(
`Note: task has been in ${currentTask.phase} for ${ageMinutes} min and is considered stale. ` +
`Run automaton_transition (or touch via status.py) if still active.`,
);
}
} else {
lines.push(`File edits are BLOCKED. Call automaton_transition to move it to 'implement' before editing.`);
}
} else {
lines.push(`State: No active tasks. When the user makes a work request, call automaton_create_task first.`);
}
const byPhase = new Map<string, number>();
for (const t of tasks) {
byPhase.set(t.phase, (byPhase.get(t.phase) || 0) + 1);
}
const summary = Array.from(byPhase.entries())
.sort((a, b) => b[1] - a[1])
.map(([p, c]) => `${p} (${c})`)
.join(", ");
lines.push(`All tasks: ${summary || "none"}`);
const loopInfo = getLoopInfo(baseDir);
for (const l of loopInfo) {
lines.push(l);
}
return {
systemPrompt: `${_event.systemPrompt}\n\n## Automaton Context\n\n${lines.join("\n")}`,
};
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-automaton-context",
"version": "1.0.0",
"description": "Auto-injects Automaton framework context into Pi Dev system prompt — no more manual copy-paste of pi-automaton.sh output",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+171
View File
@@ -0,0 +1,171 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "@sinclair/typebox";
import { execSync } from "child_process";
import { existsSync, readFileSync, readdirSync } from "fs";
import { join } from "path";
import { homedir } from "os";
const STATUS_SCRIPT = join(homedir(), ".automaton", "scripts", "status.py");
function resolveProjectDir(cwd: string): string | null {
const inFramework = cwd === join(homedir(), ".automaton") || cwd.startsWith(join(homedir(), ".automaton") + "/");
if (inFramework) return homedir() + "/.automaton";
if (existsSync(join(cwd, ".automaton"))) return cwd;
return null;
}
function getTasks(projectDir: string) {
const tasksDir = join(projectDir, ".automaton", "tasks");
if (!existsSync(tasksDir)) return [];
const tasks: { name: string; phase: string }[] = [];
for (const entry of readdirSync(tasksDir)) {
const stateFile = join(tasksDir, entry, ".state");
if (existsSync(stateFile)) {
try {
const phase = readFileSync(stateFile, "utf-8").trim();
if (phase) tasks.push({ name: entry, phase });
} catch {}
}
}
tasks.sort((a, b) => a.name.localeCompare(b.name));
return tasks;
}
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "automaton_create_task",
label: "Create Automaton Task",
description:
"Create a new automaton task and transition it to the specified phase. " +
"Use this when the user asks you to do work that requires file edits — you need a task in 'implement' or 'doc_review' phase before you can edit files.",
promptSnippet: "Create automaton tasks for work management",
promptGuidelines: [
"When the user asks you to build, implement, fix, or change something, first call automaton_create_task to create a task and transition it to 'implement' phase",
"Only after the task is in 'implement' can you edit files — the guard will block edits otherwise",
"Use a descriptive task name based on what the user wants (e.g., 'add-login-page', 'fix-api-timeout')",
"Default phase is 'implement' — omit phase for most cases",
],
parameters: Type.Object({
name: Type.String({ minLength: 1, description: "Task name (kebab-case, e.g. 'add-login-page')" }),
phase: Type.Optional(Type.String({ default: "implement", description: "Phase to transition to after creation" })),
}),
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { name, phase = "implement" } = params;
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project. No .automaton/ directory found." }],
details: {},
};
}
try {
const createOut = execSync(
`python3 ${STATUS_SCRIPT} --create-task ${JSON.stringify(name)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
const transitionOut = execSync(
`python3 ${STATUS_SCRIPT} --transition ${phase} --task ${JSON.stringify(name)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
return {
content: [{ type: "text", text: `Created task "${name}" and transitioned to "${phase}".` }],
details: { task: name, phase },
};
} catch (e: any) {
return {
content: [{
type: "text",
text: `Failed to create task: ${e.stderr || e.message || String(e)}`,
}],
details: { error: e.stderr || e.message, task: name },
};
}
},
});
pi.registerTool({
name: "automaton_transition",
label: "Transition Automaton Task",
description:
"Transition an existing automaton task to a new phase. " +
"Valid phases: research, decomposition, design, implement, test_design, testing, doc_review, complete. " +
"Use this to move a task forward (e.g. from research to implement).",
promptSnippet: "Transition automaton tasks between phases",
parameters: Type.Object({
task: Type.String({ minLength: 1, description: "Task name" }),
phase: Type.String({ minLength: 1, description: "Target phase" }),
}),
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { task, phase } = params;
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project." }],
details: {},
};
}
try {
execSync(
`python3 ${STATUS_SCRIPT} --transition ${phase} --task ${JSON.stringify(task)} --project ${JSON.stringify(projectDir)}`,
{ encoding: "utf-8", timeout: 10000 },
);
return {
content: [{ type: "text", text: `Task "${task}" transitioned to "${phase}".` }],
details: { task, phase },
};
} catch (e: any) {
return {
content: [{
type: "text",
text: `Failed to transition task: ${e.stderr || e.message || String(e)}`,
}],
details: { error: e.stderr || e.message, task, phase },
};
}
},
});
pi.registerTool({
name: "automaton_status",
label: "Automaton Status",
description:
"Show current automaton project status — all tasks, their phases, and any running loops. " +
"Call this to check what state things are in before deciding what to do.",
promptSnippet: "Check automaton project status",
parameters: Type.Object({}),
async execute(_toolCallId, _params, _signal, _onUpdate, _ctx) {
const projectDir = resolveProjectDir(process.cwd());
if (!projectDir) {
return {
content: [{ type: "text", text: "Not in an automaton-managed project." }],
details: {},
};
}
const tasks = getTasks(projectDir);
const editTasks = tasks.filter((t) => t.phase === "implement" || t.phase === "doc_review");
const currentTask = editTasks.length > 0 ? editTasks[0] : tasks.length > 0 ? tasks[0] : null;
const lines: string[] = [];
lines.push(`Project: ${projectDir.split("/").pop()}`);
if (currentTask) {
lines.push(`Current: "${currentTask.name}" (${currentTask.phase})`);
} else {
lines.push("No tasks. Create one with automaton_create_task.");
}
lines.push("");
for (const t of tasks) {
const marker = currentTask && t.name === currentTask.name ? ">" : " ";
lines.push(`${marker} ${t.name.padEnd(35)} ${t.phase}`);
}
return {
content: [{ type: "text", text: lines.join("\n") }],
details: { tasks: tasks.length, current: currentTask?.name || null },
};
},
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-automaton-tools",
"version": "1.0.0",
"description": "Registers automaton_create_task and automaton_transition tools so the agent can manage task lifecycle automatically",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+91
View File
@@ -0,0 +1,91 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "@sinclair/typebox";
const BIRD_BIN = "/opt/homebrew/bin/bird";
const TweetParams = Type.Object({
url: Type.String({ minLength: 1, description: "Tweet URL or numeric ID" }),
mode: Type.Optional(
Type.Union([
Type.Literal("read"),
Type.Literal("thread"),
Type.Literal("replies"),
]),
),
});
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "read_tweet",
label: "Read Tweet",
description:
"Read an X/Twitter tweet or thread using the local bird CLI. " +
"Returns tweet author, text, media, metrics, and timestamps. " +
"Use this instead of webfetch for x.com/twitter.com URLs.",
promptSnippet: "Read X/Twitter tweets with bird CLI",
promptGuidelines: [
"Use read_tweet when the user shares an x.com or twitter.com URL and asks what it says",
"Use mode='thread' for full conversation threads",
"Use mode='replies' to fetch replies to a tweet",
"If fetch fails with auth errors, ask the user to sign in to x.com in their browser and retry",
],
parameters: TweetParams,
async execute(_toolCallId, params, _signal, _onUpdate, _ctx) {
const { url, mode = "read" } = params;
try {
const { execSync } = await import("child_process");
const cmd = `${BIRD_BIN} ${mode} ${JSON.stringify(url)} --json`;
const stdout = execSync(cmd, { encoding: "utf-8", timeout: 15000 });
const parsed = JSON.parse(stdout);
const author = parsed.author?.name || parsed.author?.screen_name || "unknown";
const text = parsed.text || parsed.content || "";
const createdAt = parsed.created_at || "";
const retweetCount = parsed.metrics?.retweet_count ?? parsed.metrics?.retweets ?? 0;
const likeCount = parsed.metrics?.like_count ?? parsed.metrics?.likes ?? 0;
const replyCount = parsed.metrics?.reply_count ?? parsed.metrics?.replies ?? 0;
const mediaCount = parsed.media?.length ?? 0;
const threadCount = parsed.thread?.tweets?.length ?? 0;
const repliesCount = parsed.replies?.length ?? 0;
const lines: string[] = [];
lines.push(`Author: ${author}`);
if (createdAt) lines.push(`Posted: ${createdAt}`);
lines.push("");
lines.push(text);
if (retweetCount || likeCount || replyCount) {
lines.push("");
lines.push(`Retweets: ${retweetCount} Likes: ${likeCount} Replies: ${replyCount}`);
}
if (mediaCount) lines.push(`Media: ${mediaCount} attachment(s)`);
if (threadCount) lines.push(`Thread: ${threadCount} tweets total`);
if (repliesCount) lines.push(`Replies fetched: ${repliesCount}`);
return {
content: [
{ type: "text", text: lines.join("\n") },
{ type: "text", text: `\n--- raw ---\n${JSON.stringify(parsed, null, 2)}` },
],
details: {
author,
text: text.slice(0, 500),
url,
mode,
tweetCount: threadCount || repliesCount || 1,
},
};
} catch (e: any) {
const errMsg = e.stderr || e.message || String(e);
const hint = errMsg.includes("auth")
? " Sign in to x.com in your browser and retry."
: "";
return {
content: [{ type: "text", text: `Failed to fetch tweet: ${errMsg}${hint}` }],
details: { error: errMsg, url, mode },
};
}
},
});
}
+10
View File
@@ -0,0 +1,10 @@
{
"name": "pi-read-tweet",
"version": "1.0.0",
"description": "Read X/Twitter tweets using the local `bird` CLI — registered as a read_tweet tool for Pi Dev",
"main": "index.ts",
"type": "module",
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*"
}
}
+1 -1
View File
@@ -1,2 +1,2 @@
#!/usr/bin/env bash
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/Users/laptran/.automaton"
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/private/var/folders/f5/yv0dzbnx47x3yp8sc_2519gh0000gn/T/pytest-of-laptran/pytest-130/test_uninstall_via_disabled_re0"
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env python3
"""Probe opencode.json and localhost endpoints to produce a candidate models.json.
Usage:
python3 scripts/detect_models.py [--json] [--write]
Without --json, prints a human-readable report.
With --json, emits the candidate models.json to stdout as the last JSON line.
With --write, writes the candidate to ~/.automaton/models.json (idempotent,
never overwrites an existing file unless --force is also given).
Probing strategy (stdlib only):
1. Parse opencode.json (or opencode.jsonc) for configured provider+model pairs.
2. Probe localhost endpoints to find locally-running LLM servers:
- http://localhost:8080/v1/models (llama.cpp / generic OpenAI-compatible)
- http://localhost:11434/api/tags (Ollama)
- http://localhost:1234/v1/models (LM Studio)
- http://localhost:8000/v1/models (vLLM)
3. Merge results into a candidate models.json.
"""
from __future__ import annotations
import json
import os
import re
import sys
from pathlib import Path
from typing import Optional
AUTOMATON_DIR = Path.home() / ".automaton"
# ---------------------------------------------------------------------------
# opencode.json parsing
# ---------------------------------------------------------------------------
def _find_opencode_json() -> Optional[Path]:
"""Locate the opencode config file (opencode.json or opencode.jsonc)."""
candidates = [
Path.cwd() / "opencode.json",
Path.cwd() / "opencode.jsonc",
AUTOMATON_DIR / "opencode.json",
AUTOMATON_DIR / "opencode.jsonc",
Path.home() / ".opencode.json",
Path.home() / ".config" / "opencode" / "opencode.json",
Path.home() / ".config" / "opencode" / "opencode.jsonc",
]
for p in candidates:
if p.exists():
return p
return None
def _parse_opencode_models(config_path: Path) -> list[dict]:
"""Extract model entries from an opencode.json config.
Expected structure (common patterns):
{
"providers": {
"opencode": { "model": "glm-4.6", ... },
...
}
}
or a flatter:
{
"model": "glm-4.6",
...
}
"""
try:
content = config_path.read_text(encoding="utf-8")
except OSError:
return []
# Strip JSONC comments (// line comments only, sufficient for our use)
content = re.sub(r"//.*", "", content)
try:
data = json.loads(content)
except json.JSONDecodeError:
return []
if not isinstance(data, dict):
return []
models: list[dict] = []
seen: set[str] = set()
# Check top-level "model" field (single-model config)
single = data.get("model")
if isinstance(single, str) and single not in seen:
seen.add(single)
models.append({"name": single, "provider": "opencode", "context_window": None, "location": "remote"})
# Check providers dict
providers = data.get("providers") or {}
for prov_name, prov_cfg in providers.items():
if isinstance(prov_cfg, dict):
model_name = prov_cfg.get("model")
if isinstance(model_name, str) and model_name not in seen:
seen.add(model_name)
models.append({"name": model_name, "provider": prov_name, "context_window": None, "location": "remote"})
# Check "models" list (explicit model roster)
model_list = data.get("models")
if isinstance(model_list, list):
for entry in model_list:
if isinstance(entry, dict):
name = entry.get("name") or entry.get("model")
if isinstance(name, str) and name not in seen:
seen.add(name)
models.append({
"name": name,
"provider": entry.get("provider", "opencode"),
"context_window": entry.get("context_window"),
"location": entry.get("location", "remote"),
})
return models
# ---------------------------------------------------------------------------
# Localhost probing
# ---------------------------------------------------------------------------
def _fetch_json(url: str, timeout: int = 5) -> Optional[dict]:
"""Fetch a JSON response from a URL using urllib (stdlib)."""
import urllib.request
import urllib.error
try:
req = urllib.request.Request(url, method="GET")
with urllib.request.urlopen(req, timeout=timeout) as resp:
body = resp.read().decode("utf-8")
return json.loads(body)
except (OSError, urllib.error.URLError, json.JSONDecodeError, ValueError):
return None
def _probe_ollama() -> list[dict]:
"""Probe Ollama: GET http://localhost:11434/api/tags → models[].name"""
data = _fetch_json("http://localhost:11434/api/tags")
if not data:
return []
models_list = data.get("models") or []
return [
{"name": m.get("name"), "provider": "ollama", "context_window": None, "location": "http://localhost:11434"}
for m in models_list
if isinstance(m, dict) and isinstance(m.get("name"), str)
]
def _probe_openai_compatible(url: str, provider: str) -> list[dict]:
"""Probe an OpenAI-compatible /v1/models endpoint."""
data = _fetch_json(url)
if not data:
return []
model_list = data.get("data") or []
return [
{"name": m.get("id"), "provider": provider, "context_window": None, "location": url}
for m in model_list
if isinstance(m, dict) and isinstance(m.get("id"), str)
]
_ENDPOINTS = [
("http://localhost:8080/v1/models", "llama.cpp"),
("http://localhost:11434/api/tags", "ollama"), # handled separately above
("http://localhost:1234/v1/models", "lm-studio"),
("http://localhost:8000/v1/models", "vllm"),
]
def _probe_localhost() -> list[dict]:
"""Probe all known localhost endpoints and merge results."""
seen_names: set[str] = set()
models: list[dict] = []
for url, provider in _ENDPOINTS:
if provider == "ollama":
result = _probe_ollama()
else:
result = _probe_openai_compatible(url, provider)
for m in result:
n = m.get("name")
if isinstance(n, str) and n not in seen_names:
seen_names.add(n)
models.append(m)
return models
# ---------------------------------------------------------------------------
# Merge & write
# ---------------------------------------------------------------------------
def build_candidate_models(probe_local: bool = True) -> dict:
"""Build a candidate models.json dict.
1. Parse models from opencode.json
2. Optionally probe localhost endpoints
3. Merge: opencode config models come first; local probes fill in gaps.
4. Build result with default, advised, models[].
"""
opencode_path = _find_opencode_json()
config_models: list[dict] = []
if opencode_path:
config_models = _parse_opencode_models(opencode_path)
local_models: list[dict] = []
if probe_local:
local_models = _probe_localhost()
# Merge: key by name, config models take priority (unordered)
merged: dict[str, dict] = {}
for m in config_models:
n = m["name"]
if n not in merged:
merged[n] = m
for m in local_models:
n = m.get("name")
if n and n not in merged:
merged[n] = m
models_list = list(merged.values())
# Determine default: first config model, or first local model, or empty
default_name: Optional[str] = None
if config_models:
default_name = config_models[0].get("name")
elif local_models:
default_name = local_models[0].get("name")
# Determine advised: if only 0-1 models, set advised=true; else false
advised = len(models_list) <= 1
result: dict = {
"schema_version": 1,
"default": default_name,
"advised": advised,
"models": models_list,
}
return result
def write_models_file(candidate: dict, force: bool = False) -> bool:
"""Write candidate models.json to AUTOMATON_DIR.
Never overwrites an existing file unless force=True.
Returns True if written, False if skipped.
"""
target = AUTOMATON_DIR / "models.json"
if target.exists() and not force:
return False
target.write_text(json.dumps(candidate, indent=2) + "\n")
return True
def format_report(candidate: dict) -> str:
"""Human-readable report of the candidate models."""
lines = []
lines.append("=== Model Detection Report ===")
lines.append("")
source = "No opencode.json found" if not _find_opencode_json() else f"Config: {_find_opencode_json()}"
lines.append(f"Source: {source}")
lines.append("")
models = candidate.get("models", [])
if not models:
lines.append("No models detected.")
else:
lines.append(f"Detected {len(models)} model(s):")
for m in models:
loc = m.get("location", "unknown")
prov = m.get("provider", "?")
ctx = m.get("context_window")
ctx_str = f", context: {ctx}" if ctx else ""
lines.append(f" - {m['name']} ({prov}, {loc}{ctx_str})")
lines.append("")
lines.append(f"Default: {candidate.get('default', 'none')}")
lines.append(f"Advised: {candidate.get('advised', False)}")
lines.append(f"Mode: {'multi-LLM' if len(models) >= 2 else 'single-LLM'}")
lines.append("")
target = AUTOMATON_DIR / "models.json"
if target.exists():
lines.append(f"models.json already exists at {target} (use --force to overwrite)")
else:
lines.append(f"Ready to write to {target} (use --write to create)")
return "\n".join(lines)
def main() -> int:
import argparse
parser = argparse.ArgumentParser(description="Detect available LLM models and write models.json")
parser.add_argument("--json", action="store_true", help="Output candidate JSON on last line")
parser.add_argument("--write", action="store_true", help="Write candidate models.json to ~/.automaton/ (idempotent)")
parser.add_argument("--force", action="store_true", help="Overwrite existing models.json")
parser.add_argument("--no-probe", action="store_true", help="Skip localhost endpoint probing")
args = parser.parse_args()
candidate = build_candidate_models(probe_local=not args.no_probe)
if args.write:
written = write_models_file(candidate, force=args.force)
if written:
print(f"Written models.json to {AUTOMATON_DIR / 'models.json'}")
else:
print(f"Skipped: {AUTOMATON_DIR / 'models.json'} already exists (use --force to overwrite)")
if args.json:
print(json.dumps(candidate))
else:
print(format_report(candidate))
return 0
if __name__ == "__main__":
sys.exit(main())
+5
View File
@@ -3,9 +3,14 @@
#
# Usage: bash ~/.automaton/scripts/install-hooks.sh [project-path]
#
# Called automatically by onboard-project.sh. Can also be run manually
# after framework updates to refresh hooks.
#
# Installs pre-commit and pre-push hooks. The pre-commit hook blocks
# commits when no task is in implement/doc_review. The pre-push hook
# blocks pushes in the same condition, catching --no-verify bypasses.
#
# Next step: python3 ~/.automaton/scripts/status.py --create-task --project .
set -euo pipefail
+29 -23
View File
@@ -3,27 +3,39 @@ set -e
FRAMEWORK_DIR="$HOME/.automaton"
if [ -d "$FRAMEWORK_DIR" ]; then
echo "automaton already installed at $FRAMEWORK_DIR"
echo "Run './update.sh' to update."
exit 0
fi
GIT_URL="${1:-}"
if [ ! -d "$FRAMEWORK_DIR" ]; then
# Fresh install — need a Git URL to clone
if [ -z "$GIT_URL" ]; then
echo "ERROR: Git URL required."
echo "Usage: ./install.sh <git-url>"
echo "Example: ./install.sh https://github.com/user/automaton.git"
echo "ERROR: Git URL required for fresh install."
echo ""
echo "Usage:"
echo " curl -fsSL https://raw.githubusercontent.com/<user>/automaton/main/scripts/install.sh | bash -s -- <git-url>"
echo ""
echo " --- or ---"
echo ""
echo " git clone <git-url> ~/.automaton"
echo " bash ~/.automaton/scripts/install.sh"
echo ""
echo "The framework is cloned to ~/.automaton. Choose your URL carefully"
echo "as it cannot be changed later without reinstalling."
exit 1
fi
echo "Cloning automaton from $GIT_URL to $FRAMEWORK_DIR..."
git clone "$GIT_URL" "$FRAMEWORK_DIR"
echo ""
else
echo "automaton already installed at $FRAMEWORK_DIR — running setup steps..."
if [ -n "$GIT_URL" ]; then
echo "Note: Git URL argument ignored because ~/.automaton already exists."
echo "To update, run: cd ~/.automaton && ./update.sh"
fi
echo ""
fi
# --- Everything below is idempotent and runs on both fresh and existing installs ---
echo "=== VRAM / Context Detection ==="
echo "Detecting your system's VRAM to recommend task decomposition settings..."
echo ""
@@ -31,7 +43,7 @@ echo ""
# Run VRAM detection script if it exists
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
# Run in project-dir context so it can read framework overhead
detection_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>&1)
detection_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>&1 || true)
# Extract JSON output (the block after "=== JSON Output ===")
json_output=$(echo "$detection_output" | sed -n '/=== JSON Output ===/,$p' | tail -n +2)
@@ -40,12 +52,9 @@ if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
echo "$detection_output"
# Extract key values from JSON using Python
recommended_k=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["recommended_k"])')
max_peak_kb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["max_peak_context_kb"])')
headroom=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["headroom"])')
gpu_vram=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["gpu_vram_gb"])')
ram_gb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["ram_gb"])')
model_context=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["model_context_kb"])')
recommended_k=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["recommended_k"])' 2>/dev/null || echo "?")
max_peak_kb=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["max_peak_context_kb"])' 2>/dev/null || echo "?")
headroom=$(echo "$json_output" | python3 -c 'import json,sys; print(json.load(sys.stdin)["headroom"])' 2>/dev/null || echo "?")
echo ""
echo "=== Recommended VRAM Configuration ==="
@@ -72,11 +81,8 @@ echo ""
echo "Installation complete."
echo ""
echo "Next steps:"
echo " 1. cd into a project and run the onboarding prompt"
echo " 2. In each project that uses git, install the automaton hooks:"
echo " bash ~/.automaton/scripts/install-hooks.sh /path/to/project"
echo ""
echo "These hooks block commits and pushes when no task is in an edit-allowed phase."
echo " 1. Onboard a project: bash ~/.automaton/scripts/onboard-project.sh /path/to/project"
echo " 2. Or tell your agent: 'Onboard this project into automaton'"
echo ""
# Register pre-edit guards for detected harnesses
+51 -12
View File
@@ -375,6 +375,9 @@ def _invoke_harness(
"""Build the harness command from loop.json and invoke it. Returns stdout.
extras: substitution tokens specific to this role ({artifact}, {verdict}, etc).
If extras contains a "model" key, the ``{model}`` token in the harness
command is substituted. The caller is responsible for passing the model
via extras (extracted from loop.json role config or manifest default).
"""
resolved_prompt = prompt_path
if loop_path is not None:
@@ -524,6 +527,24 @@ def _role_prompt(cfg: dict, role: str) ->Optional[str]:
return role_cfg.get("prompt")
def _role_model(cfg: dict, role: str) -> Optional[str]:
"""Get the model configured for a role in loop.json, or the manifest default."""
roles = cfg.get("roles") or {}
role_cfg = roles.get(role) or {}
model = role_cfg.get("model")
if model:
return model
models_file = AUTOMATON_DIR / "models.json"
if models_file.exists():
try:
import json as _mj
manifest = _mj.loads(models_file.read_text())
return manifest.get("default")
except (OSError, _mj.JSONDecodeError):
pass
return None
# ---------------------------------------------------------------------------
# Work sources (task add-goal-mode)
# ---------------------------------------------------------------------------
@@ -809,27 +830,39 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
out_dir = _outputs_dir(loop_path)
tick_num = state.get('iteration_count', 0) + 1
impl_output = str(out_dir / f"tick{tick_num}-implement.json")
implement_stdout = _invoke_harness(
harness_cfg, "implement", implement_prompt, cwd,
extras={"output": impl_output,
impl_model = _role_model(cfg, "implement")
implement_extras = {
"output": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
"next_hint": next_hint,
}
if impl_model:
implement_extras["model"] = impl_model
implement_stdout = _invoke_harness(
harness_cfg, "implement", implement_prompt, cwd,
extras=implement_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(impl_output)).write_text(implement_stdout)
# Step 6: spawn Verify
verify_prompt = _role_prompt(cfg, "verify") or ""
verify_output = str(out_dir / f"tick{tick_num}-verify.json")
verify_stdout = _invoke_harness(
harness_cfg, "verify", verify_prompt, cwd,
extras={"output": verify_output,
verify_model = _role_model(cfg, "verify")
verify_extras = {
"output": verify_output,
"artifact": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
"next_hint": next_hint,
}
if verify_model:
verify_extras["model"] = verify_model
verify_stdout = _invoke_harness(
harness_cfg, "verify", verify_prompt, cwd,
extras=verify_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(verify_output)).write_text(verify_stdout)
@@ -854,12 +887,18 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
# Step 9: spawn Orchestrate
orch_prompt = _role_prompt(cfg, "orchestrate") or ""
orch_output = str(out_dir / f"tick{tick_num}-orchestrate.json")
orch_stdout = _invoke_harness(
harness_cfg, "orchestrate", orch_prompt, cwd,
extras={"output": orch_output,
orch_model = _role_model(cfg, "orchestrate")
orch_extras = {
"output": orch_output,
"verdict": json.dumps(verdict),
"current_task": current_task,
"current_phase": state.get("current_phase", "")},
"current_phase": state.get("current_phase", ""),
}
if orch_model:
orch_extras["model"] = orch_model
orch_stdout = _invoke_harness(
harness_cfg, "orchestrate", orch_prompt, cwd,
extras=orch_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(orch_output)).write_text(orch_stdout)
+184
View File
@@ -0,0 +1,184 @@
#!/usr/bin/env bash
# onboard-project.sh — Bootstrap automaton in a new or existing project.
#
# Usage:
# bash ~/.automaton/scripts/onboard-project.sh /path/to/project
#
# Creates .automaton/ skeleton, detects models, creates config, inits git,
# installs hooks, and verifies everything works.
set -euo pipefail
FRAMEWORK_DIR="$HOME/.automaton"
PROJECT_DIR="${1:-}"
# Colors
RED='\033[0;31m'
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
BLUE='\033[0;34m'
NC='\033[0m' # No Color
info() { echo -e "${BLUE}INFO:${NC} $1"; }
ok() { echo -e "${GREEN}OK:${NC} $1"; }
warn() { echo -e "${YELLOW}WARN:${NC} $1"; }
error() { echo -e "${RED}ERROR:${NC} $1"; }
# --- Argument checks ---
if [ -z "$PROJECT_DIR" ]; then
error "Usage: bash ~/.automaton/scripts/onboard-project.sh /path/to/project"
exit 1
fi
PROJECT_DIR="$(cd "$PROJECT_DIR" 2>/dev/null && pwd)" || true
if [ -z "$PROJECT_DIR" ] || [ ! -d "$PROJECT_DIR" ]; then
echo ""
error "'$1' does not exist."
echo " Create it first: mkdir -p '$1'"
echo " Then re-run this script."
exit 1
fi
if [ ! -d "$FRAMEWORK_DIR/scripts" ]; then
error "Framework not found at $FRAMEWORK_DIR."
echo " Install the framework first:"
echo " curl -fsSL https://raw.githubusercontent.com/<user>/automaton/main/scripts/install.sh | bash -s -- <git-url>"
echo " Or: git clone <git-url> ~/.automaton && bash ~/.automaton/scripts/install.sh"
exit 1
fi
echo ""
echo "========================================"
echo " Automaton Project Onboarding"
echo " Project: $PROJECT_DIR"
echo "========================================"
echo ""
# --- Step 1: Create .automaton/ skeleton ---
AUTO_DIR="$PROJECT_DIR/.automaton"
if [ -d "$AUTO_DIR" ]; then
warn "$AUTO_DIR already exists — skipping skeleton creation"
else
info "Creating .automaton/ skeleton..."
mkdir -p "$AUTO_DIR/tasks" "$AUTO_DIR/loops" "$AUTO_DIR/design"
ok "Created $AUTO_DIR/"
fi
# --- Step 2: Models ---
MODELS_FILE="$AUTO_DIR/models.json"
if [ -f "$MODELS_FILE" ]; then
warn "$MODELS_FILE already exists — skipping model detection"
else
info "Probing local models..."
if python3 "$FRAMEWORK_DIR/scripts/detect_models.py" --write --project "$PROJECT_DIR" 2>/dev/null; then
ok "Detected models written to $MODELS_FILE"
else
info "Auto-detection failed. Creating minimal models.json..."
cat > "$MODELS_FILE" <<- 'EOF'
{
"models": [
{"name": "default-model", "provider": "local", "context": 32768}
],
"default": "default-model"
}
EOF
warn "Edit $MODELS_FILE to set your actual model(s)."
fi
fi
# --- Step 3: Config ---
CONFIG_FILE="$AUTO_DIR/config.md"
if [ -f "$CONFIG_FILE" ]; then
warn "$CONFIG_FILE already exists — skipping"
else
info "Creating config.md..."
PROJECT_NAME="$(basename "$PROJECT_DIR")"
if [ -f "$FRAMEWORK_DIR/scripts/vram_detect.py" ]; then
vram_output=$(cd "$FRAMEWORK_DIR" && python3 "$FRAMEWORK_DIR/scripts/vram_detect.py" 2>/dev/null || true)
json_part=$(echo "$vram_output" | sed -n '/=== JSON Output ===/,$p' | tail -n +2)
recommended=$(echo "$json_part" | python3 -c 'import json,sys; print(json.load(sys.stdin).get("recommended_k","16"))' 2>/dev/null || echo "16")
else
recommended="16"
fi
cat > "$CONFIG_FILE" <<- EOF
# $PROJECT_NAME — Automaton Configuration
## VRAM Configuration
- **Auto-detect**: Yes
- **Target context**: ${recommended}k tokens
- **Headroom**: 25%
- **Max peak context per sub-task**: $((recommended * 3 / 4))k tokens
## Model Configuration
# Uses models.json for model divergence enforcement.
# Default model is read from models.json's "default" key.
EOF
ok "Created $CONFIG_FILE"
fi
# --- Step 4: Project name ---
NAME_FILE="$AUTO_DIR/project-name.md"
if [ -f "$NAME_FILE" ]; then
warn "$NAME_FILE already exists — skipping"
else
PROJECT_NAME="$(basename "$PROJECT_DIR")"
echo "$PROJECT_NAME" > "$NAME_FILE"
ok "Created $NAME_FILE ($PROJECT_NAME)"
fi
# --- Step 5: Git ---
GIT_DIR="$PROJECT_DIR/.git"
if [ -d "$GIT_DIR" ]; then
ok "Git repository already initialized"
else
info "Initializing git repository..."
cd "$PROJECT_DIR" && git init
ok "Git initialized"
fi
# --- Step 6: Git hooks ---
if [ -d "$GIT_DIR" ]; then
info "Installing git hooks..."
bash "$FRAMEWORK_DIR/scripts/install-hooks.sh" "$PROJECT_DIR"
fi
# --- Step 7: .gitignore ---
GITIGNORE="$PROJECT_DIR/.gitignore"
if [ -f "$GITIGNORE" ]; then
if ! grep -q ".automaton/tasks/" "$GITIGNORE" 2>/dev/null; then
echo "" >> "$GITIGNORE"
echo "# Automaton" >> "$GITIGNORE"
echo ".automaton/tasks/" >> "$GITIGNORE"
echo ".automaton/loops/*/worktree/" >> "$GITIGNORE"
warn "Added automaton entries to .gitignore"
fi
else
cat > "$GITIGNORE" <<- 'EOF'
# Automaton
.automaton/tasks/
.automaton/loops/*/worktree/
.automaton/loops/*/outputs/
EOF
ok "Created .gitignore with automaton entries"
fi
# --- Step 8: Verify ---
info "Verifying setup..."
cd "$FRAMEWORK_DIR"
python3 "$FRAMEWORK_DIR/scripts/status.py" --project "$PROJECT_DIR" --audit 2>&1 | head -5 || true
echo ""
echo "========================================"
echo -e "${GREEN} Onboarding complete!${NC}"
echo "========================================"
echo ""
echo " Project: $PROJECT_DIR"
echo " Config: $CONFIG_FILE"
echo " Models: $MODELS_FILE"
echo ""
echo " Next steps:"
echo " 1. cd $PROJECT_DIR"
echo " 2. Create a task:"
echo " python3 ~/.automaton/scripts/status.py --create-task my-first-task --project ."
echo " 3. Start working with your agent."
echo ""
+174
View File
@@ -0,0 +1,174 @@
#!/usr/bin/env bash
# pi-automaton.sh — Pi Dev automaton context printer.
#
# Prints automaton project/framework context for the user to paste as their
# first message to a Pi Dev agent. Does NOT launch pi.
#
# Usage:
# cd /path/to/project
# bash ~/.automaton/scripts/pi-automaton.sh
#
# Copy the output and paste it as your first message in Pi Dev.
set -euo pipefail
FRAMEWORK_DIR="$HOME/.automaton"
STATUS_PY="$FRAMEWORK_DIR/scripts/status.py"
# Colors
BOLD='\033[1m'
DIM='\033[2m'
NC='\033[0m'
info() { echo -e " $1"; }
dim() { echo -e " ${DIM}$1${NC}"; }
dim_nl(){ echo -e "${DIM}$1${NC}"; }
# --- Scope detection ---
CWD="$(pwd)"
if [ "$CWD" = "$FRAMEWORK_DIR" ] || [ "${CWD##"$FRAMEWORK_DIR"}" != "$CWD" ]; then
SCOPE="framework"
SCOPE_LABEL="framework mode (automaton itself)"
PROJECT_DIR="$FRAMEWORK_DIR"
elif [ -d "$CWD/.automaton" ]; then
SCOPE="project"
SCOPE_LABEL="project mode ($(basename "$CWD"))"
PROJECT_DIR="$CWD"
else
echo ""
echo "No automaton project detected in $CWD"
echo ""
echo "To onboard this project:"
echo " bash $FRAMEWORK_DIR/scripts/onboard-project.sh ."
echo ""
exit 1
fi
TASKS_DIR="$PROJECT_DIR/.automaton/tasks"
MODELS_FILE="$PROJECT_DIR/.automaton/models.json"
CONFIG_FILE="$PROJECT_DIR/.automaton/config.md"
AGENTS_FILE="$PROJECT_DIR/.automaton/AGENTS.md"
# --- Collect data ---
# Active tasks via status.py --audit --json
TASKS_JSON=""
if [ -f "$STATUS_PY" ] && [ -d "$TASKS_DIR" ]; then
TASKS_JSON=$(python3 "$STATUS_PY" --audit --json --project "$PROJECT_DIR" 2>/dev/null || true)
fi
# Default model
DEFAULT_MODEL=""
if [ -f "$MODELS_FILE" ]; then
DEFAULT_MODEL=$(python3 -c "
import json
with open('$MODELS_FILE') as f:
m = json.load(f)
print(m.get('default', ''))
" 2>/dev/null || echo "")
fi
# VRAM context snippet from config.md
CONFIG_SNIPPET=""
if [ -f "$CONFIG_FILE" ]; then
CONFIG_SNIPPET=$(grep -i 'context\|headroom\|target' "$CONFIG_FILE" 2>/dev/null | head -3 | sed 's/^/ /')
fi
# Project rules from AGENTS.md
RULES_TEXT=""
if [ -f "$AGENTS_FILE" ]; then
RULES_TEXT=$(grep -v -E '^#|^$' "$AGENTS_FILE" 2>/dev/null | head -10)
fi
# Can-edit status
CAN_EDIT_OUTPUT=""
if [ -f "$STATUS_PY" ]; then
CAN_EDIT_OUTPUT=$(python3 "$STATUS_PY" --can-edit --json --project "$PROJECT_DIR" 2>/dev/null || true)
fi
# --- Build task list ---
TASK_LINES=""
TASK_COUNT=0
if [ -n "$TASKS_JSON" ]; then
while IFS=$'\t' read -r name state edit_phase; do
if [ -n "$name" ]; then
FLAG=""
if [ "$edit_phase" = "true" ]; then
FLAG=" (edit allowed)"
elif [ "$state" != "backlog" ] && [ "$state" != "done" ] && [ "$state" != "blocked" ]; then
FLAG=" (read-only)"
fi
TASK_LINES+=" * $name\t\t$state$FLAG\n"
TASK_COUNT=$((TASK_COUNT + 1))
fi
done < <(echo "$TASKS_JSON" | python3 -c "
import json, sys
data = json.load(sys.stdin)
tasks = []
for t in data.get('tasks', []):
tasks.append((t['name'], t['state']))
# Map edit-eligible phases
EDIT_PHASES = {'implement', 'doc_review'}
for name, state in tasks:
ep = 'true' if state in EDIT_PHASES else 'false'
print(f'{name}\t{state}\t{ep}')
" 2>/dev/null || true)
fi
# --- Render ---
LINE="══════════════════════════════════════════════════"
SEP="──────────────────────────────────────────────────"
echo ""
echo -e "${BOLD}${LINE}${NC}"
echo -e "${BOLD} Automaton Context — $SCOPE_LABEL${NC}"
echo -e "${BOLD}${LINE}${NC}"
echo ""
echo -e "${BOLD}This project uses the automaton workflow framework.${NC}"
info "Tasks are tracked in $(basename "$PROJECT_DIR")/.automaton/tasks/"
info "and flow through phases:"
info "backlog → research → implement → code_review → bug_find → ... → complete"
echo ""
if [ -n "$TASK_LINES" ]; then
echo -e "${BOLD}Active tasks:${NC}"
echo -e "$TASK_LINES"
echo ""
fi
if [ -n "$DEFAULT_MODEL" ]; then
echo -e "${BOLD}Default model:${NC} $DEFAULT_MODEL"
echo ""
fi
if [ -n "$CONFIG_SNIPPET" ]; then
echo -e "${BOLD}Configuration:${NC}"
echo "$CONFIG_SNIPPET"
echo ""
fi
echo -e "${BOLD}To work on a task:${NC}"
info "python3 ~/.automaton/scripts/status.py --transition <phase> --task <name>"
echo ""
echo -e "${BOLD}To create a new task:${NC}"
info "python3 ~/.automaton/scripts/status.py --create-task <name>"
echo ""
if [ -n "$RULES_TEXT" ]; then
echo -e "${BOLD}Project rules (from AGENTS.md):${NC}"
echo "$RULES_TEXT" | head -5
echo ""
fi
echo -e "${BOLD}Important:${NC}"
info "Only modify files when a task is in ${BOLD}implement${NC} or ${BOLD}doc_review${NC} phase"
info "All phase transitions go through status.py"
info "The automaton-guard-pi plugin blocks edits outside allowed phases"
echo ""
echo -e "${DIM}$SEP${NC}"
dim_nl "Copy this entire block and paste it as your first message"
dim_nl "to the Pi Dev agent to provide automaton context."
echo -e "${DIM}$SEP${NC}"
echo ""
+40 -7
View File
@@ -54,17 +54,50 @@ else
echo "OpenCode: not detected (no ~/.config/opencode/opencode.json or .jsonc)"
fi
# Pi Dev guard
PI_SOURCE="$FRAMEWORK_DIR/plugins/automaton-guard-pi"
# Pi Dev extensions
PI_EXTENSIONS=(
"$FRAMEWORK_DIR/plugins/automaton-guard-pi"
"$FRAMEWORK_DIR/plugins/pi-read-tweet"
"$FRAMEWORK_DIR/plugins/pi-automaton-context"
"$FRAMEWORK_DIR/plugins/pi-automaton-tools"
)
if command -v pi &>/dev/null; then
INSTALLED=$(pi list 2>/dev/null | grep -c "automaton-guard-pi" || true)
for ext in "${PI_EXTENSIONS[@]}"; do
ext_name=$(basename "$ext")
INSTALLED=$(pi list 2>/dev/null | grep -c "$ext_name" || true)
if [ "$INSTALLED" -gt 0 ]; then
echo "Pi Dev: already installed"
echo "Pi Dev: $ext_name already installed"
else
echo "Pi Dev: installing guard extension..."
pi install "$PI_SOURCE" 2>&1 | sed 's/^/ /'
echo "Pi Dev: installing $ext_name..."
pi install "$ext" 2>&1 | sed 's/^/ /'
INSTALLED_PI=true
echo "Pi Dev: installed"
echo "Pi Dev: $ext_name installed"
fi
done
# Offer pi-automaton startup wrapper
PI_WRAPPER="$FRAMEWORK_DIR/scripts/pi-automaton.sh"
if [ -f "$PI_WRAPPER" ]; then
echo ""
echo "Pi Dev context wrapper available at:"
echo " $PI_WRAPPER"
echo ""
echo "Before starting a Pi Dev session, run this script to print"
echo "automaton project context that you can paste as your first message."
echo ""
echo " bash ~/.automaton/scripts/pi-automaton.sh"
echo ""
# Only prompt interactively if stdin is a terminal
BIN_DIR="$HOME/bin"
if [ -t 0 ] && [ ! -f "$BIN_DIR/pi-automaton" ]; then
echo -n "Symlink to ~/bin/pi-automaton for easier access? [Y/n] "
read -r REPLY
if [ -z "$REPLY" ] || [ "$REPLY" = "y" ] || [ "$REPLY" = "Y" ]; then
mkdir -p "$BIN_DIR"
ln -sf "$PI_WRAPPER" "$BIN_DIR/pi-automaton"
echo " Created $BIN_DIR/pi-automaton → $PI_WRAPPER"
echo " (ensure ~/bin is in your PATH)"
fi
fi
fi
else
echo "Pi Dev: not detected (pi not in PATH)"
+261 -2
View File
@@ -153,7 +153,19 @@ FORBIDDEN_ARTIFACTS = {
}
NON_ARTIFACT_FILES = {".state", ".state.tmp", ".state.lock", ".state.approvals",
".state.implementer", ".state.lastedit", "VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"}
".state.implementer", ".state.lastedit", ".state.models",
"VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"}
# Model-divergence enforcement
CONFLICT_MATRIX = {
"code_review": {"implement"},
"bug_find": {"implement"},
"adversarial_bug_find": {"implement", "bug_find"},
"referee": {"implement", "bug_find", "adversarial_bug_find"},
"loop-verify": {"loop-implement"},
}
MODELS_JSON_FILE = "models.json"
PHASE_PRIORITY = {
"referee": 12, "doc_review": 11, "adversarial_bug_find": 10,
@@ -438,6 +450,124 @@ def _lock_timeout_seconds(project: Optional[str] = None) -> int:
return val
# ---------------------------------------------------------------------------
# Model-divergence enforcement helpers
# ---------------------------------------------------------------------------
def _load_models_manifest(project: Optional[str] = None) -> Optional[dict]:
"""Load the models.json manifest for the given project.
Searches:
1. project/.automaton/models.json
2. ~/.automaton/models.json (fallback)
Returns None if no models.json exists (single-LLM mode, backward compatible).
"""
project_dir = _find_project_dir(project)
candidates = [
project_dir / ".automaton" / MODELS_JSON_FILE,
AUTOMATON_DIR / MODELS_JSON_FILE,
]
for path in candidates:
if path.exists():
try:
return json.loads(path.read_text())
except (OSError, json.JSONDecodeError):
return None
return None
def _get_model_mode(manifest: Optional[dict]) -> str:
"""Determine the model mode: 'single' or 'multi-llm'.
- Missing manifest → single-LLM (backward compatible)
- 0-1 models → single-LLM
- 2+ models → multi-LLM
"""
if manifest is None:
return "single"
models = manifest.get("models") or []
if len(models) >= 2:
return "multi-llm"
return "single"
def _check_conflict(state_models: dict, role: str, model: str, matrix: Optional[dict] = None) -> Optional[str]:
"""Check if the given model conflicts with already-filled roles.
state_models: dict of {role: model_name} from .state.models
role: the role being entered (e.g. 'code_review')
model: the model name being assigned
matrix: conflict matrix (defaults to CONFLICT_MATRIX)
Returns the name of the conflicting role, or None if no conflict.
"""
if matrix is None:
matrix = CONFLICT_MATRIX
if role not in matrix:
return None
conflicting_roles = matrix[role]
for filled_role, filled_model in state_models.items():
if filled_model == model and filled_role in conflicting_roles:
return filled_role
return None
def _read_state_models(task_path: Path) -> dict:
"""Read .state.models from the task directory. Returns {} if missing."""
f = task_path / ".state.models"
if not f.exists():
return {}
try:
data = json.loads(f.read_text())
if isinstance(data, dict):
return data
except (OSError, json.JSONDecodeError):
pass
return {}
def _write_state_models(task_path: Path, state_models: dict) -> None:
"""Write .state.models atomically."""
tmp = task_path / ".state.models.tmp"
tmp.write_text(json.dumps(state_models, indent=2, sort_keys=True) + "\n")
tmp.replace(task_path / ".state.models")
def _model_divergence_violations(project: Optional[str] = None) -> list[dict]:
"""Scan all tasks for model-divergence violations.
Returns list of violation dicts:
{"task": str, "message": str, "severity": "high", "resolved": False}
"""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode == "single":
return []
violations = []
tasks = _all_task_dirs(project)
for name, path in tasks:
sm = _read_state_models(path)
if not sm:
continue
for role, model in sm.items():
if model is None:
continue
# D8: doc_review, code_review, bug_find have no cross-conflicts
# with each other; only conflicts documented in CONFLICT_MATRIX apply.
conflict = _check_conflict(sm, role, str(model))
if conflict:
violations.append({
"task": name,
"severity": "high",
"message": f"model-divergence: role '{role}' uses model '{model}' "
f"which conflicts with role '{conflict}' (same model)",
"resolved": False,
})
return violations
# --- Command implementations ---
def cmd_show_task(args):
@@ -593,6 +723,60 @@ def cmd_transition(args):
return 1
if current == "human_intervention" and target == "complete":
_auto_update_verdict_on_complete(task_path)
# Model-divergence enforcement (Subtask 2)
# When entering a phase that maps to a role, record the model
target_base = _base_phase(target)
ROLE_PHASES = {"implement", "code_review", "bug_find", "adversarial_bug_find", "doc_review", "referee"}
if target_base in ROLE_PHASES and current != target:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
state_models = _read_state_models(task_path)
model_arg = getattr(args, "model", None)
if model_arg:
# --model explicitly provided — record advisory in single mode, check in multi
state_models[target_base] = model_arg
if mode == "multi-llm":
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 1
conflict = _check_conflict(state_models, target_base, model_arg)
if conflict:
# Remove the entry we just added
del state_models[target_base]
print(f"ERROR: Model '{model_arg}' assigned to role '{target_base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model> to specify a different model.")
return 1
elif mode == "multi-llm":
# Auto-assign: try default, then next-available non-conflicting
default = (manifest or {}).get("default")
assigned = False
if default and default in model_names:
conflict = _check_conflict(state_models, target_base, default)
if not conflict:
state_models[target_base] = default
assigned = True
if not assigned:
for m_name in model_names:
if m_name == default:
continue
conflict = _check_conflict(state_models, target_base, m_name)
if not conflict:
state_models[target_base] = m_name
assigned = True
break
if not assigned:
print(f"ERROR: Cannot auto-assign a model for role '{target_base}'. "
f"All available models conflict with already-filled roles. "
f"Use --model <name> to override.")
return 1
if model_arg or mode == "multi-llm":
_write_state_models(task_path, state_models)
if current == "implement" and target == "code_review":
lock_file = task_path / ".state.lock"
if lock_file.exists():
@@ -980,6 +1164,14 @@ def _audit_collect(args):
"halt_reason": lhalt, "current_task": ltask,
"violation": is_violation, "message": msg})
# Model-divergence violations (Category 6)
for mv in _model_divergence_violations(args.project):
violations.append({
"category": 6, "severity": mv["severity"],
"task": mv["task"], "message": mv["message"],
"resolved": False,
})
return {"violations": violations,
"loops": loops,
"total_tasks": len(tasks),
@@ -1118,7 +1310,16 @@ def cmd_audit(args):
if stuck_found == 0:
print(f"[PASS] No stuck tasks (threshold: {stuck_threshold} min)")
print("\n=== Category 6: Loops ===")
print("\n=== Category 6: Model-Divergence Violations ===")
md_violations = _model_divergence_violations(args.project)
if md_violations:
for v in md_violations:
print(f"[FAIL] {v['task']}: {v['message']}")
violations += 1
else:
print("[PASS] No model-divergence violations found")
print("\n=== Category 7: Loops ===")
violations += _audit_loops_block(args)
print(f"\n=== Summary ===")
@@ -1435,6 +1636,26 @@ def cmd_claim(args):
if implementer == args.agent:
print(f"ERROR: Agent '{args.agent}' implemented this task and cannot claim the code_review phase. Reviewer must be different from implementer.")
return 1
# Model-divergence check on claim (Subtask 2, multi-LLM only)
model_arg = getattr(args, "model", None)
if model_arg:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
if mode == "multi-llm":
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 2
state_models = _read_state_models(task_path)
conflict = _check_conflict(state_models, base, model_arg)
if conflict:
print(f"ERROR: Model '{model_arg}' for role '{base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model>.")
return 1
lock_file = task_path / ".state.lock"
timeout_sec = _lock_timeout_seconds(args.project)
if lock_file.exists():
@@ -2600,6 +2821,42 @@ def _gate_worktree_drift(state: dict, cfg: dict, project: Optional[str]) -> Opti
return None
def _gate_model_divergence(state: dict, cfg: dict, project: Optional[str]) -> Optional[dict]:
"""Model-divergence brake: in multi-LLM mode, verify and implement
roles must use different models. This prevents same-model verification
(rubber-stamping) within a loop tick."""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode != "multi-llm":
return None
roles = cfg.get("roles") or {}
impl_model = None
verify_model = None
impl_cfg = roles.get("implement") or {}
verify_cfg = roles.get("verify") or {}
impl_model = impl_cfg.get("model")
verify_model = verify_cfg.get("model")
# Fall back to manifest default if role has no explicit model
if not impl_model or not verify_model:
default = (manifest or {}).get("default")
if not impl_model:
impl_model = default
if not verify_model:
verify_model = default
if impl_model and verify_model and impl_model == verify_model:
return {
"ok": False,
"reason": "halted:model_conflict",
"halt_reason": "human_intervention",
"remaining_iterations": None,
"remaining_budget_usd": None,
"task_phase": None,
"task_in_halt_loop": True,
"out_of_scope_files": [],
}
return None
def _gate_score_plateau(state: dict, cfg: dict) -> Optional[dict]:
window = int(cfg.get("brakes", {}).get("score_plateau_window", 0))
if window <= 0:
@@ -2648,6 +2905,7 @@ def cmd_check_gate(args) -> int:
_gate_task_phase(state, cfg, args.project),
_gate_worktree_drift(state, cfg, args.project),
_gate_score_plateau(state, cfg),
_gate_model_divergence(state, cfg, args.project),
]
failure = next((g for g in gates if g is not None), None)
if failure is None:
@@ -2864,6 +3122,7 @@ def main():
parser.add_argument("--days", type=int, help="Cleanup age threshold in days (default 7, used with --cleanup-done / --install-cleanup-schedule)")
parser.add_argument("--dry-run", action="store_true", help="With --cleanup-done, list candidates without moving them")
parser.add_argument("--version", action="store_true", help="Print framework version and exit")
parser.add_argument("--model", metavar="NAME", help="Model name for model-divergence enforcement (used with --transition, --claim)")
args = parser.parse_args()

Some files were not shown because too many files have changed in this diff Show More