Pi Dev doesn't auto-load automaton's system prompt (unlike opencode),
so the agent has no awareness of tasks, phases, or status.py commands.
This bridges that gap:
- scripts/pi-automaton.sh: NEW — detects scope (framework/project),
reads active tasks, default model, config snippet, AGENTS.md rules,
and prints a formatted context block the user pastes as their first
message to the Pi Dev agent. Does NOT launch pi.
- scripts/register-guards.sh: EXTENDED — after Pi Dev guard install,
offers to symlink pi-automaton.sh to ~/bin/pi-automaton
(interactive prompt only when stdin is a terminal).
- README.md: added 'Using Pi Dev with automaton' FAQ subsection
documenting the context printer workflow.
- install.sh: no longer exits early when ~/.automaton exists.
Skips the clone but runs all setup (VRAM detection, guards,
self-improvement loop, virtualenv). Both curl|bash and
git-clone + ./install.sh now work correctly.
- onboard-project.sh: new script that bootstraps automaton in a
project — creates .automaton/ skeleton, detects models, writes
config.md, inits git, installs hooks, adds .gitignore entries.
- README.md: fix install flow docs (curl|bash + clone-then-run),
add onboard-project.sh as Option A for project setup
- task.py: Task dataclass gains models: dict[str, str], loaded from
.state.models in discover_tasks()
- app.py: models dict included in all task API responses
- dashboard.js: model badges rendered between artifacts and subtask
progress on kanban cards; ROLE_LABELS map for readable tooltips
- styles.css: .task-card-models and .model-badge styles
- design/loops/technical.md: document {model} substitution token,
per-role model field, and model-divergence brake gate (gate #7)
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
level (all <7 days old per the cleanup policy; premature bulk archive
was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
the board via innerHTML every 2s, destroying each column-body's
scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
view.scrollTop before rebuild and restores after (matched by
PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
section cards, transition buttons, inline artifact editor (textarea for
writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
pointing at a pytest temp dir (test isolation leak). Rewired to point
at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
but a real install creates it. Now snapshots mtime before run, asserts
unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
board renders tasks, column scroll survives auto-refresh tick.
Verified the test fails without the scroll fix (scrollTop resets to 0).
Skipped via importorskip when playwright is absent (main CI stays
green).
- **Clarify SI loop scope in README** — new-project onboarding section
documents the framework-scoped self-improvement loop and options
(leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
gap (mde tasks marked complete but per-role model binding was never
implemented).
--execute flag runs status.py --transition instead of printing suggestions.
--install-schedule creates a launchd plist that runs autopilot --drive --execute
every 60 seconds, providing full automation for task lifecycle transitions.
The autopilot loop:
- Runs on a schedule (60s default), ticks the most advanced unblocked task
- Stalls at :awaiting_approval gates until user runs --approve
- Handles all phases: new → research → decomposition → design → test_design →
implement → code_review → bug_find → adversarial_bug_find → doc_review →
referee → complete
- Fails cleanly when required artifacts are missing (agent must write them)
- Logs stdout/stderr to ~/.automaton/logs/autopilot-*.log
Also fixes test leak in scripts/automaton-cleanup.sh (pytest temp path was
being written into the real stub).
runnable-test-suite (parent) — complete. Three sub-tasks all complete:
- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
python3), add idempotent .venv install block to scripts/install.sh, and add
'from __future__ import annotations' to 3 dashboard modules using PEP 604
union syntax at definition time so they import on Python 3.9+. The
PEP 604 bug was caught by the streak verifier itself during implementation.
- vram-detect-cross-platform: scripts/vram_detect.py now branches on
platform.system() for Linux/Darwin/Windows. macOS path uses
system_profiler SPDisplaysDataType (Apple Silicon unified memory via
sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
/proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
model cards. Added _probe_ollama_model() that runs 'ollama list' as a
last-resort fallback. run_command() now wraps PowerShell cmdlets on
Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).
- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
detect_ram and detect_gpu_vram, prefix-match for unknown model names,
ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
mocked; no live hardware probes. Suite total: 235 passed, 0 errors.
Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.
Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.
Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
- Pre-commit hook: warns when autopilot is enabled and non-terminal tasks exist
- Post-commit hook: after commit, prints non-terminal task summary if autopilot on
- Post-commit exits 0 always (informational only, never blocks)
- Both hooks read .agent.md to detect Autopilot: Enabled
- task.py: remove naive substring fallback from parse_verdict_status(),
only parse ## Status: header; no header → ambiguous (REFEREE)
- status.py: add _auto_update_verdict_on_complete() — when transitioning
human_intervention→complete, auto-update VERDICT.md to PASS
- status.py: add stale-task detection to --can-edit — deny edits if
all edit tasks have .state mtime >30 min old (reason: stale_task)
- status.py: add --touch command to reset task activity clock
- guard plugin: handle stale_task reason with specific error message
- Update test to match new verdict parsing behavior
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks
Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)
Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)
Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit
- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM
- Standardize all prompts to .automaton/tasks/{task-name}/ path
- Reconcile dashboard spec with web implementation; remove themes.py
- Remove half-implemented refresh.py file watcher
- Harden dashboard static-file serving and task-name validation
- Add uncommitted-change guard to update.sh and real Gitea URLs
- Add AGENTS.md, Gitea CI workflow, and template documentation
update.sh now checks the current project after pulling latest changes.
If it detects stale prompt/contract/script copies or tasks at the
deprecated root-level location, it prompts the user to run migration.
No more needing to know about migrate-project.sh separately.
Both framework and project tasks now live in .automaton/tasks/.
Every project has a .automaton/ directory, so no need for special
scope-based path logic. Removed the dual-convention gap.
Changes:
- Dashboard _find_tasks_dir() and _get_review_path() always use
{root}/.automaton/tasks/ — no scope branching
- migrate-project.sh: moves tasks/ -> .automaton/tasks/ during upgrade
- Orchestrator prompt: all task path references updated to
{project}/.automaton/tasks/
The additive extension model refactor moved prompts/contracts/scripts into
.automaton/ but never addressed where tasks live. Two conventions existed:
framework keeps tasks in .automaton/tasks/, Orchestrator creates project
tasks at tasks/.
Fixes:
- Dashboard _find_tasks_dir() now scope-aware: framework mode prefers
.automaton/tasks/, project mode prefers tasks/ at project root
- migrate-project.sh: if tasks exist in .automaton/tasks/ (old model),
move them to tasks/ (project root)
- Orchestrator prompt: documents that tasks location depends on scope
- Add automaton.pth to user site-packages so python -m automaton.dashboard
works from any working directory, not just ~/.automaton/
- Add scripts/dashboard.sh as a convenience wrapper with PYTHONPATH