- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
level (all <7 days old per the cleanup policy; premature bulk archive
was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
the board via innerHTML every 2s, destroying each column-body's
scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
view.scrollTop before rebuild and restores after (matched by
PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
section cards, transition buttons, inline artifact editor (textarea for
writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
pointing at a pytest temp dir (test isolation leak). Rewired to point
at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
but a real install creates it. Now snapshots mtime before run, asserts
unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
board renders tasks, column scroll survives auto-refresh tick.
Verified the test fails without the scroll fix (scrollTop resets to 0).
Skipped via importorskip when playwright is absent (main CI stays
green).
- **Clarify SI loop scope in README** — new-project onboarding section
documents the framework-scoped self-improvement loop and options
(leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
gap (mde tasks marked complete but per-role model binding was never
implemented).
The review system (REVIEW.md, approve/changes_requested buttons,
review filter, review badges) was purely cosmetic — only the dashboard
read/wrote it. No workflow component (status.py, autopilot.py,
loop-runner.py, prompts) ever enforced it.
The 'Approve' button in the task detail panel confused users into
thinking it approved the task's phase gate. In reality it only wrote
to REVIEW.md, which had zero effect on transitions.
Removed:
- Review section (buttons, textarea, status badge) from detail panel
- Review badge from task cards
- Review filter from toolbar
- Pending-review counter from header
- All review-related CSS
Users now use the single '🔓 Approve Phase' button in the detail
panel, which calls status.py --approve and actually transitions the
task.
The dashboard had no way to run status.py --approve. The existing
'Approve' button only wrote to REVIEW.md (code review status), not
approve the task's phase gate.
Changes:
- /api/approve/{task_name} POST endpoint in app.py calls status.py --approve
- '🔓 Approve Phase' button in detail panel for approval-gated tasks
- approvePhase() JS function sends the POST and refreshes the board
- CSS for the new button (warning-yellow styling)
This enables the full automation flow: user clicks Approve Phase →
autopilot picks up the approved task on next tick → drives it forward.
--execute flag runs status.py --transition instead of printing suggestions.
--install-schedule creates a launchd plist that runs autopilot --drive --execute
every 60 seconds, providing full automation for task lifecycle transitions.
The autopilot loop:
- Runs on a schedule (60s default), ticks the most advanced unblocked task
- Stalls at :awaiting_approval gates until user runs --approve
- Handles all phases: new → research → decomposition → design → test_design →
implement → code_review → bug_find → adversarial_bug_find → doc_review →
referee → complete
- Fails cleanly when required artifacts are missing (agent must write them)
- Logs stdout/stderr to ~/.automaton/logs/autopilot-*.log
Also fixes test leak in scripts/automaton-cleanup.sh (pytest temp path was
being written into the real stub).
Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)
Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.
Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.
runnable-test-suite (parent) — complete. Three sub-tasks all complete:
- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
python3), add idempotent .venv install block to scripts/install.sh, and add
'from __future__ import annotations' to 3 dashboard modules using PEP 604
union syntax at definition time so they import on Python 3.9+. The
PEP 604 bug was caught by the streak verifier itself during implementation.
- vram-detect-cross-platform: scripts/vram_detect.py now branches on
platform.system() for Linux/Darwin/Windows. macOS path uses
system_profiler SPDisplaysDataType (Apple Silicon unified memory via
sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
/proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
model cards. Added _probe_ollama_model() that runs 'ollama list' as a
last-resort fallback. run_command() now wraps PowerShell cmdlets on
Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).
- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
detect_ram and detect_gpu_vram, prefix-match for unknown model names,
ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
mocked; no live hardware probes. Suite total: 235 passed, 0 errors.
Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.
Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.
Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
- Pre-commit hook: warns when autopilot is enabled and non-terminal tasks exist
- Post-commit hook: after commit, prints non-terminal task summary if autopilot on
- Post-commit exits 0 always (informational only, never blocks)
- Both hooks read .agent.md to detect Autopilot: Enabled
- Support ## Verdict: PASS, # VERDICT: PASS, VERDICT: PASS formats
- Add ## Status: PASS line to 16 old-format VERDICT.md files
- All 49 tasks now correctly detected as DONE by dashboard
- task.py: remove naive substring fallback from parse_verdict_status(),
only parse ## Status: header; no header → ambiguous (REFEREE)
- status.py: add _auto_update_verdict_on_complete() — when transitioning
human_intervention→complete, auto-update VERDICT.md to PASS
- status.py: add stale-task detection to --can-edit — deny edits if
all edit tasks have .state mtime >30 min old (reason: stale_task)
- status.py: add --touch command to reset task activity clock
- guard plugin: handle stale_task reason with specific error message
- Update test to match new verdict parsing behavior
- task.py: add phase_guidance, blocker, next_phase_name, required_artifact_name,
is_edit_phase, is_approval_gated properties to Task model
- dashboard.js: show phase_guidance and blocker in detail panel for all phases;
show status_reason on cards for all phases (was only blocked/bug_find/adv_bug_find);
add Edit Allowed and Requires Approval badges
- styles.css: style blocker warning, guidance info box, edit/approval badges
- app.py: include new fields in API JSON response
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks
Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)
Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)
Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit
- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM
- Standardize all prompts to .automaton/tasks/{task-name}/ path
- Reconcile dashboard spec with web implementation; remove themes.py
- Remove half-implemented refresh.py file watcher
- Harden dashboard static-file serving and task-name validation
- Add uncommitted-change guard to update.sh and real Gitea URLs
- Add AGENTS.md, Gitea CI workflow, and template documentation
The agent scanned tasks in invest-copilot while working on automaton
framework, wasting time on unrelated work. Added a rule with concrete
failure example limiting work to the current project's .automaton/tasks/
only.
The orchestrator only drove ONE task per invocation, then stopped with
ORCHESTRATION_COMPLETE. User had to manually re-trigger for each task.
Changes:
- Added drive_all() outer loop that scans EVERY task and drives them all
- Tasks needing user review are flagged; orchestrator continues to next
- Only stops when ALL tasks are terminal or ALL remaining need user input
- Output format now shows session summary (completed, awaiting review, blocked)
- Session-starter references 'orchestrate' as primary command
- Updated continue-from-existing to use drive_all()
Full M3 redesign with Google's color scheme, elevation system,
and component styling.
Color: primary (#8ab4f8 dark / #1a73e8 light), error (#f28b82),
surface tones with proper dark/light contrast
Elevation: M3 shadow system (levels 0-5) with proper opacity per
color scheme
Typography: Roboto / system-ui stack
Components: M3 card radi (8px), column radi (12px), modal radi (16px)
Focus: M3 focus ring (2px with primary-alpha 0.2)
Cards: accent bar, subtle hover elevation lift
update.sh now checks the current project after pulling latest changes.
If it detects stale prompt/contract/script copies or tasks at the
deprecated root-level location, it prompts the user to run migration.
No more needing to know about migrate-project.sh separately.
Both framework and project tasks now live in .automaton/tasks/.
Every project has a .automaton/ directory, so no need for special
scope-based path logic. Removed the dual-convention gap.
Changes:
- Dashboard _find_tasks_dir() and _get_review_path() always use
{root}/.automaton/tasks/ — no scope branching
- migrate-project.sh: moves tasks/ -> .automaton/tasks/ during upgrade
- Orchestrator prompt: all task path references updated to
{project}/.automaton/tasks/
The additive extension model refactor moved prompts/contracts/scripts into
.automaton/ but never addressed where tasks live. Two conventions existed:
framework keeps tasks in .automaton/tasks/, Orchestrator creates project
tasks at tasks/.
Fixes:
- Dashboard _find_tasks_dir() now scope-aware: framework mode prefers
.automaton/tasks/, project mode prefers tasks/ at project root
- migrate-project.sh: if tasks exist in .automaton/tasks/ (old model),
move them to tasks/ (project root)
- Orchestrator prompt: documents that tasks location depends on scope
Dashboard was hardcoded to read tasks from .automaton/tasks/, but the
Orchestrator creates tasks at tasks/ (project root). Added _find_tasks_dir()
that checks both locations and _get_review_path() that reads REVIEW.md
from whichever location the task lives in.
Priority: .automaton/tasks/ first, then tasks/ as fallback.
This handles framework mode (tasks in .automaton/), project mode
(tasks at root), and symlinked projects.
- Add automaton.pth to user site-packages so python -m automaton.dashboard
works from any working directory, not just ~/.automaton/
- Add scripts/dashboard.sh as a convenience wrapper with PYTHONPATH
- Drive all approved tasks to completion with VERDICT.md
- Fix state machine: IMPLEMENTATION.md was never checked in determine_task_state()
- Fix state machine: DOC_REVIEW.md priority wrong (checked after BUG_REPORT)
- Fix board display: approved planning tasks now advance to Design group
- Fix board display: rejected planning tasks move to Blocked group
- Fix path traversal: review API validated task names against ../ injection
- Fix URL encoding: unquote() task names in API path parsing
- Fix comment parsing: robust REVIEW.md read/write, handle falsy comments
- Fix dead code: KanbanBoard class missing COLUMNS and __init__
- Fix inotify: explicit error messages and polling fallback
- Fix review API: validate task names, prevent path traversal
- Update CHANGELOG.md with all changes