runnable-test-suite (parent) — complete. Three sub-tasks all complete:
- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
python3), add idempotent .venv install block to scripts/install.sh, and add
'from __future__ import annotations' to 3 dashboard modules using PEP 604
union syntax at definition time so they import on Python 3.9+. The
PEP 604 bug was caught by the streak verifier itself during implementation.
- vram-detect-cross-platform: scripts/vram_detect.py now branches on
platform.system() for Linux/Darwin/Windows. macOS path uses
system_profiler SPDisplaysDataType (Apple Silicon unified memory via
sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
/proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
model cards. Added _probe_ollama_model() that runs 'ollama list' as a
last-resort fallback. run_command() now wraps PowerShell cmdlets on
Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).
- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
detect_ram and detect_gpu_vram, prefix-match for unknown model names,
ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
mocked; no live hardware probes. Suite total: 235 passed, 0 errors.
Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.
Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.
Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
- Pre-commit hook: warns when autopilot is enabled and non-terminal tasks exist
- Post-commit hook: after commit, prints non-terminal task summary if autopilot on
- Post-commit exits 0 always (informational only, never blocks)
- Both hooks read .agent.md to detect Autopilot: Enabled
- Support ## Verdict: PASS, # VERDICT: PASS, VERDICT: PASS formats
- Add ## Status: PASS line to 16 old-format VERDICT.md files
- All 49 tasks now correctly detected as DONE by dashboard
- task.py: remove naive substring fallback from parse_verdict_status(),
only parse ## Status: header; no header → ambiguous (REFEREE)
- status.py: add _auto_update_verdict_on_complete() — when transitioning
human_intervention→complete, auto-update VERDICT.md to PASS
- status.py: add stale-task detection to --can-edit — deny edits if
all edit tasks have .state mtime >30 min old (reason: stale_task)
- status.py: add --touch command to reset task activity clock
- guard plugin: handle stale_task reason with specific error message
- Update test to match new verdict parsing behavior
- task.py: add phase_guidance, blocker, next_phase_name, required_artifact_name,
is_edit_phase, is_approval_gated properties to Task model
- dashboard.js: show phase_guidance and blocker in detail panel for all phases;
show status_reason on cards for all phases (was only blocked/bug_find/adv_bug_find);
add Edit Allowed and Requires Approval badges
- styles.css: style blocker warning, guidance info box, edit/approval badges
- app.py: include new fields in API JSON response
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks
Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)
Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)
Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit
- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM
- Standardize all prompts to .automaton/tasks/{task-name}/ path
- Reconcile dashboard spec with web implementation; remove themes.py
- Remove half-implemented refresh.py file watcher
- Harden dashboard static-file serving and task-name validation
- Add uncommitted-change guard to update.sh and real Gitea URLs
- Add AGENTS.md, Gitea CI workflow, and template documentation
- Drive all approved tasks to completion with VERDICT.md
- Fix state machine: IMPLEMENTATION.md was never checked in determine_task_state()
- Fix state machine: DOC_REVIEW.md priority wrong (checked after BUG_REPORT)
- Fix board display: approved planning tasks now advance to Design group
- Fix board display: rejected planning tasks move to Blocked group
- Fix path traversal: review API validated task names against ../ injection
- Fix URL encoding: unquote() task names in API path parsing
- Fix comment parsing: robust REVIEW.md read/write, handle falsy comments
- Fix dead code: KanbanBoard class missing COLUMNS and __init__
- Fix inotify: explicit error messages and polling fallback
- Fix review API: validate task names, prevent path traversal
- Update CHANGELOG.md with all changes