Files
automaton/CHANGELOG.md
T
Lap Tran f32f98575b
CI / build (push) Has been cancelled
Make test suite runnable from clean checkout + cross-platform vram_detect
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00

14 KiB

Changelog

[unreleased]

Added

  • Added: requirements.txt pinning pytest==7.4.4 for reproducible test runs.
  • Added: scripts/install.sh now creates .venv/ and installs pytest into it.
  • Harness pre-edit hook: --can-edit now supports project-level checks without --task, file scope checks with --file, and --json output for machine-readable harness integration
  • opencode plugin: plugins/automaton-guard/plugin.ts — intercepts edit and write tool calls, calls --can-edit before allowing modifications
  • Git pre-commit hook: scripts/git-hooks/pre-commit — blocks commits when no task is in an edit-allowed phase (universal safety net for all harnesses)
  • Pre-v2.0 task enforcement: Tasks without .state files are UNTRACKED — --transition, --can-edit, --task, and --approve all refuse to operate on them
  • New --upgrade command: Bootstraps .state files for pre-v2.0 tasks (single task with --task or all tasks at once)
  • Untracked task reporting: --list shows UNTRACKED (no .state) for tasks without .state files instead of silently bootstrapping
  • Project scoping fix: status.py errors when no project is detected instead of silently falling back to framework directory
  • Scope check fix: --scope-check marks framework files as OUT_OF_SCOPE when working on a project
  • Dashboard scope fix: Handler methods use stored project_root instead of re-detecting from CWD on every request
  • --project flag: Added to all status.py command invocations across 16+ prompt and config files
  • _infer_state_from_artifacts locked to --upgrade: Removed as silent fallback from all operational commands
  • Phase approval gates: Research, Decomposition, Design, and Test Design phases now require explicit user approval (:awaiting_approval → :approved) before proceeding
  • status.py script: Comprehensive enforcement and status tool with --task, --list, --create-task, --transition, --approve, --validate-folder, --audit, --claim, --release, --next-available, --available, --can-edit, --scope-check, --same-session, --upgrade
  • Untracked task enforcement: Tasks without .state files are UNTRACKED — --transition, --can-edit, --task, --approve all refuse to operate on them. Run --upgrade to bootstrap .state files
  • --project flag: All status.py commands now support --project for explicit project scoping when multiple projects exist on the same machine
  • --upgrade command: Bootstraps .state files for pre-v2.0 tasks that lack them (single task with --task or all tasks at once)
  • Project scoping: status.py now errors when not in a project directory and --project is not specified, instead of silently falling back to ~/.automaton/
  • Scope check fix: --scope-check now correctly marks framework files as OUT_OF_SCOPE when working on a project (was incorrectly always IN_SCOPE)
  • Dashboard scope fix: Dashboard handler methods now use stored project_root and scope instead of re-detecting from CWD on every request
  • Phase-scoped prompts: All phase prompts now include ALLOWED ACTIONS, FORBIDDEN ACTIONS, approval gates (where applicable), pre-work validation, and .state precondition checks
  • Orchestrator restructuring: Reduced from 493 lines to 143 lines; sub-task management extracted to subtask_management.md; state machine reference moved to workflow.md
  • ALLOWED/FORBIDDEN enforcement: Each phase prompt explicitly defines what agents can and cannot do, with user override resistance instructions
  • Workflow enforcement: --transition refuses illegal phase transitions; --validate-folder detects out-of-order artifacts; --audit checks all tasks for violations
  • Task creation gate: status.py --create-task is the only valid way to create tasks; --audit flags manually created folders
  • Approval log: .state.approvals file records all user approvals with timestamp and approver
  • Multi-agent support: Optional Agent Configuration section in .agent.md enables task claiming, role binding, and work discovery for multi-agent setups
  • Tool integration hooks: --can-edit, --scope-check, --same-session for agent tool integrations (optional, not called by prompts)
  • upgrade.sh script: Bootstraps .state files for existing tasks from artifact heuristic
  • Framework version marker: config.md now includes version 2.0 with state enforcement indicator

Changed

  • Changed: All documented python invocations now read python3 (stock macOS / Windows Python ship as python3).
  • orchestrate.md: Reduced from 493 to 143 lines; gate-check loop replaces soft advisory approach; approval gates enforced at research, decomposition, design, and test_design
  • workflow.md: Rewritten to reference .state as canonical phase indicator; approval sub-states documented; enforcement via status.py documented
  • All phase prompts: Added .state precondition check, pre-work validation, ALLOWED/FORBIDDEN sections, handling user overrides
  • research.md, design.md, decompose.md, test_design.md: Added approval gate sections with --transition {phase}:awaiting_approval and --approve
  • implement.md, bug_finder.md, adversarial_bug_find.md, doc_review.md, referee.md: Added no-approval-gate notes with direct --transition instructions
  • status_reason property on Task model showing human-readable explanation for each state (#task-status-reason)
  • Revoke buttons for approved/changes_requested reviews — replaces approve/request-changes with a single revoke option (#task-status-reason)
  • pytest test suite covering dashboard core, app security, and VRAM detection (#add-pytest-test-suite)
  • Structured verdict parsing: parse_verdict_status() uses ## Status: line before substring fallback, preventing false-BLOCKED classification (#fix-verdict-parsing)
  • State machine alignment: IMPLEMENTATION.md alone → Bug Find, ADVERSARIAL_BUG_REPORT alone → Bug Find (matching orchestrator spec) (#fix-verdict-parsing)
  • Filesystem task name validation: discover_tasks() and parse_sub_tasks() skip directories with invalid characters (#fix-verdict-parsing)
  • Added CORS headers, do_OPTIONS handler, X-Content-Type-Options to all dashboard API responses (#harden-dashboard-security)
  • Added POST content-length bounds (64KB) and review comment length limits (4096 chars) (#harden-dashboard-security)
  • Replaced inline onclick review handlers with data-* attributes and event delegation (#harden-dashboard-security)
  • Applied escapeHtml() to task display_name in dashboard card rendering (#harden-dashboard-security)
  • GET /api/config and PUT /api/config endpoints for reading and persisting dashboard configuration (#wire-dashboard-config)
  • Server-side task cache with 1s TTL to eliminate redundant disk I/O on every polling request (#wire-dashboard-config)
  • Dashboard JS applies config on init: theme, default_view, auto_refresh_interval, column_width, show_timelines (#wire-dashboard-config)
  • Review POSTinvalidates task cache so next poll picks up changes (#wire-dashboard-config)
  • decomposition_content, parent_spec_content, vram_config_content fields on Task model (#add-decomposition-content)
  • WaveGroup dataclass and parse_waves() for extracting wave structure from DECOMPOSITION.md (#add-decomposition-content)
  • parse_vram_config() for reading VRAM_CONFIG.md (#add-decomposition-content)
  • Dashboard JS wave statistics use parsed wave data instead of 50/50 heuristic (#add-decomposition-content)
  • Detail panel shows Decomposition, Parent Context, and VRAM Configuration sections (#add-decomposition-content)

Changed

  • Removed stale dashboard = ["inotify>=0.2"] optional dependency from pyproject.toml (#cleanup-cruft)
  • Deleted debug_root.py stray development script (#cleanup-cruft)
  • Deleted empty automaton/dashboard/ui/widgets/ directory (#cleanup-cruft)
  • Fixed config.md RAM detection description (was "via free", now "via /proc/meminfo or sysctl") (#cleanup-cruft)
  • _find_tasks_dir() returns Path instead of Path | None, removed tautological condition (#cleanup-cruft)
  • Removed sys.path.insert hack from __main__.py (#cleanup-cruft)
  • Documented scripts/dashboard.sh convenience wrapper in README.md (#cleanup-cruft)
  • Framework self-consistency test suite: 17 tests covering prompt stop conditions, hardcoded URLs, canonical paths, .rules.md sections, stale dependencies, CSS theme parity, verdict regression, and CI validation (#framework-self-consistency-tests)

Fixed

  • REFEREE state was never produced by state machine — verdict with unparseable status now correctly shows as REFEREE instead of silently falling through to earlier states (#task-status-reason)
  • Pending review count in header now excludes done/blocked tasks (#task-status-reason)
  • Critical: PASS verdicts mentioning FAIL/NEEDS_REVIEW in body text were falsely classified as BLOCKED (#fix-verdict-parsing)
  • State divergence: IMPLEMENTATION.md alone showed "Implement" instead of "Bug Find" (#fix-verdict-parsing)
  • Added mandatory stop conditions to bug_finder.md and adversarial_bug_find.md (#fix-prompt-consistency)
  • Fixed deprecated {project}/tasks/ path in onboarding.md (#fix-prompt-consistency)
  • Expanded prompt path test to catch concrete deprecated path patterns (#fix-prompt-consistency)
  • Root pyproject.toml with optional test/dashboard dependency groups (#add-pytest-test-suite)
  • AGENTS.md with build/test commands and conventions (#developer-experience-gitea-ci)
  • .gitea/workflows/ci.yml running py_compile, pytest, and shell script syntax checks (#developer-experience-gitea-ci)
  • templates/README.md documenting the task template examples (#developer-experience-gitea-ci)
  • Blocked phase column between Verification and Resolution on dashboard (#additive-extension-model)
  • Framework self-enforcement rules in .rules.md and system-prompt.md (#framework-self-enforcement)
  • Additive extension model: projects extend via extensions/ dir, never copy framework files (#additive-extension-model)
  • CHANGELOG.md for release notes tracking (#changelog)
  • Framework audit: comprehensive self-consistency check with RESEARCH.md (#framework-audit)

Changed

  • All prompts now use the canonical task path {project}/.automaton/tasks/{task-name}/ (#standardize-task-path-conventions)
  • scripts/vram_detect.sh rewritten as scripts/vram_detect.py for testability and correctness (#rewrite-vram-detection-python)
  • tasks/dashboard-spec.md reconciled with the implemented web dashboard (#reconcile-dashboard-spec)
  • automaton/dashboard/README.md and help modal shortcuts now match the web UI (#reconcile-dashboard-spec)
  • prompts/orchestrate.md: always reads prompts/contracts/scripts from global, project extensions are additive (#additive-extension-model)
  • prompts/onboarding.md: removed diff/merge upgrade, replaced with migration check (#additive-extension-model)
  • README.md: updated upgrade docs for new additive model (#additive-extension-model); added Dashboard section (#dashboard-task-review)
  • scripts/update.sh: simplified to plain git pull (#additive-extension-model)
  • .rules.md: converted from template to concrete rules with Task-Driven Development, VRAM-aware sizing, Changelog, and Self-Improvement sections (#framework-self-enforcement)
  • system-prompt.md: added instruction to read global .rules.md (#framework-self-enforcement)
  • automaton/dashboard/ui/app.py: added review API endpoints (GET/POST /api/task/{name}/review), spec_content in responses, unquote() for URL-encoded task names, path traversal fix (#dashboard-task-review, #spec-in-detail)
  • automaton/dashboard/html/dashboard.js: review UI (badges, buttons, filter), artifact badges, specification display, modal conversion, textarea replacement, display group for approved planning tasks (#dashboard-task-review, #artifact-badges, #spec-in-detail, #task-detail-modal, #review-textarea)
  • automaton/dashboard/html/styles.css: review components, artifact badges, modal layout, textarea styles (#dashboard-task-review, #artifact-badges, #task-detail-modal, #review-textarea)
  • automaton/dashboard/html/index.html: review filter, pending count, modal overlay (#dashboard-task-review, #task-detail-modal)
  • automaton/dashboard/core/task.py: fixed state machine priority — IMPLEMENTATION.md now correctly detected, DOC_REVIEW checked before BUG_REPORT (#implement-task)
  • automaton/dashboard/core/board.py: fixed KanbanBoard — added missing COLUMNS and init (#implement-task)
  • automaton/dashboard/core/refresh.py: improved inotify error handling with explicit fallback messages (#implement-task)

Fixed

  • VRAM detection: undefined headroom, hardcoded JSON headroom, and code-block config parsing (#rewrite-vram-detection-python)
  • VRAM detection: 10KB file-read limit now enforced for API config files (#rewrite-vram-detection-python)
  • Dashboard static file serving: replaced string-prefix path traversal check with Path.relative_to() (#harden-dashboard-security-scripts)
  • Dashboard task name validation: restricted to [A-Za-z0-9_-]+ (#harden-dashboard-security-scripts)
  • scripts/update.sh: now warns and aborts on uncommitted changes before pulling (#harden-dashboard-security-scripts)
  • README/install.sh: replaced placeholder repository URL with real Gitea URL (#harden-dashboard-security-scripts)
  • State machine: IMPLEMENTATION.md was never checked in determine_task_state(), tasks showed as RESEARCH (#implement-task)
  • State machine: DOC_REVIEW checked after BUG_REPORT — wrong priority order (#implement-task)
  • Path traversal: review API accepted task names with ../ allowing writes outside tasks directory (#dashboard-task-review)
  • URL encoding: task names with spaces in API paths were not decoded (#implement-task)
  • Review parsing: comment extraction used fragile conditional, falsy comments (e.g., "0") skipped (#implement-task)
  • Board display: approved planning tasks stayed in Planning column instead of advancing to Design (#dashboard-task-review)

Removed

  • automaton/dashboard/themes.py (vestigial ANSI theme stub) (#reconcile-dashboard-spec)
  • automaton/dashboard/core/refresh.py (half-implemented file watcher; dashboard uses JS polling) (#remove-file-system-watcher)
  • templates/contract-template.md (unused) (#developer-experience-gitea-ci)
  • automaton/dashboard/pyproject.toml (consolidated into root pyproject.toml) (#add-pytest-test-suite)

Migration

  • Project migration script for old-model projects: scripts/migrate-project.sh (#project-migration)
  • Project migration detection in onboarding.md (#project-migration)