Commit Graph
10 Commits
Author SHA1 Message Date
Lap Tran 715f6f9495 Fix 7 bugs in autopilot.py, add 43 tests
Fixes from functional correctness review:
- PHASE_PRIORITY: add code_review (was missing, caused wrong task selection)
- cmd_drive: implement→code_review (was illegal implement→bug_find)
- cmd_drive: add handlers for code_review and code_review:approved
- _all_tasks: skip tasks/complete/ (was showing completed tasks as phase=None)
- _all_tasks: recurse into subtasks/ (subtasks were invisible to autopilot)
- _all_tasks: use .automaton/tasks for non-framework projects (was project/tasks)
- is_terminal: only True for complete (human_intervention has legal transitions)
- needs_user_input: flag human_intervention as blocked (was auto-driving it)

Tests cover all 8 bugs: PHASE_PRIORITY ordering, every phase transition,
code_review handlers, complete/ exclusion, subtask recursion, project path,
terminal state, user-input detection, human_intervention handling.
2026-06-25 07:11:44 -04:00
Lap Tran 4a2301b077 Archive completed tasks, add cleanup commands, self-documenting dashboard UI
CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00
Lap Tran e13513faaa Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
2026-06-24 10:31:49 -04:00
Lap Tran 81ccf548e5 Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
CI / build (push) Has been cancelled
Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
2026-06-22 10:40:58 -04:00
Lap Tran f32f98575b Make test suite runnable from clean checkout + cross-platform vram_detect
CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00
gitea 3d4c0926b4 Add code_review phase with approval gate, reviewer≠implementer enforcement, and structured CODE_REVIEW.md
CI / build (push) Has been cancelled
- Insert code_review phase between implement and bug_find
- Approval gate: code_review:awaiting_approval → code_review:approved
- Read-only phase — no edits, no fixes, no returning to implement
- Reviewer≠implementer: .state.implementer tracking + --claim enforcement
- Structured CODE_REVIEW.md: spec compliance, design conformance, quality
  scorecard, items found (severity/category/location/resolution), test coverage
- Updated status.py (10 data structures), dashboard (4 files), prompts (3 files),
  agent routing, tests (6 new test classes, 19 new tests)
2026-06-16 09:00:22 -04:00
gitea 21f16b7da2 Fix verdict parsing + plug stale-task enforcement hole
CI / build (push) Has been cancelled
- task.py: remove naive substring fallback from parse_verdict_status(),
  only parse ## Status: header; no header → ambiguous (REFEREE)
- status.py: add _auto_update_verdict_on_complete() — when transitioning
  human_intervention→complete, auto-update VERDICT.md to PASS
- status.py: add stale-task detection to --can-edit — deny edits if
  all edit tasks have .state mtime >30 min old (reason: stale_task)
- status.py: add --touch command to reset task activity clock
- guard plugin: handle stale_task reason with specific error message
- Update test to match new verdict parsing behavior
2026-06-15 21:55:17 -04:00
gitea af66f5081d Fix pre-existing issues and update README for v2.0
CI / build (push) Has been cancelled
- Fix _infer_state_from_artifacts: SPEC-only maps to research (was bug_find)
- Fix cmd_validate_folder: corrupted .state files now error instead of silent fallback
- Fix upgrade.sh: remove || true, add python3 check, use git rev-parse --git-dir
- Update README.md with --project flag, --can-edit modes, --upgrade, enforcement layers
- Add 6 new tests for inference and validation fixes
- Complete hook-install-process and readme-upgrade-docs tasks
2026-06-15 14:56:33 -04:00
gitea 05c76852a2 v2.0: state enforcement, project scoping, harness integration
CI / build (push) Has been cancelled
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks

Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)

Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)

Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2026-06-15 14:16:46 -04:00
gitea 79b783864e Harden framework: tests, VRAM Python, dashboard spec, security, CI
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit

- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM

- Standardize all prompts to .automaton/tasks/{task-name}/ path

- Reconcile dashboard spec with web implementation; remove themes.py

- Remove half-implemented refresh.py file watcher

- Harden dashboard static-file serving and task-name validation

- Add uncommitted-change guard to update.sh and real Gitea URLs

- Add AGENTS.md, Gitea CI workflow, and template documentation
2026-06-14 11:24:36 -04:00