25 Commits
Author SHA1 Message Date
Lap Tran 35e449b03e feat(model-divergence): full enforcement — manifest, transition, claim, audit, loop gates, detect script
Completes all 3 model-divergence enforcement subtasks:

- scripts/detect_models.py: probes opencode.json + localhost endpoints,
  builds models.json with --json/--write/--force
- scripts/status.py: CONFLICT_MATRIX, --model flag, --transition --model,
  --claim --model, --audit Category 6, model-divergence brake gate in
  --check-gate, helpers for manifest loading and conflict checking
- scripts/loop-runner.py: _role_model() helper + {model} passed via extras
  dict to _invoke_harness for implement, verify, orchestrate roles
- tests/test_model_divergence.py: 33 tests covering all enforcement layers
- Single-LLM mode: record model advisory, no conflict check
- Multi-LLM mode (2+ models): conflict matrix enforced at transition, claim,
  and loop brake gate
- Project-level models.json preferred over global ~/.automaton/models.json
2026-06-26 13:23:17 -04:00
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00
Lap Tran fe43b9e1fc Remove old review system from dashboard (cosmetic only)
CI / build (push) Has been cancelled
The review system (REVIEW.md, approve/changes_requested buttons,
review filter, review badges) was purely cosmetic — only the dashboard
read/wrote it. No workflow component (status.py, autopilot.py,
loop-runner.py, prompts) ever enforced it.

The 'Approve' button in the task detail panel confused users into
thinking it approved the task's phase gate. In reality it only wrote
to REVIEW.md, which had zero effect on transitions.

Removed:
- Review section (buttons, textarea, status badge) from detail panel
- Review badge from task cards
- Review filter from toolbar
- Pending-review counter from header
- All review-related CSS

Users now use the single '🔓 Approve Phase' button in the detail
panel, which calls status.py --approve and actually transitions the
task.
2026-06-25 11:23:55 -04:00
Lap Tran d325963644 Flatten model-divergence subtasks into 3 independent tasks
CI / build (push) Has been cancelled
Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)

Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.

Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.
2026-06-25 07:25:10 -04:00
Lap Tran 7336db282d Bootstrap self-improvement loop, decompose model-divergence-enforcement
CI / build (push) Has been cancelled
Self-improvement loop:
- Created via --create-loop --from-template self-improvement
- Scheduled via launchd (3600s interval)
- State: running

model-divergence-enforcement task:
- Research approved, decomposed into 3 sequential subtasks:
  1. mde-manifest-detection (models.json + detect_models.py)
  2. mde-interactive-enforcement (conflict matrix + --model args + audit)
  3. mde-loop-enforcement (loop.json roles + {model} substitution + check-gate)
- Parent at decomposition:approved (stays active until subtasks complete)
- Each subtask has BRIEF.md with scope, deliverables, acceptance criteria
2026-06-25 07:15:25 -04:00
Lap Tran 4a2301b077 Archive completed tasks, add cleanup commands, self-documenting dashboard UI
CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00
Lap Tran e13513faaa Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
2026-06-24 10:31:49 -04:00
Lap Tran 81ccf548e5 Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
CI / build (push) Has been cancelled
Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
2026-06-22 10:40:58 -04:00
Lap Tran f32f98575b Make test suite runnable from clean checkout + cross-platform vram_detect
CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00
gitea 1d36c0e4ad Add autopilot-aware pre-commit warnings + post-commit driver reminder
CI / build (push) Has been cancelled
- Pre-commit hook: warns when autopilot is enabled and non-terminal tasks exist
- Post-commit hook: after commit, prints non-terminal task summary if autopilot on
- Post-commit exits 0 always (informational only, never blocks)
- Both hooks read .agent.md to detect Autopilot: Enabled
2026-06-16 12:44:21 -04:00
gitea 42ccc7e2b7 Fix verdict parser to handle old VERDICT.md formats
CI / build (push) Has been cancelled
- Support ## Verdict: PASS, # VERDICT: PASS, VERDICT: PASS formats
- Add ## Status: PASS line to 16 old-format VERDICT.md files
- All 49 tasks now correctly detected as DONE by dashboard
2026-06-16 07:31:18 -04:00
gitea 99ddb98861 Drive all 4 remaining tasks to completion through full lifecycle
CI / build (push) Has been cancelled
- actionable-phase-guidance: lifecycle artifacts + .state->complete
- harden-enforcement-layers: pre-push hook, install-hooks.sh, register-guards.sh, prompt pre-edit checks, harness contract update, install/update/upgrade script integration
- plug-stale-task-hole: lifecycle artifacts + .state->complete
- port-pi-guard: pi dev guard plugin, package.json, register-guards integration

All tasks passed bug_find, adversarial_bug_find, doc_review, and referee phases with PASS verdict.
2026-06-16 07:28:45 -04:00
gitea 21f16b7da2 Fix verdict parsing + plug stale-task enforcement hole
CI / build (push) Has been cancelled
- task.py: remove naive substring fallback from parse_verdict_status(),
  only parse ## Status: header; no header → ambiguous (REFEREE)
- status.py: add _auto_update_verdict_on_complete() — when transitioning
  human_intervention→complete, auto-update VERDICT.md to PASS
- status.py: add stale-task detection to --can-edit — deny edits if
  all edit tasks have .state mtime >30 min old (reason: stale_task)
- status.py: add --touch command to reset task activity clock
- guard plugin: handle stale_task reason with specific error message
- Update test to match new verdict parsing behavior
2026-06-15 21:55:17 -04:00
gitea cc412a51bb Add actionable next-step guidance to all dashboard phases
CI / build (push) Has been cancelled
- task.py: add phase_guidance, blocker, next_phase_name, required_artifact_name,
  is_edit_phase, is_approval_gated properties to Task model
- dashboard.js: show phase_guidance and blocker in detail panel for all phases;
  show status_reason on cards for all phases (was only blocked/bug_find/adv_bug_find);
  add Edit Allowed and Requires Approval badges
- styles.css: style blocker warning, guidance info box, edit/approval badges
- app.py: include new fields in API JSON response
2026-06-15 18:30:46 -04:00
gitea 3480e4ecba Fix 11 automation gaps: dead-end phases, autopilot runtime, guard plugin, status.py bugs, Category 3 audit
CI / build (push) Has been cancelled
- Fix decomposition:approved and human_intervention dead-end phases
- Add scripts/autopilot.py: real drive_all() implementation
- Fix guard plugin: throw Error instead of injecting user messages
- Fix status.py: double continue, _require_state, --list-states
- Add Category 3 (git-based modification) audit
- Add Category 5 (stuck-task detection) audit
- All 206 tests pass
2026-06-15 17:42:17 -04:00
gitea af66f5081d Fix pre-existing issues and update README for v2.0
CI / build (push) Has been cancelled
- Fix _infer_state_from_artifacts: SPEC-only maps to research (was bug_find)
- Fix cmd_validate_folder: corrupted .state files now error instead of silent fallback
- Fix upgrade.sh: remove || true, add python3 check, use git rev-parse --git-dir
- Update README.md with --project flag, --can-edit modes, --upgrade, enforcement layers
- Add 6 new tests for inference and validation fixes
- Complete hook-install-process and readme-upgrade-docs tasks
2026-06-15 14:56:33 -04:00
gitea 06e47c5503 Add pre-commit hook installation to onboarding and upgrade processes
CI / build (push) Has been cancelled
- prompts/onboarding.md: Step 2c installs git pre-commit hook for new projects
- scripts/upgrade.sh: Installs pre-commit hook during upgrade (with symlink detection)
- scripts/install.sh: Mentions hook installation in post-install next steps
2026-06-15 14:20:52 -04:00
gitea 05c76852a2 v2.0: state enforcement, project scoping, harness integration
CI / build (push) Has been cancelled
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks

Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)

Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)

Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2026-06-15 14:16:46 -04:00
gitea 79b783864e Harden framework: tests, VRAM Python, dashboard spec, security, CI
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit

- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM

- Standardize all prompts to .automaton/tasks/{task-name}/ path

- Reconcile dashboard spec with web implementation; remove themes.py

- Remove half-implemented refresh.py file watcher

- Harden dashboard static-file serving and task-name validation

- Add uncommitted-change guard to update.sh and real Gitea URLs

- Add AGENTS.md, Gitea CI workflow, and template documentation
2026-06-14 11:24:36 -04:00
gitea b15b495d2e Close all 10 tasks through full lifecycle (Implementation → Bug Find → Adversarial → Doc Review → Referee)
- Drive all approved tasks to completion with VERDICT.md
- Fix state machine: IMPLEMENTATION.md was never checked in determine_task_state()
- Fix state machine: DOC_REVIEW.md priority wrong (checked after BUG_REPORT)
- Fix board display: approved planning tasks now advance to Design group
- Fix board display: rejected planning tasks move to Blocked group
- Fix path traversal: review API validated task names against ../ injection
- Fix URL encoding: unquote() task names in API path parsing
- Fix comment parsing: robust REVIEW.md read/write, handle falsy comments
- Fix dead code: KanbanBoard class missing COLUMNS and __init__
- Fix inotify: explicit error messages and polling fallback
- Fix review API: validate task names, prevent path traversal
- Update CHANGELOG.md with all changes
2026-06-13 21:09:57 -04:00
gitea 9b8f527776 Replace prompt() with inline textarea in review section, close modal on submit 2026-06-13 18:04:10 -04:00
gitea 0ac84a8d41 Add retroactive task specs for ad-hoc changes 2026-06-13 17:54:14 -04:00
gitea 52dcf8e309 Implement all 5 tasks: additive extension model, framework self-enforcement, changelog, project migration, dashboard review 2026-06-13 12:26:27 -04:00
gitea 46ceed0122 Add Blocked dashboard column, framework audit, and 5 new task specs 2026-06-13 12:16:01 -04:00
gitea 5ffcb4b624 Add dashboard, tasks, and template structure 2026-06-12 23:59:21 -04:00