26 Commits
Author SHA1 Message Date
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00
Lap Tran e13513faaa Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
2026-06-24 10:31:49 -04:00
Lap Tran f32f98575b Make test suite runnable from clean checkout + cross-platform vram_detect
CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00
gitea 3d4c0926b4 Add code_review phase with approval gate, reviewer≠implementer enforcement, and structured CODE_REVIEW.md
CI / build (push) Has been cancelled
- Insert code_review phase between implement and bug_find
- Approval gate: code_review:awaiting_approval → code_review:approved
- Read-only phase — no edits, no fixes, no returning to implement
- Reviewer≠implementer: .state.implementer tracking + --claim enforcement
- Structured CODE_REVIEW.md: spec compliance, design conformance, quality
  scorecard, items found (severity/category/location/resolution), test coverage
- Updated status.py (10 data structures), dashboard (4 files), prompts (3 files),
  agent routing, tests (6 new test classes, 19 new tests)
2026-06-16 09:00:22 -04:00
gitea 99ddb98861 Drive all 4 remaining tasks to completion through full lifecycle
CI / build (push) Has been cancelled
- actionable-phase-guidance: lifecycle artifacts + .state->complete
- harden-enforcement-layers: pre-push hook, install-hooks.sh, register-guards.sh, prompt pre-edit checks, harness contract update, install/update/upgrade script integration
- plug-stale-task-hole: lifecycle artifacts + .state->complete
- port-pi-guard: pi dev guard plugin, package.json, register-guards integration

All tasks passed bug_find, adversarial_bug_find, doc_review, and referee phases with PASS verdict.
2026-06-16 07:28:45 -04:00
gitea 06e47c5503 Add pre-commit hook installation to onboarding and upgrade processes
CI / build (push) Has been cancelled
- prompts/onboarding.md: Step 2c installs git pre-commit hook for new projects
- scripts/upgrade.sh: Installs pre-commit hook during upgrade (with symlink detection)
- scripts/install.sh: Mentions hook installation in post-install next steps
2026-06-15 14:20:52 -04:00
gitea 05c76852a2 v2.0: state enforcement, project scoping, harness integration
CI / build (push) Has been cancelled
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks

Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)

Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)

Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2026-06-15 14:16:46 -04:00
gitea 79b783864e Harden framework: tests, VRAM Python, dashboard spec, security, CI
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit

- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM

- Standardize all prompts to .automaton/tasks/{task-name}/ path

- Reconcile dashboard spec with web implementation; remove themes.py

- Remove half-implemented refresh.py file watcher

- Harden dashboard static-file serving and task-name validation

- Add uncommitted-change guard to update.sh and real Gitea URLs

- Add AGENTS.md, Gitea CI workflow, and template documentation
2026-06-14 11:24:36 -04:00
gitea a1dcf09d67 Fix autopilot drive-all loop — agent was stopping after one task
The orchestrator only drove ONE task per invocation, then stopped with
ORCHESTRATION_COMPLETE. User had to manually re-trigger for each task.

Changes:
- Added drive_all() outer loop that scans EVERY task and drives them all
- Tasks needing user review are flagged; orchestrator continues to next
- Only stops when ALL tasks are terminal or ALL remaining need user input
- Output format now shows session summary (completed, awaiting review, blocked)
- Session-starter references 'orchestrate' as primary command
- Updated continue-from-existing to use drive_all()
2026-06-13 22:11:44 -04:00
gitea 03d9ed5a9c Update 'upgrade automaton for this project' flow to handle task migration
The onboarding Migration Check now also detects tasks at the deprecated
{project}/tasks/ location and moves them to .automaton/tasks/.
2026-06-13 21:41:42 -04:00
gitea 7c8c606d73 Consolidate task location: always .automaton/tasks/
Both framework and project tasks now live in .automaton/tasks/.
Every project has a .automaton/ directory, so no need for special
scope-based path logic. Removed the dual-convention gap.

Changes:
- Dashboard _find_tasks_dir() and _get_review_path() always use
  {root}/.automaton/tasks/ — no scope branching
- migrate-project.sh: moves tasks/ -> .automaton/tasks/ during upgrade
- Orchestrator prompt: all task path references updated to
  {project}/.automaton/tasks/
2026-06-13 21:37:13 -04:00
gitea 741364654c Fix task directory gap in upgrade process
The additive extension model refactor moved prompts/contracts/scripts into
.automaton/ but never addressed where tasks live. Two conventions existed:
framework keeps tasks in .automaton/tasks/, Orchestrator creates project
tasks at tasks/.

Fixes:
- Dashboard _find_tasks_dir() now scope-aware: framework mode prefers
  .automaton/tasks/, project mode prefers tasks/ at project root
- migrate-project.sh: if tasks exist in .automaton/tasks/ (old model),
  move them to tasks/ (project root)
- Orchestrator prompt: documents that tasks location depends on scope
2026-06-13 21:33:23 -04:00
gitea 52dcf8e309 Implement all 5 tasks: additive extension model, framework self-enforcement, changelog, project migration, dashboard review 2026-06-13 12:26:27 -04:00
gitea a1391e7364 Rename product from agent-framework to automaton; move scripts to scripts/ folder 2026-06-12 13:22:10 -04:00
gitea 502f47eb21 Refactor: rename framework files to dot-prefixed lowercase, fix onboarding references, validate VRAM detection
- Rename AGENT.md -> .agent.md, RULES.md -> .rules.md, ONBOARDING.md -> .onboarding.md
- Rename BUG_REPORT.md -> .bug_report.md, ADVERSARIAL_BUG_REPORT.md -> .adversarial_bug_report.md, VERDICT.md -> .verdict.md
- Fix onboarding.md references to use new .onboarding.md path
- Fix stop-hook-pattern.md reference to use .onboarding.md
- Update README.md, config.md, install.sh, update.sh, prompts/*, references/*
- VRAM detection script validated and working
2026-06-12 12:40:15 -04:00
gitea c629661b28 Redesign file system to avoid conflicts: layered approach with project overrides and global defaults, proper upgrade process 2026-06-11 22:25:09 -04:00
gitea 8897852cc0 Fix all 30 bugs: Orchestrator state determination, auto-execution loop, sub-task management, VRAM detection, and more 2026-06-11 21:54:15 -04:00
gitea a74eadfb86 Split config: VRAM/model settings to config.md, agent behavior to AGENT.md 2026-06-11 20:25:27 -04:00
gitea f4587886b9 Add Test Design phase: test_design.md prompt, workflow state machine, orchestrator updates, implement/referee reference TEST_PLAN.md 2026-06-11 09:25:45 -04:00
gitea fec11d29dc Enable Autopilot by default: AGENT.md, README, onboarding, orchestrate updates 2026-06-10 23:59:37 -04:00
gitea a17cbbe304 Add update mechanism: update.sh script, project upgrade option in onboarding, update instructions in README, Bug 12 2026-06-10 23:48:43 -04:00
gitea 3ac1b0858b Fix bugs 1-11: interactive research/design, doc review phase, orchestrator state machine, design phase in workflow, fix verdict task creation contradiction 2026-06-10 23:44:55 -04:00
gitea 3cdb95e083 Integrate TDD and Grill with Docs into Design, Implementation, and Verification workflows 2026-06-10 14:46:11 -04:00
gitea 59b6339765 Implement Autopilot mode: state machine, adversarial bug finding, and proactive orchestration 2026-06-09 23:58:48 -04:00
gitea d235c12ab0 Add design phase + 3D visualization + orchestration improvements 2026-05-31 00:29:10 -04:00
gitea 72ae07c310 Initial commit: minimal agent framework 2026-05-30 23:27:09 -04:00