Commit Graph
24 Commits
Author SHA1 Message Date
Lap Tran c2355954b9 feat(pi): pi-automaton.sh startup context printer for Pi Dev
Pi Dev doesn't auto-load automaton's system prompt (unlike opencode),
so the agent has no awareness of tasks, phases, or status.py commands.
This bridges that gap:

- scripts/pi-automaton.sh: NEW — detects scope (framework/project),
  reads active tasks, default model, config snippet, AGENTS.md rules,
  and prints a formatted context block the user pastes as their first
  message to the Pi Dev agent. Does NOT launch pi.
- scripts/register-guards.sh: EXTENDED — after Pi Dev guard install,
  offers to symlink pi-automaton.sh to ~/bin/pi-automaton
  (interactive prompt only when stdin is a terminal).
- README.md: added 'Using Pi Dev with automaton' FAQ subsection
  documenting the context printer workflow.
2026-06-26 20:47:08 -04:00
Lap Tran 0437cbae6c docs: architecture section, FAQ, cross-references, vault-memory update
README.md:
  - Added §1.5 'Architecture: Framework vs Project' with directory tree
    and lifecycle flow diagram
  - Added §2.5 FAQ covering coexistence, hooks, loops, multi-project
  - Updated install section already done in prior commit

AGENTS.md:
  - Added 'Script Cross-References' table mapping install.sh →
    onboard-project.sh → status.py --create-task

scripts/install-hooks.sh:
  - Updated header to reference onboard-project.sh as caller
  - Added 'Next step' line pointing to --create-task

vault-memory CONTEXT.md:
  - Updated test count (518→611)
  - Added new scripts (detect_models.py, onboard-project.sh)
  - Documented install/onboard flow and split architecture
2026-06-26 16:51:39 -04:00
Lap Tran b880f2535a fix(install): make install.sh idempotent + add onboard-project.sh
- install.sh: no longer exits early when ~/.automaton exists.
  Skips the clone but runs all setup (VRAM detection, guards,
  self-improvement loop, virtualenv). Both curl|bash and
  git-clone + ./install.sh now work correctly.
- onboard-project.sh: new script that bootstraps automaton in a
  project — creates .automaton/ skeleton, detects models, writes
  config.md, inits git, installs hooks, adds .gitignore entries.
- README.md: fix install flow docs (curl|bash + clone-then-run),
  add onboard-project.sh as Option A for project setup
2026-06-26 13:54:28 -04:00
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00
Lap Tran e13513faaa Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
2026-06-24 10:31:49 -04:00
Lap Tran f32f98575b Make test suite runnable from clean checkout + cross-platform vram_detect
CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00
gitea af66f5081d Fix pre-existing issues and update README for v2.0
CI / build (push) Has been cancelled
- Fix _infer_state_from_artifacts: SPEC-only maps to research (was bug_find)
- Fix cmd_validate_folder: corrupted .state files now error instead of silent fallback
- Fix upgrade.sh: remove || true, add python3 check, use git rev-parse --git-dir
- Update README.md with --project flag, --can-edit modes, --upgrade, enforcement layers
- Add 6 new tests for inference and validation fixes
- Complete hook-install-process and readme-upgrade-docs tasks
2026-06-15 14:56:33 -04:00
gitea 05c76852a2 v2.0: state enforcement, project scoping, harness integration
CI / build (push) Has been cancelled
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks

Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)

Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)

Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2026-06-15 14:16:46 -04:00
gitea 79b783864e Harden framework: tests, VRAM Python, dashboard spec, security, CI
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit

- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM

- Standardize all prompts to .automaton/tasks/{task-name}/ path

- Reconcile dashboard spec with web implementation; remove themes.py

- Remove half-implemented refresh.py file watcher

- Harden dashboard static-file serving and task-name validation

- Add uncommitted-change guard to update.sh and real Gitea URLs

- Add AGENTS.md, Gitea CI workflow, and template documentation
2026-06-14 11:24:36 -04:00
gitea b15b495d2e Close all 10 tasks through full lifecycle (Implementation → Bug Find → Adversarial → Doc Review → Referee)
- Drive all approved tasks to completion with VERDICT.md
- Fix state machine: IMPLEMENTATION.md was never checked in determine_task_state()
- Fix state machine: DOC_REVIEW.md priority wrong (checked after BUG_REPORT)
- Fix board display: approved planning tasks now advance to Design group
- Fix board display: rejected planning tasks move to Blocked group
- Fix path traversal: review API validated task names against ../ injection
- Fix URL encoding: unquote() task names in API path parsing
- Fix comment parsing: robust REVIEW.md read/write, handle falsy comments
- Fix dead code: KanbanBoard class missing COLUMNS and __init__
- Fix inotify: explicit error messages and polling fallback
- Fix review API: validate task names, prevent path traversal
- Update CHANGELOG.md with all changes
2026-06-13 21:09:57 -04:00
gitea 52dcf8e309 Implement all 5 tasks: additive extension model, framework self-enforcement, changelog, project migration, dashboard review 2026-06-13 12:26:27 -04:00
gitea a1391e7364 Rename product from agent-framework to automaton; move scripts to scripts/ folder 2026-06-12 13:22:10 -04:00
gitea 502f47eb21 Refactor: rename framework files to dot-prefixed lowercase, fix onboarding references, validate VRAM detection
- Rename AGENT.md -> .agent.md, RULES.md -> .rules.md, ONBOARDING.md -> .onboarding.md
- Rename BUG_REPORT.md -> .bug_report.md, ADVERSARIAL_BUG_REPORT.md -> .adversarial_bug_report.md, VERDICT.md -> .verdict.md
- Fix onboarding.md references to use new .onboarding.md path
- Fix stop-hook-pattern.md reference to use .onboarding.md
- Update README.md, config.md, install.sh, update.sh, prompts/*, references/*
- VRAM detection script validated and working
2026-06-12 12:40:15 -04:00
gitea c629661b28 Redesign file system to avoid conflicts: layered approach with project overrides and global defaults, proper upgrade process 2026-06-11 22:25:09 -04:00
gitea a74eadfb86 Split config: VRAM/model settings to config.md, agent behavior to AGENT.md 2026-06-11 20:25:27 -04:00
gitea f4587886b9 Add Test Design phase: test_design.md prompt, workflow state machine, orchestrator updates, implement/referee reference TEST_PLAN.md 2026-06-11 09:25:45 -04:00
gitea fec11d29dc Enable Autopilot by default: AGENT.md, README, onboarding, orchestrate updates 2026-06-10 23:59:37 -04:00
gitea a17cbbe304 Add update mechanism: update.sh script, project upgrade option in onboarding, update instructions in README, Bug 12 2026-06-10 23:48:43 -04:00
gitea 3ac1b0858b Fix bugs 1-11: interactive research/design, doc review phase, orchestrator state machine, design phase in workflow, fix verdict task creation contradiction 2026-06-10 23:44:55 -04:00
gitea 3cdb95e083 Integrate TDD and Grill with Docs into Design, Implementation, and Verification workflows 2026-06-10 14:46:11 -04:00
gitea 59b6339765 Implement Autopilot mode: state machine, adversarial bug finding, and proactive orchestration 2026-06-09 23:58:48 -04:00
gitea d235c12ab0 Add design phase + 3D visualization + orchestration improvements 2026-05-31 00:29:10 -04:00
gitea 72ae07c310 Initial commit: minimal agent framework 2026-05-30 23:27:09 -04:00