Commit Graph
48 Commits
Author SHA1 Message Date
Lap Tran dd2726c0dd Add memory/ with audit findings and framework knowledge
CI / build (push) Has been cancelled
7 memory files covering:
- audit-bug-patterns: recurring status.py bug patterns
- vram-model-matching: three-tier model prefix matching
- dashboard-security: .state priority, CORS removal, register-guards
- state-machine-workflow: legal transitions and approval gates
- testing-conventions: pytest patterns and helpers
- audit-process: report conventions and batch processing
- framework-architecture: enforcement layers and key files
2026-06-22 10:48:19 -04:00
Lap Tran 81ccf548e5 Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
CI / build (push) Has been cancelled
Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
2026-06-22 10:40:58 -04:00
Lap Tran f32f98575b Make test suite runnable from clean checkout + cross-platform vram_detect
CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00
gitea 1d36c0e4ad Add autopilot-aware pre-commit warnings + post-commit driver reminder
CI / build (push) Has been cancelled
- Pre-commit hook: warns when autopilot is enabled and non-terminal tasks exist
- Post-commit hook: after commit, prints non-terminal task summary if autopilot on
- Post-commit exits 0 always (informational only, never blocks)
- Both hooks read .agent.md to detect Autopilot: Enabled
2026-06-16 12:44:21 -04:00
gitea 02c7a4649b Add code_review to dashboard frontend (JS, HTML, CSS)
CI / build (push) Has been cancelled
2026-06-16 12:28:25 -04:00
gitea 3d4c0926b4 Add code_review phase with approval gate, reviewer≠implementer enforcement, and structured CODE_REVIEW.md
CI / build (push) Has been cancelled
- Insert code_review phase between implement and bug_find
- Approval gate: code_review:awaiting_approval → code_review:approved
- Read-only phase — no edits, no fixes, no returning to implement
- Reviewer≠implementer: .state.implementer tracking + --claim enforcement
- Structured CODE_REVIEW.md: spec compliance, design conformance, quality
  scorecard, items found (severity/category/location/resolution), test coverage
- Updated status.py (10 data structures), dashboard (4 files), prompts (3 files),
  agent routing, tests (6 new test classes, 19 new tests)
2026-06-16 09:00:22 -04:00
gitea 42ccc7e2b7 Fix verdict parser to handle old VERDICT.md formats
CI / build (push) Has been cancelled
- Support ## Verdict: PASS, # VERDICT: PASS, VERDICT: PASS formats
- Add ## Status: PASS line to 16 old-format VERDICT.md files
- All 49 tasks now correctly detected as DONE by dashboard
2026-06-16 07:31:18 -04:00
gitea 99ddb98861 Drive all 4 remaining tasks to completion through full lifecycle
CI / build (push) Has been cancelled
- actionable-phase-guidance: lifecycle artifacts + .state->complete
- harden-enforcement-layers: pre-push hook, install-hooks.sh, register-guards.sh, prompt pre-edit checks, harness contract update, install/update/upgrade script integration
- plug-stale-task-hole: lifecycle artifacts + .state->complete
- port-pi-guard: pi dev guard plugin, package.json, register-guards integration

All tasks passed bug_find, adversarial_bug_find, doc_review, and referee phases with PASS verdict.
2026-06-16 07:28:45 -04:00
gitea 21f16b7da2 Fix verdict parsing + plug stale-task enforcement hole
CI / build (push) Has been cancelled
- task.py: remove naive substring fallback from parse_verdict_status(),
  only parse ## Status: header; no header → ambiguous (REFEREE)
- status.py: add _auto_update_verdict_on_complete() — when transitioning
  human_intervention→complete, auto-update VERDICT.md to PASS
- status.py: add stale-task detection to --can-edit — deny edits if
  all edit tasks have .state mtime >30 min old (reason: stale_task)
- status.py: add --touch command to reset task activity clock
- guard plugin: handle stale_task reason with specific error message
- Update test to match new verdict parsing behavior
2026-06-15 21:55:17 -04:00
gitea cc412a51bb Add actionable next-step guidance to all dashboard phases
CI / build (push) Has been cancelled
- task.py: add phase_guidance, blocker, next_phase_name, required_artifact_name,
  is_edit_phase, is_approval_gated properties to Task model
- dashboard.js: show phase_guidance and blocker in detail panel for all phases;
  show status_reason on cards for all phases (was only blocked/bug_find/adv_bug_find);
  add Edit Allowed and Requires Approval badges
- styles.css: style blocker warning, guidance info box, edit/approval badges
- app.py: include new fields in API JSON response
2026-06-15 18:30:46 -04:00
gitea 3480e4ecba Fix 11 automation gaps: dead-end phases, autopilot runtime, guard plugin, status.py bugs, Category 3 audit
CI / build (push) Has been cancelled
- Fix decomposition:approved and human_intervention dead-end phases
- Add scripts/autopilot.py: real drive_all() implementation
- Fix guard plugin: throw Error instead of injecting user messages
- Fix status.py: double continue, _require_state, --list-states
- Add Category 3 (git-based modification) audit
- Add Category 5 (stuck-task detection) audit
- All 206 tests pass
2026-06-15 17:42:17 -04:00
gitea af66f5081d Fix pre-existing issues and update README for v2.0
CI / build (push) Has been cancelled
- Fix _infer_state_from_artifacts: SPEC-only maps to research (was bug_find)
- Fix cmd_validate_folder: corrupted .state files now error instead of silent fallback
- Fix upgrade.sh: remove || true, add python3 check, use git rev-parse --git-dir
- Update README.md with --project flag, --can-edit modes, --upgrade, enforcement layers
- Add 6 new tests for inference and validation fixes
- Complete hook-install-process and readme-upgrade-docs tasks
2026-06-15 14:56:33 -04:00
gitea 06e47c5503 Add pre-commit hook installation to onboarding and upgrade processes
CI / build (push) Has been cancelled
- prompts/onboarding.md: Step 2c installs git pre-commit hook for new projects
- scripts/upgrade.sh: Installs pre-commit hook during upgrade (with symlink detection)
- scripts/install.sh: Mentions hook installation in post-install next steps
2026-06-15 14:20:52 -04:00
gitea 05c76852a2 v2.0: state enforcement, project scoping, harness integration
CI / build (push) Has been cancelled
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks

Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)

Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)

Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2026-06-15 14:16:46 -04:00
gitea 79b783864e Harden framework: tests, VRAM Python, dashboard spec, security, CI
- Rewrite vram_detect in Python with fixed config parsing and 10KB read limit

- Add pytest suite (72 tests) covering dashboard core, app security, and VRAM

- Standardize all prompts to .automaton/tasks/{task-name}/ path

- Reconcile dashboard spec with web implementation; remove themes.py

- Remove half-implemented refresh.py file watcher

- Harden dashboard static-file serving and task-name validation

- Add uncommitted-change guard to update.sh and real Gitea URLs

- Add AGENTS.md, Gitea CI workflow, and template documentation
2026-06-14 11:24:36 -04:00
gitea cc2e97bd43 Add scope confinement rule to prevent cross-project task scanning
The agent scanned tasks in invest-copilot while working on automaton
framework, wasting time on unrelated work. Added a rule with concrete
failure example limiting work to the current project's .automaton/tasks/
only.
2026-06-13 22:25:43 -04:00
gitea a1dcf09d67 Fix autopilot drive-all loop — agent was stopping after one task
The orchestrator only drove ONE task per invocation, then stopped with
ORCHESTRATION_COMPLETE. User had to manually re-trigger for each task.

Changes:
- Added drive_all() outer loop that scans EVERY task and drives them all
- Tasks needing user review are flagged; orchestrator continues to next
- Only stops when ALL tasks are terminal or ALL remaining need user input
- Output format now shows session summary (completed, awaiting review, blocked)
- Session-starter references 'orchestrate' as primary command
- Updated continue-from-existing to use drive_all()
2026-06-13 22:11:44 -04:00
gitea da1799cd5c Redesign dashboard with Material Design 3 aesthetic
Full M3 redesign with Google's color scheme, elevation system,
and component styling.

Color: primary (#8ab4f8 dark / #1a73e8 light), error (#f28b82),
surface tones with proper dark/light contrast
Elevation: M3 shadow system (levels 0-5) with proper opacity per
color scheme
Typography: Roboto / system-ui stack
Components: M3 card radi (8px), column radi (12px), modal radi (16px)
Focus: M3 focus ring (2px with primary-alpha 0.2)
Cards: accent bar, subtle hover elevation lift
2026-06-13 21:57:21 -04:00
gitea 1531021783 Redesign dashboard with Moon Design System aesthetic
Complete visual refresh using Moon Design's Dragon Ball Z themed
color tokens: Goku (bg), Gohan (cards), Piccolo (surfaces),
Beerus (borders), Trunks (accent/cyan), Goten (green),
Bulma (red), ChiChi (yellow).

Changes:
- Color tokens: full Moon palette with proper dark/light variants
- Typography: Inter font, improved hierarchy and spacing
- Cards: left accent bar instead of border, hover lift effect
- Buttons: shadow glow on primary, consistent radius
- Modals: slideUp animation, deeper backdrop blur
- Stats: larger value text, uppercase labels
- Scrollbar: transparent track, Moon border color
- Help modal: bordered table rows, monospace kbd
- Filter inputs: focus glow matching accent
- Layout: 24px gutters, 14px header padding
2026-06-13 21:50:28 -04:00
gitea 03d9ed5a9c Update 'upgrade automaton for this project' flow to handle task migration
The onboarding Migration Check now also detects tasks at the deprecated
{project}/tasks/ location and moves them to .automaton/tasks/.
2026-06-13 21:41:42 -04:00
gitea d49dc88e61 Integrate migration detection into update.sh
update.sh now checks the current project after pulling latest changes.
If it detects stale prompt/contract/script copies or tasks at the
deprecated root-level location, it prompts the user to run migration.
No more needing to know about migrate-project.sh separately.
2026-06-13 21:39:29 -04:00
gitea 7c8c606d73 Consolidate task location: always .automaton/tasks/
Both framework and project tasks now live in .automaton/tasks/.
Every project has a .automaton/ directory, so no need for special
scope-based path logic. Removed the dual-convention gap.

Changes:
- Dashboard _find_tasks_dir() and _get_review_path() always use
  {root}/.automaton/tasks/ — no scope branching
- migrate-project.sh: moves tasks/ -> .automaton/tasks/ during upgrade
- Orchestrator prompt: all task path references updated to
  {project}/.automaton/tasks/
2026-06-13 21:37:13 -04:00
gitea 741364654c Fix task directory gap in upgrade process
The additive extension model refactor moved prompts/contracts/scripts into
.automaton/ but never addressed where tasks live. Two conventions existed:
framework keeps tasks in .automaton/tasks/, Orchestrator creates project
tasks at tasks/.

Fixes:
- Dashboard _find_tasks_dir() now scope-aware: framework mode prefers
  .automaton/tasks/, project mode prefers tasks/ at project root
- migrate-project.sh: if tasks exist in .automaton/tasks/ (old model),
  move them to tasks/ (project root)
- Orchestrator prompt: documents that tasks location depends on scope
2026-06-13 21:33:23 -04:00
gitea e3e45a23f2 Fix dashboard task directory resolution for projects
Dashboard was hardcoded to read tasks from .automaton/tasks/, but the
Orchestrator creates tasks at tasks/ (project root). Added _find_tasks_dir()
that checks both locations and _get_review_path() that reads REVIEW.md
from whichever location the task lives in.

Priority: .automaton/tasks/ first, then tasks/ as fallback.
This handles framework mode (tasks in .automaton/), project mode
(tasks at root), and symlinked projects.
2026-06-13 21:30:08 -04:00
gitea 2c03abb7ec Fix dashboard module resolution from any project directory
- Add automaton.pth to user site-packages so python -m automaton.dashboard
  works from any working directory, not just ~/.automaton/
- Add scripts/dashboard.sh as a convenience wrapper with PYTHONPATH
2026-06-13 21:18:05 -04:00
gitea b15b495d2e Close all 10 tasks through full lifecycle (Implementation → Bug Find → Adversarial → Doc Review → Referee)
- Drive all approved tasks to completion with VERDICT.md
- Fix state machine: IMPLEMENTATION.md was never checked in determine_task_state()
- Fix state machine: DOC_REVIEW.md priority wrong (checked after BUG_REPORT)
- Fix board display: approved planning tasks now advance to Design group
- Fix board display: rejected planning tasks move to Blocked group
- Fix path traversal: review API validated task names against ../ injection
- Fix URL encoding: unquote() task names in API path parsing
- Fix comment parsing: robust REVIEW.md read/write, handle falsy comments
- Fix dead code: KanbanBoard class missing COLUMNS and __init__
- Fix inotify: explicit error messages and polling fallback
- Fix review API: validate task names, prevent path traversal
- Update CHANGELOG.md with all changes
2026-06-13 21:09:57 -04:00
gitea 9b8f527776 Replace prompt() with inline textarea in review section, close modal on submit 2026-06-13 18:04:10 -04:00
gitea 48d8fb4f44 Add session discipline rule with concrete failure example 2026-06-13 17:59:39 -04:00
gitea 0ac84a8d41 Add retroactive task specs for ad-hoc changes 2026-06-13 17:54:14 -04:00
gitea 6630577bb9 Convert task detail panel to centered modal 2026-06-13 17:52:35 -04:00
gitea 01d8451f3a Display SPEC.md content in task detail panel 2026-06-13 14:51:18 -04:00
gitea 75cb4a14eb Show artifact badges on pending review task cards 2026-06-13 14:38:08 -04:00
gitea 52dcf8e309 Implement all 5 tasks: additive extension model, framework self-enforcement, changelog, project migration, dashboard review 2026-06-13 12:26:27 -04:00
gitea 46ceed0122 Add Blocked dashboard column, framework audit, and 5 new task specs 2026-06-13 12:16:01 -04:00
gitea 5ffcb4b624 Add dashboard, tasks, and template structure 2026-06-12 23:59:21 -04:00
gitea a1391e7364 Rename product from agent-framework to automaton; move scripts to scripts/ folder 2026-06-12 13:22:10 -04:00
gitea 502f47eb21 Refactor: rename framework files to dot-prefixed lowercase, fix onboarding references, validate VRAM detection
- Rename AGENT.md -> .agent.md, RULES.md -> .rules.md, ONBOARDING.md -> .onboarding.md
- Rename BUG_REPORT.md -> .bug_report.md, ADVERSARIAL_BUG_REPORT.md -> .adversarial_bug_report.md, VERDICT.md -> .verdict.md
- Fix onboarding.md references to use new .onboarding.md path
- Fix stop-hook-pattern.md reference to use .onboarding.md
- Update README.md, config.md, install.sh, update.sh, prompts/*, references/*
- VRAM detection script validated and working
2026-06-12 12:40:15 -04:00
gitea c629661b28 Redesign file system to avoid conflicts: layered approach with project overrides and global defaults, proper upgrade process 2026-06-11 22:25:09 -04:00
gitea 8897852cc0 Fix all 30 bugs: Orchestrator state determination, auto-execution loop, sub-task management, VRAM detection, and more 2026-06-11 21:54:15 -04:00
gitea a74eadfb86 Split config: VRAM/model settings to config.md, agent behavior to AGENT.md 2026-06-11 20:25:27 -04:00
gitea f4587886b9 Add Test Design phase: test_design.md prompt, workflow state machine, orchestrator updates, implement/referee reference TEST_PLAN.md 2026-06-11 09:25:45 -04:00
gitea fec11d29dc Enable Autopilot by default: AGENT.md, README, onboarding, orchestrate updates 2026-06-10 23:59:37 -04:00
gitea a17cbbe304 Add update mechanism: update.sh script, project upgrade option in onboarding, update instructions in README, Bug 12 2026-06-10 23:48:43 -04:00
gitea 3ac1b0858b Fix bugs 1-11: interactive research/design, doc review phase, orchestrator state machine, design phase in workflow, fix verdict task creation contradiction 2026-06-10 23:44:55 -04:00
gitea 3cdb95e083 Integrate TDD and Grill with Docs into Design, Implementation, and Verification workflows 2026-06-10 14:46:11 -04:00
gitea 59b6339765 Implement Autopilot mode: state machine, adversarial bug finding, and proactive orchestration 2026-06-09 23:58:48 -04:00
gitea d235c12ab0 Add design phase + 3D visualization + orchestration improvements 2026-05-31 00:29:10 -04:00
gitea 72ae07c310 Initial commit: minimal agent framework 2026-05-30 23:27:09 -04:00