From dd2726c0ddcb7ef453d24201597982ec6812e860 Mon Sep 17 00:00:00 2001 From: Lap Tran Date: Mon, 22 Jun 2026 10:48:19 -0400 Subject: [PATCH] Add memory/ with audit findings and framework knowledge 7 memory files covering: - audit-bug-patterns: recurring status.py bug patterns - vram-model-matching: three-tier model prefix matching - dashboard-security: .state priority, CORS removal, register-guards - state-machine-workflow: legal transitions and approval gates - testing-conventions: pytest patterns and helpers - audit-process: report conventions and batch processing - framework-architecture: enforcement layers and key files --- memory/audit-bug-patterns.md | 18 ++++++++++++++++++ memory/audit-process.md | 19 +++++++++++++++++++ memory/dashboard-security.md | 10 ++++++++++ memory/framework-architecture.md | 19 +++++++++++++++++++ memory/state-machine-workflow.md | 20 ++++++++++++++++++++ memory/testing-conventions.md | 12 ++++++++++++ memory/vram-model-matching.md | 9 +++++++++ 7 files changed, 107 insertions(+) create mode 100644 memory/audit-bug-patterns.md create mode 100644 memory/audit-process.md create mode 100644 memory/dashboard-security.md create mode 100644 memory/framework-architecture.md create mode 100644 memory/state-machine-workflow.md create mode 100644 memory/testing-conventions.md create mode 100644 memory/vram-model-matching.md diff --git a/memory/audit-bug-patterns.md b/memory/audit-bug-patterns.md new file mode 100644 index 0000000..4b8ad5e --- /dev/null +++ b/memory/audit-bug-patterns.md @@ -0,0 +1,18 @@ +# Audit Bug Patterns in status.py + +Recurring issues found during the 10-bug audit (2026-06-22). + +## 1. Path prefix matching without separator boundary +`startswith(proj_str)` matches sibling directories (e.g. `/foo/project-evil` matches `/foo/project`). Always use `startswith(proj_str + os.sep)` with exact-match fallback. + +## 2. Substring verdict parsing +Using `"PASS" in content` falsely matches PASS mentioned in FAIL body text. Always use structured parsing (e.g. `## Status:` line) before substring fallback. + +## 3. mtime as activity proxy +`.state` file mtime reflects phase transitions, not actual edit activity. Use a separate `.state.lastedit` file touched on ALLOWED `--can-edit` responses for stale-task detection. + +## 4. Artifact-to-phase mapping mismatches +TEST_PLAN.md is produced during `test_design`, not `implement`. Always cross-reference artifact filenames with the workflow phase that produces them. + +## 5. Untracked task enforcement +Tasks without `.state` files must be refused by all operational commands. Run `--upgrade` to bootstrap them. diff --git a/memory/audit-process.md b/memory/audit-process.md new file mode 100644 index 0000000..8c4f951 --- /dev/null +++ b/memory/audit-process.md @@ -0,0 +1,19 @@ +# Audit process and report conventions + +## Report location +Root-level dotfiles (`.bug_report.md`, `.adversarial_bug_report.md`, `.verdict.md`) following precedent from previous reviews — NOT task artifacts. + +## Bug finder workflow +Read all source files in parallel, identify bugs, score by severity (1-10), write `.bug_report.md` with Bug N, Location, Description, Suggested Fix. + +## Adversarial bug find +Attack-vector focused — test edge cases, security implications, race conditions, input manipulation. Write `.adversarial_bug_report.md`. + +## Batch processing +Group bugs by severity. High-severity (batch 1) first, then Medium/Low (batch 2). Each bug gets its own task with full lifecycle (SPEC → IMPLEMENTATION → CODE_REVIEW → BUG_REPORT → ADVERSARIAL_BUG_REPORT → DOC_REVIEW → VERDICT). + +## Verdict format +Uses `## Status: PASS` or `## Status: FAIL` structured line (parsed by `_parse_verdict_status_line()` in status.py and `parse_verdict_status()` in dashboard task.py). + +## Pre-approved workflow +User can pre-approve all tasks upfront, allowing the agent to drive all phases to completion without stopping for approval gates. diff --git a/memory/dashboard-security.md b/memory/dashboard-security.md new file mode 100644 index 0000000..4ee8ea9 --- /dev/null +++ b/memory/dashboard-security.md @@ -0,0 +1,10 @@ +# Dashboard state inference and security + +## .state file is source of truth +`determine_task_state()` in `automaton/dashboard/core/task.py` must read `.state` file BEFORE falling back to artifact heuristic. Added `_state_string_to_task_state()` to map phase strings (including sub-states like `code_review:awaiting_approval`) to TaskState enum. + +## No wildcard CORS on local apps +Dashboard is single-origin (serves HTML + API from same origin). `Access-Control-Allow-Origin: *` allows any malicious webpage to call the API. Remove CORS headers entirely; use `SECURITY_HEADERS` with `X-Content-Type-Options: nosniff` and `X-Frame-Options: DENY` instead. + +## register-guards.sh pitfalls +Must check both `.json` and `.jsonc` opencode config files, write to `plugin` (singular) key not `plugins`, and strip `//` comments before `json.loads()`. diff --git a/memory/framework-architecture.md b/memory/framework-architecture.md new file mode 100644 index 0000000..1aec63f --- /dev/null +++ b/memory/framework-architecture.md @@ -0,0 +1,19 @@ +# Framework architecture — key files and enforcement layers + +## Three enforcement layers (in order of strength) +1. **Harness pre-edit hook** (`--can-edit`) — blocks edits before they happen. opencode plugin at `plugins/automaton-guard/`. +2. **Git pre-commit hook** (`scripts/git-hooks/pre-commit`) — blocks commits when no task in edit-allowed phase. Works for ALL harnesses. +3. **Prompt-based rules** (ALLOWED/FORBIDDEN sections) — advisory only. + +## Key files +- `scripts/status.py` (~1430 lines) — core enforcement: phase transitions, approvals, can-edit, scope-check, audit, claim/release +- `scripts/vram_detect.py` (~700 lines) — VRAM detection and model context lookup +- `automaton/dashboard/core/task.py` — dashboard state inference (`.state` first, artifacts fallback) +- `automaton/dashboard/ui/app.py` — HTTP server with security headers + +## NON_ARTIFACT_FILES +Files that should NOT be treated as phase artifacts: +`.state`, `.state.tmp`, `.state.lock`, `.state.approvals`, `.state.implementer`, `.state.lastedit`, `VRAM_CONFIG.md`, `PARENT_SPEC.md`, `REVIEW.md` + +## Stale-task mechanism +`.state.lastedit` file touched by `--can-edit` on ALLOWED. Falls back to `.state` mtime if `.state.lastedit` doesn't exist. 30-minute threshold. `--touch` resets the timer. diff --git a/memory/state-machine-workflow.md b/memory/state-machine-workflow.md new file mode 100644 index 0000000..569a071 --- /dev/null +++ b/memory/state-machine-workflow.md @@ -0,0 +1,20 @@ +# State machine workflow and approval gates + +LEGAL_TRANSITIONS in status.py:85-108 defines phase flow: + +``` +new → research → research:awaiting_approval → research:approved + → decomposition → decomposition:awaiting_approval → decomposition:approved + → design → design:awaiting_approval → design:approved + → test_design → test_design:awaiting_approval → test_design:approved + → implement → code_review → code_review:awaiting_approval → code_review:approved + → bug_find → adversarial_bug_find → doc_review → referee → complete +``` + +## Key points +- Approval gates exist for: research, decomposition, design, test_design, code_review +- To approve: must first transition to `{phase}:awaiting_approval`, then `--approve`, then transition to next phase +- `implement → code_review → bug_find → adversarial_bug_find → doc_review → referee → complete` has NO approval gates (direct transitions) +- All transitions must go through `status.py --transition` +- `.state` file is the single source of truth for current phase +- `--upgrade` bootstraps `.state` for pre-v2.0 tasks using artifact heuristic diff --git a/memory/testing-conventions.md b/memory/testing-conventions.md new file mode 100644 index 0000000..5de278c --- /dev/null +++ b/memory/testing-conventions.md @@ -0,0 +1,12 @@ +# Testing conventions for automaton framework + +- Test suite: `python3 -m pytest tests/ -v` (249 tests as of audit completion) +- Tests call status.py via subprocess: `_run_status(["--flag", "val"], project)` pattern in test_status.py +- `_create_task(project, task_name, phase)` helper creates task via `--create-task` and writes `.state` file +- `tmp_project` fixture creates `.automaton/tasks/` structure in tmp_path +- Dashboard tests in test_app.py use `_make_task()` helper and `determine_task_state()` directly +- When changing phase mappings (e.g. TEST_PLAN.md → test_design), update BOTH status.py tests AND test_task.py tests +- vram_detect tests import module directly: `import vram_detect as vram` +- Framework self-consistency tests in test_prompt_paths.py verify canonical task paths in prompts +- Compile check: `python3 -m py_compile automaton/**/*.py scripts/*.py` +- Shell syntax check: `bash -n scripts/*.sh` diff --git a/memory/vram-model-matching.md b/memory/vram-model-matching.md new file mode 100644 index 0000000..78c3075 --- /dev/null +++ b/memory/vram-model-matching.md @@ -0,0 +1,9 @@ +# vram_detect.py model prefix matching + +`_lookup_model_context()` must avoid false prefix matches. The correct approach is three-tier: + +1. **Exact match** — `name_lower == key_lower` +2. **Ollama parameter tag** — `name_lower.startswith(key_lower + ":")` (e.g. `deepseek-r1:7b` matches `deepseek-r1`) +3. **Known instruction-tuning suffix** — `name_lower.startswith(key_lower + "-")` only if next segment is in `_KNOWN_MODEL_SUFFIXES = {"instruct", "chat", "it", "fp16", "f16", "bf16"}` + +Keys sorted by length descending so most specific match wins first. This prevents `phi-4` matching `phi-4-mini-instruct` or `gpt-4o` matching `gpt-4o-foo-unknown`.