Add memory/ with audit findings and framework knowledge
CI / build (push) Has been cancelled

7 memory files covering:
- audit-bug-patterns: recurring status.py bug patterns
- vram-model-matching: three-tier model prefix matching
- dashboard-security: .state priority, CORS removal, register-guards
- state-machine-workflow: legal transitions and approval gates
- testing-conventions: pytest patterns and helpers
- audit-process: report conventions and batch processing
- framework-architecture: enforcement layers and key files
This commit is contained in:
Lap Tran
2026-06-22 10:48:19 -04:00
parent 81ccf548e5
commit dd2726c0dd
7 changed files with 107 additions and 0 deletions
+18
View File
@@ -0,0 +1,18 @@
# Audit Bug Patterns in status.py
Recurring issues found during the 10-bug audit (2026-06-22).
## 1. Path prefix matching without separator boundary
`startswith(proj_str)` matches sibling directories (e.g. `/foo/project-evil` matches `/foo/project`). Always use `startswith(proj_str + os.sep)` with exact-match fallback.
## 2. Substring verdict parsing
Using `"PASS" in content` falsely matches PASS mentioned in FAIL body text. Always use structured parsing (e.g. `## Status:` line) before substring fallback.
## 3. mtime as activity proxy
`.state` file mtime reflects phase transitions, not actual edit activity. Use a separate `.state.lastedit` file touched on ALLOWED `--can-edit` responses for stale-task detection.
## 4. Artifact-to-phase mapping mismatches
TEST_PLAN.md is produced during `test_design`, not `implement`. Always cross-reference artifact filenames with the workflow phase that produces them.
## 5. Untracked task enforcement
Tasks without `.state` files must be refused by all operational commands. Run `--upgrade` to bootstrap them.
+19
View File
@@ -0,0 +1,19 @@
# Audit process and report conventions
## Report location
Root-level dotfiles (`.bug_report.md`, `.adversarial_bug_report.md`, `.verdict.md`) following precedent from previous reviews — NOT task artifacts.
## Bug finder workflow
Read all source files in parallel, identify bugs, score by severity (1-10), write `.bug_report.md` with Bug N, Location, Description, Suggested Fix.
## Adversarial bug find
Attack-vector focused — test edge cases, security implications, race conditions, input manipulation. Write `.adversarial_bug_report.md`.
## Batch processing
Group bugs by severity. High-severity (batch 1) first, then Medium/Low (batch 2). Each bug gets its own task with full lifecycle (SPEC → IMPLEMENTATION → CODE_REVIEW → BUG_REPORT → ADVERSARIAL_BUG_REPORT → DOC_REVIEW → VERDICT).
## Verdict format
Uses `## Status: PASS` or `## Status: FAIL` structured line (parsed by `_parse_verdict_status_line()` in status.py and `parse_verdict_status()` in dashboard task.py).
## Pre-approved workflow
User can pre-approve all tasks upfront, allowing the agent to drive all phases to completion without stopping for approval gates.
+10
View File
@@ -0,0 +1,10 @@
# Dashboard state inference and security
## .state file is source of truth
`determine_task_state()` in `automaton/dashboard/core/task.py` must read `.state` file BEFORE falling back to artifact heuristic. Added `_state_string_to_task_state()` to map phase strings (including sub-states like `code_review:awaiting_approval`) to TaskState enum.
## No wildcard CORS on local apps
Dashboard is single-origin (serves HTML + API from same origin). `Access-Control-Allow-Origin: *` allows any malicious webpage to call the API. Remove CORS headers entirely; use `SECURITY_HEADERS` with `X-Content-Type-Options: nosniff` and `X-Frame-Options: DENY` instead.
## register-guards.sh pitfalls
Must check both `.json` and `.jsonc` opencode config files, write to `plugin` (singular) key not `plugins`, and strip `//` comments before `json.loads()`.
+19
View File
@@ -0,0 +1,19 @@
# Framework architecture — key files and enforcement layers
## Three enforcement layers (in order of strength)
1. **Harness pre-edit hook** (`--can-edit`) — blocks edits before they happen. opencode plugin at `plugins/automaton-guard/`.
2. **Git pre-commit hook** (`scripts/git-hooks/pre-commit`) — blocks commits when no task in edit-allowed phase. Works for ALL harnesses.
3. **Prompt-based rules** (ALLOWED/FORBIDDEN sections) — advisory only.
## Key files
- `scripts/status.py` (~1430 lines) — core enforcement: phase transitions, approvals, can-edit, scope-check, audit, claim/release
- `scripts/vram_detect.py` (~700 lines) — VRAM detection and model context lookup
- `automaton/dashboard/core/task.py` — dashboard state inference (`.state` first, artifacts fallback)
- `automaton/dashboard/ui/app.py` — HTTP server with security headers
## NON_ARTIFACT_FILES
Files that should NOT be treated as phase artifacts:
`.state`, `.state.tmp`, `.state.lock`, `.state.approvals`, `.state.implementer`, `.state.lastedit`, `VRAM_CONFIG.md`, `PARENT_SPEC.md`, `REVIEW.md`
## Stale-task mechanism
`.state.lastedit` file touched by `--can-edit` on ALLOWED. Falls back to `.state` mtime if `.state.lastedit` doesn't exist. 30-minute threshold. `--touch` resets the timer.
+20
View File
@@ -0,0 +1,20 @@
# State machine workflow and approval gates
LEGAL_TRANSITIONS in status.py:85-108 defines phase flow:
```
new → research → research:awaiting_approval → research:approved
→ decomposition → decomposition:awaiting_approval → decomposition:approved
→ design → design:awaiting_approval → design:approved
→ test_design → test_design:awaiting_approval → test_design:approved
→ implement → code_review → code_review:awaiting_approval → code_review:approved
→ bug_find → adversarial_bug_find → doc_review → referee → complete
```
## Key points
- Approval gates exist for: research, decomposition, design, test_design, code_review
- To approve: must first transition to `{phase}:awaiting_approval`, then `--approve`, then transition to next phase
- `implement → code_review → bug_find → adversarial_bug_find → doc_review → referee → complete` has NO approval gates (direct transitions)
- All transitions must go through `status.py --transition`
- `.state` file is the single source of truth for current phase
- `--upgrade` bootstraps `.state` for pre-v2.0 tasks using artifact heuristic
+12
View File
@@ -0,0 +1,12 @@
# Testing conventions for automaton framework
- Test suite: `python3 -m pytest tests/ -v` (249 tests as of audit completion)
- Tests call status.py via subprocess: `_run_status(["--flag", "val"], project)` pattern in test_status.py
- `_create_task(project, task_name, phase)` helper creates task via `--create-task` and writes `.state` file
- `tmp_project` fixture creates `.automaton/tasks/` structure in tmp_path
- Dashboard tests in test_app.py use `_make_task()` helper and `determine_task_state()` directly
- When changing phase mappings (e.g. TEST_PLAN.md → test_design), update BOTH status.py tests AND test_task.py tests
- vram_detect tests import module directly: `import vram_detect as vram`
- Framework self-consistency tests in test_prompt_paths.py verify canonical task paths in prompts
- Compile check: `python3 -m py_compile automaton/**/*.py scripts/*.py`
- Shell syntax check: `bash -n scripts/*.sh`
+9
View File
@@ -0,0 +1,9 @@
# vram_detect.py model prefix matching
`_lookup_model_context()` must avoid false prefix matches. The correct approach is three-tier:
1. **Exact match** — `name_lower == key_lower`
2. **Ollama parameter tag** — `name_lower.startswith(key_lower + ":")` (e.g. `deepseek-r1:7b` matches `deepseek-r1`)
3. **Known instruction-tuning suffix** — `name_lower.startswith(key_lower + "-")` only if next segment is in `_KNOWN_MODEL_SUFFIXES = {"instruct", "chat", "it", "fp16", "f16", "bf16"}`
Keys sorted by length descending so most specific match wins first. This prevents `phi-4` matching `phi-4-mini-instruct` or `gpt-4o` matching `gpt-4o-foo-unknown`.