Files
automaton/.bug_report.md
T
Lap Tran 81ccf548e5
CI / build (push) Has been cancelled
Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
2026-06-22 10:40:58 -04:00

8.6 KiB

Bug Report: Full-Codebase Audit (v2.0)

Summary

Full-codebase audit of the automaton framework covering scripts/status.py, scripts/vram_detect.py, scripts/migrate-project.sh, scripts/register-guards.sh, automaton/dashboard/, plugins/, and git hooks. This report supersedes the previous orchestrate-only audit (all prior bugs marked fixed). 7 bugs found across enforcement logic, shell scripts, and the dashboard.

Bugs Found

Bug 1: Category 3 audit false-positives for regular projects

  • Severity: High
  • Location: scripts/status.py:696
  • Description: The Category 3 (unauthorized modifications) audit checks if changed files are inside task folders using parts[0] == "tasks". This only works when the project directory IS ~/.automaton/ (framework mode), where git paths are tasks/mytask/.... For regular projects, task files have git paths like .automaton/tasks/mytask/SPEC.md, where parts[0] is .automaton, not tasks. The check fails, so ALL task-folder changes are flagged as "unauthorized modifications outside task folders."
  • Reproduction: In a regular project (not ~/.automaton/), create a task, transition to implement, make a change inside .automaton/tasks/mytask/, then run python ~/.automaton/scripts/status.py --audit --project /path/to/project. The task file is reported as unauthorized.
  • Suggested Fix: Check both parts[0] == "tasks" (framework mode) and len(parts) >= 3 and parts[0] == ".automaton" and parts[1] == "tasks" and parts[2] in task_names (regular project mode).

Bug 2: find precedence bug silently skips .md files in migrate-project.sh

  • Severity: Medium
  • Location: scripts/migrate-project.sh:108
  • Description: find "$PROJECT_AUTOMATON" -maxdepth 1 -type f -name "*.md" -o -name "*.sh" -print0 lacks parentheses around the -o group. Without parens, find parses this as (-name "*.md") -o (-name "*.sh" -print0). The -print0 only applies to the .sh branch; .md files are found but never printed. This means customized .md files (.agent.md, .rules.md, config.md, etc.) are silently skipped during migration — they are never moved to extensions/ or deleted.
  • Reproduction: Create a temp dir with both .md and .sh files. Run find "$dir" -maxdepth 1 -type f -name "*.md" -o -name "*.sh" -print0 | xargs -0 -I{} basename {}. Only .sh files appear. With parentheses \( -name "*.md" -o -name "*.sh" \) -print0, both appear.
  • Suggested Fix: Add parentheses: find "$PROJECT_AUTOMATON" -maxdepth 1 -type f \( -name "*.md" -o -name "*.sh" \) -print0

Bug 3: Model name startswith matching gives wrong context windows for unknown models

  • Severity: Medium
  • Location: scripts/vram_detect.py:396
  • Description: _lookup_model_context() uses model_name.lower().startswith(key.lower()) to match model names against the MODEL_CONTEXT_WINDOWS dict. This prefix matching causes false matches: a model phi-4-mini matches phi-4 (16000 tokens), gpt-4o-foo-unknown matches gpt-4o (128000), and any unknown model starting with a known prefix gets that prefix's context window instead of the fallback (128000). The comment on line 394 says "Strip common version/date suffixes for lookup" but the code doesn't strip anything — it uses prefix matching. Additionally, iteration order determines which key wins when multiple prefixes match, which may not be the most specific match.
  • Reproduction: python3 -c "import sys; sys.path.insert(0, 'scripts'); from vram_detect import _lookup_model_context; _lookup_model_context('phi-4-mini-instruct')" — prints "Context window: 16k" instead of the fallback 128k.
  • Suggested Fix: Try exact match first, then longest-prefix match (sort keys by length descending), or validate that the model name followed by - or end-of-string matches the key.

Bug 4: VERDICT.md PASS substring inference misclassifies FAIL/NEEDS_REVIEW verdicts

  • Severity: High
  • Location: scripts/status.py:311
  • Description: _infer_state_from_artifacts() checks if "PASS" in content: to determine if a verdict is PASS. This substring search matches "PASS" anywhere in the file — including in body text like "All unit tests PASS" or "NEEDS_REVIEW — but 2 tests PASS." A FAIL or NEEDS_REVIEW verdict containing the word "PASS" in its body is incorrectly classified as complete. The dashboard's parse_verdict_status() (task.py:63) correctly uses structured header-line parsing and explicitly avoids substring search (line 70-72), making this an inconsistency between status.py and the dashboard.
  • Reproduction: Create a VERDICT.md with ## Status: FAIL and body text "Note: 3 tests PASS." Run the audit — the task is classified as complete instead of human_intervention.
  • Suggested Fix: Use the same structured-line parsing as the dashboard's parse_verdict_status(): look for ## Status: header lines and check the value after the colon, not a substring search.

Bug 5: register-guards.sh has three bugs preventing guard registration

  • Severity: High
  • Location: scripts/register-guards.sh:23,34,33
  • Description: Three bugs in the OpenCode guard registration:
    • 5a (line 23): Only checks for opencode.jsonc, not opencode.json. OpenCode supports both .json and .jsonc config files. If the user has opencode.json (the default), the guard is never registered. (Verified: this machine has opencode.json.)
    • 5b (line 34): Writes to the plugins key (cfg.setdefault('plugins', [])), but the OpenCode config uses plugin (singular). The actual config has "plugin": ["opencode-mem"]. Even if registration ran, it would add a plugins key that OpenCode ignores.
    • 5c (line 33): Uses json.load() to parse .jsonc files (JSON with Comments), which fails on files containing // comments. Since the file is .jsonc, comments are expected.
  • Reproduction: Run bash ~/.automaton/scripts/register-guards.sh on a machine with opencode.json (not .jsonc). Output: "OpenCode: not detected (no ~/.config/opencode/opencode.jsonc)".
  • Suggested Fix: Check for both .json and .jsonc. Write to the plugin key (singular). Use a JSONC-aware parser (strip comments before json.load) or use json5 if available.

Bug 6: Dashboard task.py ignores .state files — uses artifact heuristics only

  • Severity: Medium
  • Location: automaton/dashboard/core/task.py:462 (determine_task_state)
  • Description: The dashboard's determine_task_state() infers task phase purely from which artifact files exist, never reading the .state file. This contradicts prompts/workflow.md which declares .state as the "single source of truth." Consequences: (1) A task in implement phase that hasn't written IMPLEMENTATION.md yet shows as an earlier phase. (2) A task in test_design with TEST_PLAN.md shows as IMPLEMENT (line 510-511 maps TEST_PLAN.md → IMPLEMENT). (3) A task in bug_find that already wrote BUG_REPORT.md shows as BUG_FIND even if .state says adversarial_bug_find. The dashboard cannot reflect the actual enforced state.
  • Reproduction: Create a task, transition to implement via status.py, but don't write IMPLEMENTATION.md yet. Open the dashboard — the task shows as an earlier phase (e.g., TEST_DESIGN or DESIGN), not IMPLEMENT.
  • Suggested Fix: Read the .state file first (as status.py does with _read_state()). Fall back to artifact heuristics only if .state doesn't exist (pre-v2.0 tasks).

Bug 7: Path prefix startswith allows scope bypass to sibling directories

  • Severity: High
  • Location: scripts/status.py:951, 1004, 1010, 1047, 1052
  • Description: The --can-edit and --scope-check commands use str(file_path).startswith(proj_str) to verify a file is within the project directory. String startswith matches sibling directories: if proj_str = "/home/user/project", then /home/user/project-evil/file.py matches because it starts with /home/user/project. This allows editing files outside the project boundary if a sibling directory with a similar name exists. The same bug affects the framework directory check (auto_str).
  • Reproduction: python3 -c "print('/Users/laptran/.automaton-evil/file'.startswith('/Users/laptran/.automaton'))" → True. With trailing slash: startswith('/Users/laptran/.automaton/') → False.
  • Suggested Fix: Append a trailing path separator: str(file_path).startswith(proj_str + os.sep) or use Path.relative_to() which correctly resolves path boundaries.

Score

  • Bug 1 (High): +10
  • Bug 2 (Medium): +5
  • Bug 3 (Medium): +5
  • Bug 4 (High): +10
  • Bug 5 (High): +10
  • Bug 6 (Medium): +5
  • Bug 7 (High): +10

Total: 55