Files
Lap Tran 81ccf548e5
CI / build (push) Has been cancelled
Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
2026-06-22 10:40:58 -04:00

3.5 KiB
Raw Permalink Blame History

Verdict: Full-Codebase Audit (v2.0)

Status: NEEDS_REVIEW

Completion Date: 2026-06-22

Summary

Full-codebase audit of the automaton framework v2.0. The Bug Finder identified 7 bugs (score: 55); the Adversarial Bug Finder identified 3 additional bugs (score: 7). No overlapping or contradictory findings. The framework's core state machine, phase transitions, and approval gates work correctly. However, several bugs exist in the enforcement code itself (status.py), the guard registration script, and the dashboard — undermining the "state-enforced" guarantee in specific scenarios.

Findings

What Passed

  • Core state machine: .state file reading, _base_phase() approval substate handling, and --transition validation are correct
  • _check_forbidden_artifacts() correctly detects out-of-order artifacts per phase
  • Dashboard static file serving has proper path-traversal protection (resolve() + relative_to())
  • Dashboard review file writing validates task names against [A-Za-z0-9_-]+
  • POST body size limits (64KB) and review comment length limits (4096) are enforced
  • Test suite: 235 tests pass
  • Git hooks (pre-commit, post-commit, pre-push) have correct logic

What Failed

  • Bug 1 (High): Category 3 audit is broken for ALL regular (non-framework) projects — task folder changes are always flagged as unauthorized
  • Bug 4 (High): Verdict inference uses "PASS" in content substring search, which can misclassify FAIL/NEEDS_REVIEW verdicts as complete
  • Bug 5 (High): register-guards.sh never successfully registers the OpenCode guard due to 3 compounding bugs (wrong filename, wrong config key, JSONC parsing)
  • Bug 7 (High): --can-edit file scope check uses string startswith without trailing separator, allowing edits in sibling directories

What Needs Review

  • Bug 2 (Medium): migrate-project.sh find precedence silently skips .md files during migration
  • Bug 3 (Medium): vram_detect.py prefix matching gives wrong context windows for unknown models
  • Bug 6 (Medium): Dashboard ignores .state files, contradicting the "single source of truth" design
  • Bug 8 (Medium): Dashboard CORS * allows cross-origin writes from any website
  • Bug 9 (Low): Stale-task mtime proxy is unreliable after transitions
  • Bug 10 (Low): Artifact inference maps TEST_PLAN.md to implement instead of test_design

Bug Finder vs Adversarial Bug Finder Comparison

  • Found by both: None (reports are complementary by design)
  • Found only by Bug Finder: Bugs 1–7 (enforcement logic, shell scripts, model lookup, guard registration)
  • Found only by Adversarial Bug Finder: Bugs 8–10 (security, session tracking, phase inference edge case)
  • Contradictions: None. The two reports are consistent and non-overlapping.

Tasks for Review / Tie-Breaks

  • None. No contradictions between Bug Finder and Adversarial Bug Finder. All 10 bugs are independently verified with reproduction steps.

Remaining Issues

  • Bugs 1, 4, 5, 7 (High severity) should be fixed before relying on the framework for production enforcement. Bug 5 in particular means the primary enforcement layer (harness pre-edit guard) is likely not registered for most users.
  • Bug 6 means the dashboard display may not match the actual enforced state — users could make decisions based on stale dashboard info.
  • Bug 7 is a security issue: the --can-edit scope check can be bypassed via sibling directory names.

Score

+5 (NEEDS_REVIEW)

Reviewer Comments