Files
Lap Tran 81ccf548e5
CI / build (push) Has been cancelled
Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
2026-06-22 10:40:58 -04:00

53 lines
3.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Verdict: Full-Codebase Audit (v2.0)
## Status: NEEDS_REVIEW
**Completion Date**: 2026-06-22
## Summary
Full-codebase audit of the automaton framework v2.0. The Bug Finder identified 7 bugs (score: 55); the Adversarial Bug Finder identified 3 additional bugs (score: 7). No overlapping or contradictory findings. The framework's core state machine, phase transitions, and approval gates work correctly. However, several bugs exist in the enforcement code itself (`status.py`), the guard registration script, and the dashboard — undermining the "state-enforced" guarantee in specific scenarios.
## Findings
### What Passed
- Core state machine: `.state` file reading, `_base_phase()` approval substate handling, and `--transition` validation are correct
- `_check_forbidden_artifacts()` correctly detects out-of-order artifacts per phase
- Dashboard static file serving has proper path-traversal protection (`resolve()` + `relative_to()`)
- Dashboard review file writing validates task names against `[A-Za-z0-9_-]+`
- POST body size limits (64KB) and review comment length limits (4096) are enforced
- Test suite: 235 tests pass
- Git hooks (pre-commit, post-commit, pre-push) have correct logic
### What Failed
- **Bug 1** (High): Category 3 audit is broken for ALL regular (non-framework) projects — task folder changes are always flagged as unauthorized
- **Bug 4** (High): Verdict inference uses `"PASS" in content` substring search, which can misclassify FAIL/NEEDS_REVIEW verdicts as complete
- **Bug 5** (High): `register-guards.sh` never successfully registers the OpenCode guard due to 3 compounding bugs (wrong filename, wrong config key, JSONC parsing)
- **Bug 7** (High): `--can-edit` file scope check uses string `startswith` without trailing separator, allowing edits in sibling directories
### What Needs Review
- **Bug 2** (Medium): `migrate-project.sh` `find` precedence silently skips `.md` files during migration
- **Bug 3** (Medium): `vram_detect.py` prefix matching gives wrong context windows for unknown models
- **Bug 6** (Medium): Dashboard ignores `.state` files, contradicting the "single source of truth" design
- **Bug 8** (Medium): Dashboard CORS `*` allows cross-origin writes from any website
- **Bug 9** (Low): Stale-task mtime proxy is unreliable after transitions
- **Bug 10** (Low): Artifact inference maps TEST_PLAN.md to `implement` instead of `test_design`
## Bug Finder vs Adversarial Bug Finder Comparison
- **Found by both**: None (reports are complementary by design)
- **Found only by Bug Finder**: Bugs 1–7 (enforcement logic, shell scripts, model lookup, guard registration)
- **Found only by Adversarial Bug Finder**: Bugs 8–10 (security, session tracking, phase inference edge case)
- **Contradictions**: None. The two reports are consistent and non-overlapping.
## Tasks for Review / Tie-Breaks
- None. No contradictions between Bug Finder and Adversarial Bug Finder. All 10 bugs are independently verified with reproduction steps.
## Remaining Issues
- Bugs 1, 4, 5, 7 (High severity) should be fixed before relying on the framework for production enforcement. Bug 5 in particular means the primary enforcement layer (harness pre-edit guard) is likely not registered for most users.
- Bug 6 means the dashboard display may not match the actual enforced state — users could make decisions based on stale dashboard info.
- Bug 7 is a security issue: the `--can-edit` scope check can be bypassed via sibling directory names.
## Score
+5 (NEEDS_REVIEW)
## Reviewer Comments