CI / build (push) Has been cancelled
Batch 1 (High severity): - Bug 1: --audit cat3 now checks .automaton/tasks/ paths - Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing - Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments - Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary Batch 2 (Medium/Low severity): - Bug 2: migrate-project.sh find command parentheses for -prune binding - Bug 3: vram_detect model prefix matching with known-suffix whitelist - Bug 6: dashboard reads .state file before artifact heuristic fallback - Bug 8: removed wildcard CORS, added security headers (nosniff, DENY) - Bug 9: stale-task detection uses .state.lastedit instead of .state mtime - Bug 10: TEST_PLAN.md maps to test_design (was implement) 249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
53 lines
3.5 KiB
Markdown
53 lines
3.5 KiB
Markdown
# Verdict: Full-Codebase Audit (v2.0)
|
||
|
||
## Status: NEEDS_REVIEW
|
||
**Completion Date**: 2026-06-22
|
||
|
||
## Summary
|
||
Full-codebase audit of the automaton framework v2.0. The Bug Finder identified 7 bugs (score: 55); the Adversarial Bug Finder identified 3 additional bugs (score: 7). No overlapping or contradictory findings. The framework's core state machine, phase transitions, and approval gates work correctly. However, several bugs exist in the enforcement code itself (`status.py`), the guard registration script, and the dashboard — undermining the "state-enforced" guarantee in specific scenarios.
|
||
|
||
## Findings
|
||
|
||
### What Passed
|
||
- Core state machine: `.state` file reading, `_base_phase()` approval substate handling, and `--transition` validation are correct
|
||
- `_check_forbidden_artifacts()` correctly detects out-of-order artifacts per phase
|
||
- Dashboard static file serving has proper path-traversal protection (`resolve()` + `relative_to()`)
|
||
- Dashboard review file writing validates task names against `[A-Za-z0-9_-]+`
|
||
- POST body size limits (64KB) and review comment length limits (4096) are enforced
|
||
- Test suite: 235 tests pass
|
||
- Git hooks (pre-commit, post-commit, pre-push) have correct logic
|
||
|
||
### What Failed
|
||
- **Bug 1** (High): Category 3 audit is broken for ALL regular (non-framework) projects — task folder changes are always flagged as unauthorized
|
||
- **Bug 4** (High): Verdict inference uses `"PASS" in content` substring search, which can misclassify FAIL/NEEDS_REVIEW verdicts as complete
|
||
- **Bug 5** (High): `register-guards.sh` never successfully registers the OpenCode guard due to 3 compounding bugs (wrong filename, wrong config key, JSONC parsing)
|
||
- **Bug 7** (High): `--can-edit` file scope check uses string `startswith` without trailing separator, allowing edits in sibling directories
|
||
|
||
### What Needs Review
|
||
- **Bug 2** (Medium): `migrate-project.sh` `find` precedence silently skips `.md` files during migration
|
||
- **Bug 3** (Medium): `vram_detect.py` prefix matching gives wrong context windows for unknown models
|
||
- **Bug 6** (Medium): Dashboard ignores `.state` files, contradicting the "single source of truth" design
|
||
- **Bug 8** (Medium): Dashboard CORS `*` allows cross-origin writes from any website
|
||
- **Bug 9** (Low): Stale-task mtime proxy is unreliable after transitions
|
||
- **Bug 10** (Low): Artifact inference maps TEST_PLAN.md to `implement` instead of `test_design`
|
||
|
||
## Bug Finder vs Adversarial Bug Finder Comparison
|
||
- **Found by both**: None (reports are complementary by design)
|
||
- **Found only by Bug Finder**: Bugs 1–7 (enforcement logic, shell scripts, model lookup, guard registration)
|
||
- **Found only by Adversarial Bug Finder**: Bugs 8–10 (security, session tracking, phase inference edge case)
|
||
- **Contradictions**: None. The two reports are consistent and non-overlapping.
|
||
|
||
## Tasks for Review / Tie-Breaks
|
||
- None. No contradictions between Bug Finder and Adversarial Bug Finder. All 10 bugs are independently verified with reproduction steps.
|
||
|
||
## Remaining Issues
|
||
- Bugs 1, 4, 5, 7 (High severity) should be fixed before relying on the framework for production enforcement. Bug 5 in particular means the primary enforcement layer (harness pre-edit guard) is likely not registered for most users.
|
||
- Bug 6 means the dashboard display may not match the actual enforced state — users could make decisions based on stale dashboard info.
|
||
- Bug 7 is a security issue: the `--can-edit` scope check can be bypassed via sibling directory names.
|
||
|
||
## Score
|
||
+5 (NEEDS_REVIEW)
|
||
|
||
## Reviewer Comments
|
||
|