Files
automaton/tasks/project-scoping-enforcement/ADVERSARIAL_BUG_REPORT.md
T
gitea 05c76852a2
CI / build (push) Has been cancelled
v2.0: state enforcement, project scoping, harness integration
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks

Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)

Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)

Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2026-06-15 14:16:46 -04:00

2.3 KiB

Adversarial Bug Report: project-scoping-enforcement

Summary

Adversarial review of the project scoping enforcement implementation. While the code changes are correct and well-tested, the process violation (implementing before tasking) reveals a deeper trust model issue.

Bugs Found

Bug 1: Agent can bypass the entire framework by editing files directly (Critical)

  • Severity: Critical
  • Description: All enforcement in the framework is prompt-based or tool-based (status.py). But nothing prevents an agent from directly editing scripts/status.py or any other file without a task. The --can-edit check only works if the agent chooses to call it. This is the same category of issue as the one we were fixing — the framework trusts the agent to follow its own rules.
  • Suggested Fix: This is inherent to prompt-driven frameworks. The fix is discipline, not code. However, we could add a git pre-commit hook that checks for .state file existence for modified files.

Bug 2: _infer_state_from_artifacts heuristic is still slightly wrong (Low)

  • Severity: Low
  • Description: In _infer_state_from_artifacts, when SPEC.md exists without BUG_REPORT.md, it returns bug_find instead of research. The logic at line 284 (if "SPEC.md" in artifacts and "BUG_REPORT.md" not in artifacts: return "bug_find") is incorrect — a task with only SPEC.md should be in research phase. However, since this heuristic is now only used by --upgrade (for migrating pre-v2.0 tasks), the impact is limited — the upgrade might assign a slightly wrong phase that the user can manually correct in .state.
  • Suggested Fix: Change line 284 to if "SPEC.md" in artifacts and "BUG_REPORT.md" not in artifacts and "IMPLEMENTATION.md" not in artifacts: return "research"

Bug 3: cmd_validate_folder still uses _infer_state_from_artifacts after .state exists (Low)

  • Severity: Low
  • Description: After confirming .state exists, cmd_validate_folder reads it and then falls back to _infer_state_from_artifacts if the read returns None (line 538). This shouldn't happen in practice — if .state exists, _read_state should return a value. But the fallback is unnecessary.
  • Suggested Fix: Remove the fallback and error instead.

Score

+5 (Bug 1 is a known limitation, Bug 2 and 3 are minor)