CI / build (push) Has been cancelled
State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks
Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)
Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)
Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
2.3 KiB
2.3 KiB
Adversarial Bug Report: project-scoping-enforcement
Summary
Adversarial review of the project scoping enforcement implementation. While the code changes are correct and well-tested, the process violation (implementing before tasking) reveals a deeper trust model issue.
Bugs Found
Bug 1: Agent can bypass the entire framework by editing files directly (Critical)
- Severity: Critical
- Description: All enforcement in the framework is prompt-based or tool-based (
status.py). But nothing prevents an agent from directly editingscripts/status.pyor any other file without a task. The--can-editcheck only works if the agent chooses to call it. This is the same category of issue as the one we were fixing — the framework trusts the agent to follow its own rules. - Suggested Fix: This is inherent to prompt-driven frameworks. The fix is discipline, not code. However, we could add a git pre-commit hook that checks for
.statefile existence for modified files.
Bug 2: _infer_state_from_artifacts heuristic is still slightly wrong (Low)
- Severity: Low
- Description: In
_infer_state_from_artifacts, whenSPEC.mdexists withoutBUG_REPORT.md, it returnsbug_findinstead ofresearch. The logic at line 284 (if "SPEC.md" in artifacts and "BUG_REPORT.md" not in artifacts: return "bug_find") is incorrect — a task with only SPEC.md should be inresearchphase. However, since this heuristic is now only used by--upgrade(for migrating pre-v2.0 tasks), the impact is limited — the upgrade might assign a slightly wrong phase that the user can manually correct in.state. - Suggested Fix: Change line 284 to
if "SPEC.md" in artifacts and "BUG_REPORT.md" not in artifacts and "IMPLEMENTATION.md" not in artifacts: return "research"
Bug 3: cmd_validate_folder still uses _infer_state_from_artifacts after .state exists (Low)
- Severity: Low
- Description: After confirming
.stateexists,cmd_validate_folderreads it and then falls back to_infer_state_from_artifactsif the read returns None (line 538). This shouldn't happen in practice — if.stateexists,_read_stateshould return a value. But the fallback is unnecessary. - Suggested Fix: Remove the fallback and error instead.
Score
+5 (Bug 1 is a known limitation, Bug 2 and 3 are minor)