Files
automaton/tasks/complete/status-script/ADVERSARIAL_BUG_REPORT.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

1.7 KiB

Adversarial Bug Report: Status Script

Deep Review

The 980-line status.py is comprehensive. State machine transitions are correctly validated. Approval gates are enforced. Atomic writes prevent corruption. Error messages are clear and actionable.

Potential Issues

  1. Audit Category 3 stub: The git-based unauthorized modification check is stubbed. In projects that are git repos, the audit cannot detect code edits made during non-implement phases. This is a gap — an agent could edit code during research and the audit wouldn't catch it unless the artifacts reveal it.

  2. --same-session heuristic is weak: The 30-minute window for session detection is a best-effort heuristic. It cannot reliably distinguish sessions across agent restarts. The spec acknowledges this, but the implementation doesn't add much beyond mtime comparison.

  3. No --force flag for coordinator overrides: The multi-agent spec defines --force for coordinator force-transitions, but the current implementation returns "ERROR: unknown flag" for --force. This is expected to be added in the multi-agent-support task but a truly adversarial agent could use this gap.

  4. Symlink attack on .state.tmp: An attacker with filesystem access could create a symlink at .state.tmp pointing to a sensitive file, causing the atomic write to overwrite it. This is a local-privilege scenario, not a remote attack.

  5. No rate limiting on --create-task: A script could create thousands of task folders. Mitigation exists via kebab-case validation, but no limit on creation count.

Verdict: PASS — the audit stub is a known gap. No logic flaws that would compromise phase enforcement.