Files
automaton/.adversarial_bug_report.md
Lap Tran 81ccf548e5
CI / build (push) Has been cancelled
Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
2026-06-22 10:40:58 -04:00

4.3 KiB

Adversarial Bug Report: Full-Codebase Audit (v2.0)

Summary

Adversarial review targeting difficult-to-spot bugs: security gaps, logic errors in edge cases, and inconsistencies between enforcement layers. This report complements the Bug Report with findings that require deeper analysis. 3 additional bugs found.

Bugs Found

Bug 8: CORS * on writable dashboard API allows cross-origin modification

  • Severity: Medium
  • Location: automaton/dashboard/ui/app.py:40-44, 93-108, 213-239
  • Description: The dashboard HTTP server sets Access-Control-Allow-Origin: * on all responses, including POST and PUT endpoints. The dashboard binds to localhost, but any website open in the user's browser can send cross-origin requests to localhost:8080. The POST /api/task/{name}/review endpoint writes REVIEW.md files, and the PUT /api/config endpoint overwrites the dashboard config. A malicious webpage could silently modify task reviews or corrupt the config while the dashboard is running. The preflight OPTIONS handler (line 110-116) also returns Access-Control-Allow-Origin: * with Access-Control-Allow-Methods: GET, POST, PUT, OPTIONS, explicitly enabling these cross-origin writes.
  • Reproduction: With the dashboard running on localhost:8080, open a browser console on any website and run: fetch('http://localhost:8080/api/task/mytask/review', {method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify({status:'approved', comment:'hacked'})}). The review is written successfully.
  • Suggested Fix: Set Access-Control-Allow-Origin to http://localhost:8080 only (same-origin), or remove CORS headers entirely since the dashboard is a local single-origin app. Do not return * on POST/PUT endpoints.

Bug 9: Stale-task detection uses .state mtime as session proxy — any transition resets the timer

  • Severity: Low
  • Location: scripts/status.py:965, 1072 (--can-edit and --same-session)
  • Description: The stale-task check (line 965) and --same-session command (line 1072) both use the .state file's mtime to determine if a task is being actively worked on. However, every --transition command rewrites the .state file, resetting its mtime. This means: (1) A task that was transitioned to implement 29 minutes ago and then had a trivial --transition (e.g., to implement:awaiting_approval and back) would appear "fresh" even though no actual editing happened. (2) The --same-session check (which returns SAME_SESSION if mtime < 30 min) gives false positives after any transition, even by a different agent/session. The mtime is a proxy for "last state change," not "last edit activity."
  • Reproduction: Transition a task to implement. Wait 31 minutes. Run --same-session → DIFFERENT_SESSION. Now run any --transition (e.g., --transition implement). Run --same-session again → SAME_SESSION (even though no editing occurred).
  • Suggested Fix: Track the last edit activity separately (e.g., a .state.lastedit timestamp updated by --can-edit when editing is allowed), or use the task folder's newest file mtime instead of just .state.

Bug 10: _infer_state_from_artifacts maps TEST_PLAN.md to implement — skips test_design phase

  • Severity: Low
  • Location: scripts/status.py:324-325
  • Description: In _infer_state_from_artifacts(), the presence of TEST_PLAN.md without IMPLEMENTATION.md returns "implement". But TEST_PLAN.md is the artifact of the test_design phase (per PHASE_REQUIRED_ARTIFACTS at line 114). A task that has written a TEST_PLAN but hasn't started implementation should be in test_design, not implement. This causes the audit's fallback inference (for pre-v2.0 tasks without .state) to incorrectly report the task as being in a later phase than it actually is, potentially masking out-of-order artifact violations. The dashboard has the same mapping at task.py:510-511.
  • Reproduction: Create a task folder with only SPEC.md and TEST_PLAN.md (no .state, no IMPLEMENTATION.md). Run --validate-folder — the inferred phase is implement instead of test_design.
  • Suggested Fix: Map TEST_PLAN.md (without IMPLEMENTATION.md) to "test_design", not "implement".

Score

  • Bug 8 (Medium): +5
  • Bug 9 (Low): +1
  • Bug 10 (Low): +1

Total: 7

ADVERSARIAL_BUG_FIND_COMPLETE