Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
CI / build (push) Has been cancelled

Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
This commit is contained in:
Lap Tran
2026-06-22 10:40:58 -04:00
parent f32f98575b
commit 81ccf548e5
106 changed files with 1643 additions and 436 deletions
+26 -91
View File
@@ -1,101 +1,36 @@
# Adversarial Bug Report: automaton (Adversarial Review — Post-Fix)
# Adversarial Bug Report: Full-Codebase Audit (v2.0)
## Summary
Adversarial review targeting difficult-to-spot bugs: security gaps, logic errors in edge cases, and inconsistencies between enforcement layers. This report complements the Bug Report with findings that require deeper analysis. 3 additional bugs found.
A deep adversarial review of automaton identified **12 bugs** that are difficult to spot — complex logic errors, race conditions, infinite loops, and memory/resource exhaustion issues. All bugs have been fixed. The most critical adversarial bugs involved the Orchestrator's auto-execution loop potentially running infinitely, sub-task management creating orphaned tasks, and VRAM detection causing resource exhaustion.
## Bugs Found
---
## Bugs Found and Fixed
### Bug 1: Orchestrator — Auto-execution loop can run infinitely (CRITICAL — FIXED)
- **Severity**: Critical
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Loop" section
- **Description**: The auto-execution loop has no maximum iteration count or timeout. If the agent produces an artifact but doesn't output CONTRACT_MET (e.g., the agent crashes), the loop will spin forever.
- **Fix Applied**: Added to the loop: "if iteration_count >= MAX_ITERATIONS (default: 10): break", "if total_time_elapsed >= MAX_TOTAL_TIME (default: 24 hours): break", "if phase_time_elapsed >= MAX_PHASE_TIME (default: 1 hour): break"
### Bug 2: Orchestrator — Sub-task creation doesn't prevent duplicate sub-tasks (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Management" section
- **Description**: When the Orchestrator creates sub-task folders, it doesn't check if they already exist.
- **Fix Applied**: Added: "Check for existing sub-task folders: For each sub-task, check if the folder `tasks/{parent-task-name}/subtasks/{sub-task-name}/` already exists. If it does, skip the creation and report that the sub-task has already been created."
### Bug 3: Orchestrator — Auto-detect VRAM can cause resource exhaustion (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "VRAM Detection" section
- **Description**: If the VRAM detection script is run in a loop (e.g., the Orchestrator is invoked multiple times), it will repeatedly probe the GPU and RAM, causing performance degradation.
- **Fix Applied**: Added VRAM detection caching: "When the Orchestrator is invoked multiple times (e.g., the user says 'orchestrate' twice), it MUST cache the VRAM detection results and reuse them instead of running the detection script again."
### Bug 4: Orchestrator — Sub-task completion doesn't check for orphaned sub-tasks (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Completion and Parent Task" section
- **Description**: The Orchestrator doesn't check if there are orphaned sub-tasks — sub-tasks that were created by the Orchestrator but are no longer referenced in the DECOMPOSITION.md.
- **Fix Applied**: Added: "Check for orphaned sub-tasks: Before checking completion, verify that each sub-task is still referenced in the DECOMPOSITION.md. If a sub-task is orphaned (no longer in the DECOMPOSITION.md), remove it from the parent's completion check."
### Bug 5: Orchestrator — Auto-execution loop doesn't handle concurrent sub-tasks (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Loop" section
- **Description**: The auto-execution loop only drives one sub-task at a time, even when sub-tasks are in the same wave and can run in parallel.
- **Fix Applied**: Added: "Sub-task Parallel Execution: When sub-tasks are in the same wave and can run in parallel, the Orchestrator should drive them simultaneously instead of sequentially."
### Bug 6: Orchestrator — State Determination can produce ambiguous states (MEDIUM — FIXED)
### Bug 8: CORS `*` on writable dashboard API allows cross-origin modification
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "State Determination" section
- **Description**: The state determination has multiple overlapping conditions that can produce ambiguous states.
- **Fix Applied**: Added: "Note on overlapping conditions: If a task has both `TEST_PLAN.md` and `DESIGN.md`, the Orchestrator should prioritize the more advanced state (TEST_PLAN.md → Implement) over the optional state (DESIGN.md → Test Design)."
- **Location**: automaton/dashboard/ui/app.py:40-44, 93-108, 213-239
- **Description**: The dashboard HTTP server sets `Access-Control-Allow-Origin: *` on all responses, including POST and PUT endpoints. The dashboard binds to `localhost`, but any website open in the user's browser can send cross-origin requests to `localhost:8080`. The POST `/api/task/{name}/review` endpoint writes REVIEW.md files, and the PUT `/api/config` endpoint overwrites the dashboard config. A malicious webpage could silently modify task reviews or corrupt the config while the dashboard is running. The preflight OPTIONS handler (line 110-116) also returns `Access-Control-Allow-Origin: *` with `Access-Control-Allow-Methods: GET, POST, PUT, OPTIONS`, explicitly enabling these cross-origin writes.
- **Reproduction**: With the dashboard running on localhost:8080, open a browser console on any website and run: `fetch('http://localhost:8080/api/task/mytask/review', {method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify({status:'approved', comment:'hacked'})})`. The review is written successfully.
- **Suggested Fix**: Set `Access-Control-Allow-Origin` to `http://localhost:8080` only (same-origin), or remove CORS headers entirely since the dashboard is a local single-origin app. Do not return `*` on POST/PUT endpoints.
### Bug 7: Orchestrator — Sub-task PARENT_SPEC.md can cause circular references (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Parent Specification" section
- **Description**: The PARENT_SPEC.md contains the parent task's SPEC.md content. If the parent's SPEC.md references the sub-task's SPEC.md files, a circular reference is created.
- **Fix Applied**: Changed the PARENT_SPEC.md content to include "The sub-task's own scope/acceptance criteria from the DECOMPOSITION.md" and added: "Do NOT include the parent task's full SPEC.md — this can cause circular references if the parent's SPEC.md references the sub-task's SPEC.md files."
### Bug 9: Stale-task detection uses `.state` mtime as session proxy — any transition resets the timer
- **Severity**: Low
- **Location**: scripts/status.py:965, 1072 (`--can-edit` and `--same-session`)
- **Description**: The stale-task check (line 965) and `--same-session` command (line 1072) both use the `.state` file's mtime to determine if a task is being actively worked on. However, every `--transition` command rewrites the `.state` file, resetting its mtime. This means: (1) A task that was transitioned to `implement` 29 minutes ago and then had a trivial `--transition` (e.g., to `implement:awaiting_approval` and back) would appear "fresh" even though no actual editing happened. (2) The `--same-session` check (which returns `SAME_SESSION` if mtime < 30 min) gives false positives after any transition, even by a different agent/session. The mtime is a proxy for "last state change," not "last edit activity."
- **Reproduction**: Transition a task to implement. Wait 31 minutes. Run `--same-session` → `DIFFERENT_SESSION`. Now run any `--transition` (e.g., `--transition implement`). Run `--same-session` again → `SAME_SESSION` (even though no editing occurred).
- **Suggested Fix**: Track the last edit activity separately (e.g., a `.state.lastedit` timestamp updated by `--can-edit` when editing is allowed), or use the task folder's newest file mtime instead of just `.state`.
### Bug 8: Orchestrator — Auto-detect VRAM can cause memory exhaustion (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "VRAM Detection" section
- **Description**: If the detection script doesn't exist, the Orchestrator tries to read multiple config files to detect the model name. If the .env file is large, reading it could cause memory exhaustion.
- **Fix Applied**: Added: "Only read the specific lines needed (e.g., the model name line), not the entire file. Limit file reads to 10KB to prevent memory exhaustion."
### Bug 9: Orchestrator — Sub-task completion doesn't handle sub-task failures gracefully (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Completion and Parent Task" section
- **Description**: When a sub-task FAILs during the Referee phase (before producing BUG_REPORT.md and ADVERSARIAL_BUG_REPORT.md), the Orchestrator wouldn't have the bug reports needed to create a fix task.
- **Fix Applied**: Added: "If the sub-task failed during the Referee phase (before producing BUG_REPORT.md and ADVERSARIAL_BUG_REPORT.md), the Orchestrator still creates the fix task but only copies the artifacts that exist (BUG_REPORT.md if it exists, ADVERSARIAL_BUG_REPORT.md if it exists)."
### Bug 10: Orchestrator — Sub-task creation doesn't handle DECOMPOSITION.md updates (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Management" section
- **Description**: If the DECOMPOSITION.md is updated after the Orchestrator has already created sub-task folders, the Orchestrator doesn't handle the update.
- **Fix Applied**: Added: "Check for DECOMPOSITION.md updates: Compare the DECOMPOSITION.md with the existing sub-task folders. If the DECOMPOSITION.md has been updated (new sub-tasks added or existing sub-tasks removed), update the sub-task folders accordingly."
### Bug 11: Orchestrator — Auto-execution loop doesn't handle phase timeouts (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Loop" section
- **Description**: The auto-execution loop doesn't have a timeout for each phase.
- **Fix Applied**: Added to the loop: "if phase_time_elapsed >= MAX_PHASE_TIME (default: 1 hour): break (human intervention needed — phase took too long)"
### Bug 12: Orchestrator — Sub-task VRAM_CONFIG.md doesn't include sub-task-specific VRAM limits (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Verdict Reporting" section
- **Description**: The VRAM_CONFIG.md includes "Max peak context per sub-task: {from detection script or config.md override}" which is the global max peak context from the detection script. But it doesn't include the sub-task's own estimated peak context from the DECOMPOSITION.md.
- **Fix Applied**: Added to the VRAM_CONFIG.md template: "This sub-task's estimated peak context: {from DECOMPOSITION.md}k tokens (e.g., "10k")" and "Fits within VRAM: Yes/No"
---
### Bug 10: `_infer_state_from_artifacts` maps TEST_PLAN.md to `implement` — skips test_design phase
- **Severity**: Low
- **Location**: scripts/status.py:324-325
- **Description**: In `_infer_state_from_artifacts()`, the presence of `TEST_PLAN.md` without `IMPLEMENTATION.md` returns `"implement"`. But `TEST_PLAN.md` is the artifact of the `test_design` phase (per `PHASE_REQUIRED_ARTIFACTS` at line 114). A task that has written a TEST_PLAN but hasn't started implementation should be in `test_design`, not `implement`. This causes the audit's fallback inference (for pre-v2.0 tasks without `.state`) to incorrectly report the task as being in a later phase than it actually is, potentially masking out-of-order artifact violations. The dashboard has the same mapping at task.py:510-511.
- **Reproduction**: Create a task folder with only `SPEC.md` and `TEST_PLAN.md` (no `.state`, no `IMPLEMENTATION.md`). Run `--validate-folder` — the inferred phase is `implement` instead of `test_design`.
- **Suggested Fix**: Map `TEST_PLAN.md` (without `IMPLEMENTATION.md`) to `"test_design"`, not `"implement"`.
## Score
- Bug 8 (Medium): +5
- Bug 9 (Low): +1
- Bug 10 (Low): +1
| Bug | Severity | Score | Status |
|-----|----------|-------|--------|
| 1 | Critical | +10 | **FIXED** — Auto-execution loop timeout/iteration limit added |
| 2 | High | +5 | **FIXED** — Duplicate sub-task prevention added |
| 3 | High | +5 | **FIXED** — VRAM detection caching added |
| 4 | High | +5 | **FIXED** — Orphaned sub-tasks check added |
| 5 | High | +5 | **FIXED** — Sub-task parallel execution added |
| 6 | Medium | +5 | **FIXED** — Overlapping conditions note added |
| 7 | Medium | +5 | **FIXED** — Circular reference prevention added |
| 8 | Medium | +5 | **FIXED** — Memory exhaustion prevention added |
| 9 | High | +5 | **FIXED** — Graceful sub-task failure handling added |
| 10 | Medium | +5 | **FIXED** — DECOMPOSITION.md update handling added |
| 11 | Medium | +5 | **FIXED** — Phase timeout added |
| 12 | Medium | +5 | **FIXED** — Sub-task VRAM limit added to VRAM_CONFIG.md |
| **Total** | | **65** | |
**Total: 7**
ADVERSARIAL_BUG_FIND_COMPLETE