Fix 10 audit bugs: path prefix matching, verdict parsing, CORS, stale-task detection, phase mapping
CI / build (push) Has been cancelled

Batch 1 (High severity):
- Bug 1: --audit cat3 now checks .automaton/tasks/ paths
- Bug 4: Verdict PASS/FAIL uses structured ## Status: line parsing
- Bug 5: register-guards.sh checks .json/.jsonc, writes plugin key, strips comments
- Bug 7: --can-edit/--scope-check path prefix uses os.sep boundary

Batch 2 (Medium/Low severity):
- Bug 2: migrate-project.sh find command parentheses for -prune binding
- Bug 3: vram_detect model prefix matching with known-suffix whitelist
- Bug 6: dashboard reads .state file before artifact heuristic fallback
- Bug 8: removed wildcard CORS, added security headers (nosniff, DENY)
- Bug 9: stale-task detection uses .state.lastedit instead of .state mtime
- Bug 10: TEST_PLAN.md maps to test_design (was implement)

249 tests pass (up from 235). All 10 tasks driven through full workflow to completion.
This commit is contained in:
Lap Tran
2026-06-22 10:40:58 -04:00
parent f32f98575b
commit 81ccf548e5
106 changed files with 1643 additions and 436 deletions
+53 -233
View File
@@ -1,249 +1,69 @@
# Bug Report: automaton (Bug Finder Review — Post-Fix)
# Bug Report: Full-Codebase Audit (v2.0)
## Summary
Full-codebase audit of the automaton framework covering `scripts/status.py`, `scripts/vram_detect.py`, `scripts/migrate-project.sh`, `scripts/register-guards.sh`, `automaton/dashboard/`, `plugins/`, and git hooks. This report supersedes the previous orchestrate-only audit (all prior bugs marked fixed). 7 bugs found across enforcement logic, shell scripts, and the dashboard.
A comprehensive bug finder review of automaton identified **18 bugs** (17 new + 1 re-reported from the previous review). All bugs have been fixed. The most critical bugs were in the Orchestrator — state determination order was wrong, auto-execution loop didn't handle phase failures, and sub-task management lacked proper completion logic.
## Bugs Found
---
## Bugs Found and Fixed
### Bug 1: Orchestrator — State Determination order is wrong (CRITICAL — FIXED)
- **Severity**: Critical
- **Location**: `prompts/orchestrate.md`, "State Determination" section
- **Description**: The state determination checks from "most advanced state backward" but the order was inconsistent with the workflow.md state machine. The Orchestrator checked for IMPLEMENTATION.md (line 82) BEFORE checking for BUG_REPORT.md + SPEC.md (line 80), which meant if a task had both IMPLEMENTATION.md and BUG_REPORT.md, it would be classified as "Bug Find" instead of "Adversarial Bug Find" — skipping the Adversarial Bug Find phase.
- **Reproduction**: A task that has `BUG_REPORT.md`, `SPEC.md`, and `IMPLEMENTATION.md` would be classified as "Bug Find" instead of "Adversarial Bug Find".
- **Fix Applied**: Reordered the state determination to match the workflow.md exactly — from most advanced backward: VERDICT.md with PASS → Complete, VERDICT.md with FAIL/NEEDS_REVIEW → Human Intervention, DOC_REVIEW.md → Referee, ADVERSARIAL_BUG_REPORT.md + BUG_REPORT.md + SPEC.md → Doc Review, BUG_REPORT.md + SPEC.md without ADVERSARIAL_BUG_REPORT.md → Adversarial Bug Find, IMPLEMENTATION.md → Bug Find, etc. Also added overlapping conditions note.
### Bug 2: Orchestrator — Duplicate "In Autopilot mode" paragraph (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Default Mode — Autopilot" section
- **Description**: The paragraph "In Autopilot mode, the Orchestrator auto-executes all phases until completion or human intervention:" appeared twice consecutively (copy-paste duplication).
- **Reproduction**: Read the file; observe the duplicated sentence.
- **Fix Applied**: Removed the duplicate sentence.
### Bug 3: Orchestrator — Autopilot mode task creation creates empty IMPLEMENTATION.md (HIGH — FIXED)
### Bug 1: Category 3 audit false-positives for regular projects
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "New tasks from user input" section
- **Description**: When creating a new task from user input, the Orchestrator was creating the task folder with an empty `IMPLEMENTATION.md`. But the Orchestrator comment explicitly says "Do not pre-create it." This was inconsistent and could confuse the Research phase.
- **Reproduction**: Start a new task with "Research add user auth". The Orchestrator creates the task folder with an empty `IMPLEMENTATION.md`.
- **Fix Applied**: Added explicit note: "**Do not create an empty `IMPLEMENTATION.md` for new tasks** — this is inconsistent with the Orchestrator's own rule and can confuse the Research phase."
- **Location**: scripts/status.py:696
- **Description**: The Category 3 (unauthorized modifications) audit checks if changed files are inside task folders using `parts[0] == "tasks"`. This only works when the project directory IS `~/.automaton/` (framework mode), where git paths are `tasks/mytask/...`. For regular projects, task files have git paths like `.automaton/tasks/mytask/SPEC.md`, where `parts[0]` is `.automaton`, not `tasks`. The check fails, so ALL task-folder changes are flagged as "unauthorized modifications outside task folders."
- **Reproduction**: In a regular project (not `~/.automaton/`), create a task, transition to implement, make a change inside `.automaton/tasks/mytask/`, then run `python ~/.automaton/scripts/status.py --audit --project /path/to/project`. The task file is reported as unauthorized.
- **Suggested Fix**: Check both `parts[0] == "tasks"` (framework mode) and `len(parts) >= 3 and parts[0] == ".automaton" and parts[1] == "tasks" and parts[2] in task_names` (regular project mode).
### Bug 4: Orchestrator — Sub-task completion doesn't check if sub-task has VERDICT.md (HIGH — FIXED)
### Bug 2: `find` precedence bug silently skips `.md` files in migrate-project.sh
- **Severity**: Medium
- **Location**: scripts/migrate-project.sh:108
- **Description**: `find "$PROJECT_AUTOMATON" -maxdepth 1 -type f -name "*.md" -o -name "*.sh" -print0` lacks parentheses around the `-o` group. Without parens, `find` parses this as `(-name "*.md") -o (-name "*.sh" -print0)`. The `-print0` only applies to the `.sh` branch; `.md` files are found but never printed. This means customized `.md` files (`.agent.md`, `.rules.md`, `config.md`, etc.) are silently skipped during migration — they are never moved to `extensions/` or deleted.
- **Reproduction**: Create a temp dir with both `.md` and `.sh` files. Run `find "$dir" -maxdepth 1 -type f -name "*.md" -o -name "*.sh" -print0 | xargs -0 -I{} basename {}`. Only `.sh` files appear. With parentheses `\( -name "*.md" -o -name "*.sh" \) -print0`, both appear.
- **Suggested Fix**: Add parentheses: `find "$PROJECT_AUTOMATON" -maxdepth 1 -type f \( -name "*.md" -o -name "*.sh" \) -print0`
### Bug 3: Model name `startswith` matching gives wrong context windows for unknown models
- **Severity**: Medium
- **Location**: scripts/vram_detect.py:396
- **Description**: `_lookup_model_context()` uses `model_name.lower().startswith(key.lower())` to match model names against the `MODEL_CONTEXT_WINDOWS` dict. This prefix matching causes false matches: a model `phi-4-mini` matches `phi-4` (16000 tokens), `gpt-4o-foo-unknown` matches `gpt-4o` (128000), and any unknown model starting with a known prefix gets that prefix's context window instead of the fallback (128000). The comment on line 394 says "Strip common version/date suffixes for lookup" but the code doesn't strip anything — it uses prefix matching. Additionally, iteration order determines which key wins when multiple prefixes match, which may not be the most specific match.
- **Reproduction**: `python3 -c "import sys; sys.path.insert(0, 'scripts'); from vram_detect import _lookup_model_context; _lookup_model_context('phi-4-mini-instruct')"` — prints "Context window: 16k" instead of the fallback 128k.
- **Suggested Fix**: Try exact match first, then longest-prefix match (sort keys by length descending), or validate that the model name followed by `-` or end-of-string matches the key.
### Bug 4: VERDICT.md PASS substring inference misclassifies FAIL/NEEDS_REVIEW verdicts
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Completion and Parent Task" section
- **Description**: When a sub-task reaches a terminal state, the Orchestrator doesn't verify that the VERDICT.md exists before checking its verdict. If a sub-task somehow reaches a terminal state without a VERDICT.md (e.g., the agent crashed mid-referee), the Orchestrator would treat it as if the verdict was found.
- **Reproduction**: Sub-task folder has `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `DOC_REVIEW.md` but no `VERDICT.md`. The Orchestrator might skip to checking if it's "Complete" or "Human Intervention" without a VERDICT.md.
- **Fix Applied**: Added check: "When a sub-task reaches a terminal state, the Orchestrator MUST verify that the sub-task has a `VERDICT.md` before evaluating the verdict. If the sub-task does not have a `VERDICT.md`, it is NOT in a terminal state."
- **Location**: scripts/status.py:311
- **Description**: `_infer_state_from_artifacts()` checks `if "PASS" in content:` to determine if a verdict is PASS. This substring search matches "PASS" anywhere in the file — including in body text like "All unit tests PASS" or "NEEDS_REVIEW — but 2 tests PASS." A FAIL or NEEDS_REVIEW verdict containing the word "PASS" in its body is incorrectly classified as `complete`. The dashboard's `parse_verdict_status()` (task.py:63) correctly uses structured header-line parsing and explicitly avoids substring search (line 70-72), making this an inconsistency between status.py and the dashboard.
- **Reproduction**: Create a VERDICT.md with `## Status: FAIL` and body text "Note: 3 tests PASS." Run the audit — the task is classified as `complete` instead of `human_intervention`.
- **Suggested Fix**: Use the same structured-line parsing as the dashboard's `parse_verdict_status()`: look for `## Status:` header lines and check the value after the colon, not a substring search.
### Bug 5: Orchestrator — Parent task completion logic doesn't check all sub-tasks (HIGH — FIXED)
### Bug 5: register-guards.sh has three bugs preventing guard registration
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Completion and Parent Task" section
- **Description**: The Orchestrator says "When ALL sub-tasks are in terminal state: If ALL sub-tasks PASS: The parent task is Complete." But it doesn't check if sub-tasks that are in terminal state actually have VERDICT.md with PASS. It only checks if the verdict is PASS/FAIL/NEEDS_REVIEW.
- **Reproduction**: Parent task has 3 sub-tasks. Two have VERDICT.md with PASS. The third has DOC_REVIEW.md but no VERDICT.md (agent crashed). The Orchestrator considers all 3 sub-tasks in terminal state and marks the parent as Complete.
- **Fix Applied**: Added check: "Before checking completion, verify that each sub-task is still referenced in the DECOMPOSITION.md. If a sub-task is orphaned (no longer in the DECOMPOSITION.md), remove it from the parent's completion check."
- **Location**: scripts/register-guards.sh:23,34,33
- **Description**: Three bugs in the OpenCode guard registration:
- **5a** (line 23): Only checks for `opencode.jsonc`, not `opencode.json`. OpenCode supports both `.json` and `.jsonc` config files. If the user has `opencode.json` (the default), the guard is never registered. (Verified: this machine has `opencode.json`.)
- **5b** (line 34): Writes to the `plugins` key (`cfg.setdefault('plugins', [])`), but the OpenCode config uses `plugin` (singular). The actual config has `"plugin": ["opencode-mem"]`. Even if registration ran, it would add a `plugins` key that OpenCode ignores.
- **5c** (line 33): Uses `json.load()` to parse `.jsonc` files (JSON with Comments), which fails on files containing `//` comments. Since the file is `.jsonc`, comments are expected.
- **Reproduction**: Run `bash ~/.automaton/scripts/register-guards.sh` on a machine with `opencode.json` (not `.jsonc`). Output: "OpenCode: not detected (no ~/.config/opencode/opencode.jsonc)".
- **Suggested Fix**: Check for both `.json` and `.jsonc`. Write to the `plugin` key (singular). Use a JSONC-aware parser (strip comments before `json.load`) or use `json5` if available.
### Bug 6: Orchestrator — Sub-task creation doesn't handle DECOMPOSITION.md edge cases (MEDIUM — FIXED)
### Bug 6: Dashboard task.py ignores `.state` files — uses artifact heuristics only
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Management" section
- **Description**: When creating sub-task folders, the Orchestrator doesn't check if the DECOMPOSITION.md has sub-tasks with dependencies that span different waves. If a sub-task in Wave 2 depends on a sub-task in Wave 1, the Orchestrator should ensure Wave 1 sub-tasks are driven to completion before starting Wave 2.
- **Reproduction**: Parent task has Wave 1 (sub-task A, sub-task B) and Wave 2 (sub-task C depends on A and B). The Orchestrator creates all three sub-task folders and tries to run them all in parallel.
- **Fix Applied**: Added wave enforcement: "The Orchestrator MUST check sub-task dependencies and only start driving Wave 2 sub-tasks when all Wave 1 sub-tasks are in terminal state."
- **Location**: automaton/dashboard/core/task.py:462 (`determine_task_state`)
- **Description**: The dashboard's `determine_task_state()` infers task phase purely from which artifact files exist, never reading the `.state` file. This contradicts `prompts/workflow.md` which declares `.state` as the "single source of truth." Consequences: (1) A task in `implement` phase that hasn't written IMPLEMENTATION.md yet shows as an earlier phase. (2) A task in `test_design` with TEST_PLAN.md shows as `IMPLEMENT` (line 510-511 maps TEST_PLAN.md → IMPLEMENT). (3) A task in `bug_find` that already wrote BUG_REPORT.md shows as `BUG_FIND` even if `.state` says `adversarial_bug_find`. The dashboard cannot reflect the actual enforced state.
- **Reproduction**: Create a task, transition to `implement` via status.py, but don't write IMPLEMENTATION.md yet. Open the dashboard — the task shows as an earlier phase (e.g., `TEST_DESIGN` or `DESIGN`), not `IMPLEMENT`.
- **Suggested Fix**: Read the `.state` file first (as `status.py` does with `_read_state()`). Fall back to artifact heuristics only if `.state` doesn't exist (pre-v2.0 tasks).
### Bug 7: Orchestrator — "Continue" command doesn't handle sub-tasks (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Continue from existing tasks" section
- **Description**: When the user says "orchestrate" or "continue", the Orchestrator scans for the "most advanced task." But it doesn't distinguish between parent tasks and sub-tasks. A parent task with completed sub-tasks might be more "advanced" than a sub-task that's still in the Research phase.
- **Reproduction**: Parent task has 3 sub-tasks. Two sub-tasks are in the Research phase, one is in the Referee phase. The parent task has a VERDICT.md with FAIL. The Orchestrator picks the parent task instead of continuing the sub-task in the Referee phase.
- **Fix Applied**: Added: "prioritize sub-tasks over parent tasks" and "When choosing among sub-tasks in the same wave, prioritize those in later phases (e.g., Bug Find over Research) because they are closer to completion."
### Bug 8: Orchestrator — Auto-execution loop doesn't handle phase failures (HIGH — FIXED)
### Bug 7: Path prefix `startswith` allows scope bypass to sibling directories
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Loop" section
- **Description**: The auto-execution loop doesn't check for phase failures that are NOT verdicts — for example, if the agent crashes mid-phase or a phase doesn't produce the expected artifact. The loop assumes every phase produces its artifact and then checks the next phase.
- **Reproduction**: Bug Find phase produces an empty `BUG_REPORT.md`. The Orchestrator checks for the artifact, sees it exists, and proceeds to Adversarial Bug Find.
- **Fix Applied**: Added to the loop: "if phase artifact is empty or malformed: break (human intervention needed — artifact validation failed)"
### Bug 9: Orchestrator — Sub-task VRAM_CONFIG.md creation uses wrong units (LOW — FIXED)
- **Severity**: Low
- **Location**: `prompts/orchestrate.md`, "Sub-Task Verdict Reporting" section
- **Description**: The VRAM_CONFIG.md template uses `recommended_k` from the detection script output (which is in k units, e.g., "16" for 16k), but the template says "Target VRAM context: {from detection script or config.md override}" without specifying the unit.
- **Reproduction**: The detection script outputs "recommended_k: 16" and "max_peak_context_kb: 12000". The VRAM_CONFIG.md template uses "Target VRAM context: 16" (without the "k" suffix) and "Max peak context per sub-task: 12000" (without the "k" suffix), leading to ambiguity about units.
- **Fix Applied**: Clarified the units: "Target VRAM context: {value}k tokens (e.g., "16k")", "Max peak context per sub-task: {value}k tokens (e.g., "12k")", "This sub-task's estimated peak context: {value}k tokens (e.g., "10k")". Also added note: "The units must be clarified."
### Bug 10: Orchestrator — Sub-task PARENT_SPEC.md doesn't include sub-task scope (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Parent Specification" section
- **Description**: The PARENT_SPEC.md is supposed to contain "the parent task's SPEC.md content" and "any context the sub-task needs from the parent." But it doesn't include the sub-task's own scope/acceptance criteria from the DECOMPOSITION.md.
- **Reproduction**: Parent task's SPEC.md has 5 requirements. The DECOMPOSITION.md says sub-task A is only for requirements 1-2. The PARENT_SPEC.md only contains the parent's SPEC.md (all 5 requirements).
- **Fix Applied**: Changed the PARENT_SPEC.md content to include "The sub-task's own scope/acceptance criteria from the DECOMPOSITION.md" and added: "Do NOT include the parent task's full SPEC.md — this can cause circular references if the parent's SPEC.md references the sub-task's SPEC.md files."
### Bug 11: Orchestrator — No mechanism to handle sub-task failures in parent (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Completion and Parent Task" section
- **Description**: When a sub-task FAILs or NEEDS_REVIEW, the Orchestrator reports "human intervention is required" but doesn't create fix tasks for the failing sub-task.
- **Reproduction**: Sub-task A FAILs. The Orchestrator reports "human intervention is required." The parent task is stuck.
- **Fix Applied**: Added fix task creation for sub-tasks in Manual Mode: "Task name: `{parent-task-name}-fix-{sub-task-name}` (e.g., `add-user-auth-fix-auth-gateway`) — The Orchestrator creates the folder with an empty `IMPLEMENTATION.md` — The task starts at the **Bug Find** phase — The Orchestrator copies the sub-task's `SPEC.md`, `BUG_REPORT.md`, and `ADVERSARIAL_BUG_REPORT.md` (if they exist) into the new task folder"
### Bug 12: Orchestrator — Auto-detect VRAM doesn't handle missing detection script (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "VRAM Detection" section
- **Description**: The Orchestrator's VRAM detection priority says "Auto-detect via script: Run {project}/.automaton/scripts/vram_detect.sh". If the script is not available, it falls back to "Auto-detect via API config." But the Orchestrator doesn't check if the detection script exists before trying to run it.
- **Reproduction**: User starts a new task. The Orchestrator tries to run `{project}/.automaton/scripts/vram_detect.sh` but the script doesn't exist.
- **Fix Applied**: Added: "Check if `{project}/.automaton/scripts/vram_detect.sh` exists. If it does, run it..." and "If the detection script does not exist, skip to the next detection method."
### Bug 13: Orchestrator — Auto-detect VRAM doesn't handle script failure (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "VRAM Detection" section
- **Description**: Even if the detection script exists, it might fail (e.g., `nvidia-smi` is not installed, the GPU is busy, etc.). The Orchestrator doesn't handle script failures gracefully.
- **Reproduction**: The detection script exists but `nvidia-smi` is not installed. The script fails with an error.
- **Fix Applied**: Added: "If the detection script fails (e.g., `nvidia-smi` is not installed, the GPU is busy, etc.), check the exit code and fall back to the next detection method."
### Bug 14: Orchestrator — Sub-task completion doesn't aggregate verdicts for parent (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Verdict Reporting" section
- **Description**: The Orchestrator says "The Orchestrator should aggregate sub-task verdicts when reporting the parent task's status." But it doesn't actually implement this aggregation.
- **Reproduction**: Parent task has 3 sub-tasks. Two PASS, one FAIL. The Orchestrator reports the parent task as "Complete" because it doesn't aggregate sub-task verdicts.
- **Fix Applied**: Added: "The Orchestrator MUST aggregate sub-task verdicts when reporting the parent task's status: If ANY sub-task FAILs or NEEDS_REVIEW, the parent task should be marked as **Human Intervention** regardless of whether the parent's own VERDICT.md says PASS."
### Bug 15: Orchestrator — No mechanism to handle sub-task "Tie-Breaks" (LOW — FIXED)
- **Severity**: Low
- **Location**: `prompts/referee.md`
- **Description**: The referee has "Tasks for Review / Tie-Breaks" but the Orchestrator doesn't have logic to handle sub-task tie-breaks.
- **Reproduction**: Sub-task A has tie-breaks. The Orchestrator doesn't create tie-break tasks for the sub-task.
- **Fix Applied**: Added: "If a sub-task has 'Tasks for Review / Tie-Breaks' in its VERDICT.md, the Orchestrator should create tie-break tasks for the sub-task: Task name: `{parent-task-name}-tiebreak-{sub-task-name}`"
### Bug 16: Orchestrator — State Determination doesn't check for empty artifacts (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "State Determination" section
- **Description**: The Orchestrator checks for the existence of artifacts (e.g., "Has `SPEC.md`") but doesn't check if they are empty.
- **Reproduction**: Research phase produces an empty `SPEC.md`. The Orchestrator checks for `SPEC.md` and sees it exists, so it classifies the task as "Design" (optional) or "Implement".
- **Fix Applied**: Added "(non-empty)" after every artifact check: "Has `VERDICT.md` with `PASS` (non-empty)", "Has `DOC_REVIEW.md` (non-empty)", etc.
### Bug 17: Orchestrator — "Continue" doesn't prioritize sub-tasks in the same wave (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Continue from existing tasks" section
- **Description**: When the user says "orchestrate" or "continue", the Orchestrator should prioritize sub-tasks in the same wave that are still in progress.
- **Reproduction**: Wave 1 has sub-tasks A, B, and C. A is in the Research phase, B is in the Implement phase, and C is in the Bug Find phase. The Orchestrator picks A (Research phase) instead of C (Bug Find phase).
- **Fix Applied**: Added: "When choosing among sub-tasks in the same wave, prioritize those in later phases (e.g., Bug Find over Research) because they are closer to completion."
### Bug 18: Orchestrator — Auto-Execution Rules section is duplicated (LOW — FIXED)
- **Severity**: Low
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Rules (Autopilot Mode Only)" section
- **Description**: The section "Auto-Execution Rules (Autopilot Mode Only)" appears after the "Auto-Execution Loop" section and contains overlapping rules.
- **Reproduction**: Read the file; observe that the Auto-Execution Loop and Auto-Execution Rules sections contain overlapping rules.
- **Fix Applied**: Consolidated the Auto-Execution Loop and Auto-Execution Rules sections into a single section with additional rules for artifact validation, timeout, iteration limit, and sub-task parallel execution.
---
## Adversarial Bugs Found and Fixed
### Adversarial Bug 1: Orchestrator — Auto-execution loop can run infinitely (CRITICAL — FIXED)
- **Severity**: Critical
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Loop" section
- **Description**: The auto-execution loop has no maximum iteration count or timeout. If the agent produces an artifact but doesn't output CONTRACT_MET, the loop will spin forever.
- **Fix Applied**: Added to the loop: "if iteration_count >= MAX_ITERATIONS (default: 10): break (human intervention needed — too many iterations)", "if total_time_elapsed >= MAX_TOTAL_TIME (default: 24 hours): break (human intervention needed — too much time elapsed)", "if phase_time_elapsed >= MAX_PHASE_TIME (default: 1 hour): break (human intervention needed — phase took too long)"
### Adversarial Bug 2: Orchestrator — Sub-task creation doesn't prevent duplicate sub-tasks (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Management" section
- **Description**: When the Orchestrator creates sub-task folders, it doesn't check if they already exist.
- **Fix Applied**: Added: "Check for existing sub-task folders: For each sub-task, check if the folder `tasks/{parent-task-name}/subtasks/{sub-task-name}/` already exists. If it does, skip the creation and report that the sub-task has already been created."
### Adversarial Bug 3: Orchestrator — Auto-detect VRAM can cause resource exhaustion (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "VRAM Detection" section
- **Description**: If the VRAM detection script is run in a loop (e.g., the Orchestrator is invoked multiple times), it will repeatedly probe the GPU and RAM, causing performance degradation.
- **Fix Applied**: Added VRAM detection caching: "When the Orchestrator is invoked multiple times (e.g., the user says 'orchestrate' twice), it MUST cache the VRAM detection results and reuse them instead of running the detection script again."
### Adversarial Bug 4: Orchestrator — Sub-task completion doesn't check for orphaned sub-tasks (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Completion and Parent Task" section
- **Description**: The Orchestrator doesn't check if there are orphaned sub-tasks — sub-tasks that were created by the Orchestrator but are no longer referenced in the DECOMPOSITION.md.
- **Fix Applied**: Added: "Check for orphaned sub-tasks: Before checking completion, verify that each sub-task is still referenced in the DECOMPOSITION.md. If a sub-task is orphaned (no longer in the DECOMPOSITION.md), remove it from the parent's completion check."
### Adversarial Bug 5: Orchestrator — Auto-execution loop doesn't handle concurrent sub-tasks (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Loop" section
- **Description**: The auto-execution loop only drives one sub-task at a time, even when sub-tasks are in the same wave and can run in parallel.
- **Fix Applied**: Added: "Sub-task Parallel Execution: When sub-tasks are in the same wave and can run in parallel, the Orchestrator should drive them simultaneously instead of sequentially."
### Adversarial Bug 6: Orchestrator — State Determination can produce ambiguous states (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "State Determination" section
- **Description**: The state determination has multiple overlapping conditions that can produce ambiguous states.
- **Fix Applied**: Added: "Note on overlapping conditions: If a task has both `TEST_PLAN.md` and `DESIGN.md`, the Orchestrator should prioritize the more advanced state (TEST_PLAN.md → Implement) over the optional state (DESIGN.md → Test Design). Similarly, if a task has both `IMPLEMENTATION.md` and `BUG_REPORT.md`, the Orchestrator should prioritize the more advanced state (BUG_REPORT.md → Adversarial Bug Find) over the earlier state (IMPLEMENTATION.md → Bug Find)."
### Adversarial Bug 7: Orchestrator — Sub-task PARENT_SPEC.md can cause circular references (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Parent Specification" section
- **Description**: The PARENT_SPEC.md contains the parent task's SPEC.md content. If the parent's SPEC.md references the sub-task's SPEC.md files, a circular reference is created.
- **Fix Applied**: Changed the PARENT_SPEC.md content to include "The sub-task's own scope/acceptance criteria from the DECOMPOSITION.md" and added: "Do NOT include the parent task's full SPEC.md — this can cause circular references if the parent's SPEC.md references the sub-task's SPEC.md files."
### Adversarial Bug 8: Orchestrator — Auto-detect VRAM can cause memory exhaustion (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "VRAM Detection" section
- **Description**: If the detection script doesn't exist, the Orchestrator tries to read multiple config files to detect the model name. If the .env file is large, reading it could cause memory exhaustion.
- **Fix Applied**: Added: "Only read the specific lines needed (e.g., the model name line), not the entire file. Limit file reads to 10KB to prevent memory exhaustion."
### Adversarial Bug 9: Orchestrator — Sub-task completion doesn't handle sub-task failures gracefully (HIGH — FIXED)
- **Severity**: High
- **Location**: `prompts/orchestrate.md`, "Sub-Task Completion and Parent Task" section
- **Description**: When a sub-task FAILs during the Referee phase (before producing BUG_REPORT.md and ADVERSARIAL_BUG_REPORT.md), the Orchestrator wouldn't have the bug reports needed to create a fix task.
- **Fix Applied**: Added: "If the sub-task failed during the Referee phase (before producing BUG_REPORT.md and ADVERSARIAL_BUG_REPORT.md), the Orchestrator still creates the fix task but only copies the artifacts that exist (BUG_REPORT.md if it exists, ADVERSARIAL_BUG_REPORT.md if it exists)."
### Adversarial Bug 10: Orchestrator — Sub-task creation doesn't handle DECOMPOSITION.md updates (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Management" section
- **Description**: If the DECOMPOSITION.md is updated after the Orchestrator has already created sub-task folders, the Orchestrator doesn't handle the update.
- **Fix Applied**: Added: "Check for DECOMPOSITION.md updates: Compare the DECOMPOSITION.md with the existing sub-task folders. If the DECOMPOSITION.md has been updated (new sub-tasks added or existing sub-tasks removed), update the sub-task folders accordingly."
### Adversarial Bug 11: Orchestrator — Auto-execution loop doesn't handle phase timeouts (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Auto-Execution Loop" section
- **Description**: The auto-execution loop doesn't have a timeout for each phase.
- **Fix Applied**: Added to the loop: "if phase_time_elapsed >= MAX_PHASE_TIME (default: 1 hour): break (human intervention needed — phase took too long)"
### Adversarial Bug 12: Orchestrator — Sub-task VRAM_CONFIG.md doesn't include sub-task-specific VRAM limits (MEDIUM — FIXED)
- **Severity**: Medium
- **Location**: `prompts/orchestrate.md`, "Sub-Task Verdict Reporting" section
- **Description**: The VRAM_CONFIG.md includes "Max peak context per sub-task: {from detection script or config.md override}" which is the global max peak context from the detection script. But it doesn't include the sub-task's own estimated peak context from the DECOMPOSITION.md.
- **Fix Applied**: Added to the VRAM_CONFIG.md template: "This sub-task's estimated peak context: {from DECOMPOSITION.md}k tokens (e.g., "10k")" and "Fits within VRAM: Yes/No"
---
- **Location**: scripts/status.py:951, 1004, 1010, 1047, 1052
- **Description**: The `--can-edit` and `--scope-check` commands use `str(file_path).startswith(proj_str)` to verify a file is within the project directory. String `startswith` matches sibling directories: if `proj_str = "/home/user/project"`, then `/home/user/project-evil/file.py` matches because it starts with `/home/user/project`. This allows editing files outside the project boundary if a sibling directory with a similar name exists. The same bug affects the framework directory check (`auto_str`).
- **Reproduction**: `python3 -c "print('/Users/laptran/.automaton-evil/file'.startswith('/Users/laptran/.automaton'))"` → `True`. With trailing slash: `startswith('/Users/laptran/.automaton/')` → `False`.
- **Suggested Fix**: Append a trailing path separator: `str(file_path).startswith(proj_str + os.sep)` or use `Path.relative_to()` which correctly resolves path boundaries.
## Score
- Bug 1 (High): +10
- Bug 2 (Medium): +5
- Bug 3 (Medium): +5
- Bug 4 (High): +10
- Bug 5 (High): +10
- Bug 6 (Medium): +5
- Bug 7 (High): +10
| Bug | Severity | Score | Status |
|-----|----------|-------|--------|
| 1 | Critical | +10 | **FIXED** — State determination reordered |
| 2 | Medium | +5 | **FIXED** — Duplicate paragraph removed |
| 3 | High | +5 | **FIXED** — Empty IMPLEMENTATION.md creation removed |
| 4 | High | +5 | **FIXED** — VERDICT.md check added |
| 5 | High | +5 | **FIXED** — Orphaned sub-tasks check added |
| 6 | Medium | +5 | **FIXED** — Wave enforcement added |
| 7 | Medium | +5 | **FIXED** — Sub-task prioritization added |
| 8 | High | +5 | **FIXED** — Artifact validation added |
| 9 | Low | +1 | **FIXED** — Units clarified in VRAM_CONFIG.md |
| 10 | Medium | +5 | **FIXED** — Sub-task scope added to PARENT_SPEC.md |
| 11 | High | +5 | **FIXED** — Sub-task fix task creation added |
| 12 | Medium | +5 | **FIXED** — Script existence check added |
| 13 | Medium | +5 | **FIXED** — Script failure handling added |
| 14 | Medium | +5 | **FIXED** — Sub-task verdict aggregation added |
| 15 | Low | +1 | **FIXED** — Sub-task tie-break task creation added |
| 16 | Medium | +5 | **FIXED** — Empty artifact checks added |
| 17 | Medium | +5 | **FIXED** — Sub-task wave prioritization added |
| 18 | Low | +1 | **FIXED** — Auto-Execution Rules consolidated |
| Adv1 | Critical | +10 | **FIXED** — Auto-execution loop timeout/iteration limit added |
| Adv2 | High | +5 | **FIXED** — Duplicate sub-task prevention added |
| Adv3 | High | +5 | **FIXED** — VRAM detection caching added |
| Adv4 | High | +5 | **FIXED** — Orphaned sub-tasks check added |
| Adv5 | High | +5 | **FIXED** — Sub-task parallel execution added |
| Adv6 | Medium | +5 | **FIXED** — Overlapping conditions note added |
| Adv7 | Medium | +5 | **FIXED** — Circular reference prevention added |
| Adv8 | Medium | +5 | **FIXED** — Memory exhaustion prevention added |
| Adv9 | High | +5 | **FIXED** — Graceful sub-task failure handling added |
| Adv10 | Medium | +5 | **FIXED** — DECOMPOSITION.md update handling added |
| Adv11 | Medium | +5 | **FIXED** — Phase timeout added |
| Adv12 | Medium | +5 | **FIXED** — Sub-task VRAM limit added to VRAM_CONFIG.md |
| **Total** | | **143** | |
**Total: 55**