State Enforcement (v2.0):
- .state file as single source of truth for task phase
- Approval gates for research, decomposition, design, test_design
- status.py --transition refuses illegal phase transitions
- status.py --validate-folder detects out-of-order artifacts
- status.py --audit checks all tasks for violations
- status.py --create-task is the only valid way to create tasks
- Pre-v2.0 tasks without .state are UNTRACKED -- all commands refuse them
- New --upgrade command bootstraps .state files for existing tasks
Project Scoping:
- --project flag added to all status.py commands across 16+ files
- _find_project_dir errors instead of silently falling back to ~/.automaton/
- --scope-check marks framework files OUT_OF_SCOPE when working on a project
- Dashboard handlers use stored project_root instead of re-detecting from CWD
- Prompts reference ~/.automaton/scripts/vram_detect.py (not {project}/.automaton/)
Harness Integration:
- status.py --can-edit now supports project-level checks (no --task required)
- --can-edit --file checks file scope without --task
- --json output for machine-readable harness integration
- opencode plugin (plugins/automaton-guard/plugin.ts) intercepts edit/write
- Git pre-commit hook (scripts/git-hooks/pre-commit) blocks commits without task
- Formal integration contract (contracts/harness-integration.md)
Other:
- upgrade.sh delegates to status.py --upgrade instead of manual heuristics
- Phase prompts reference --project {project} for multi-project scoping
- 200 tests passing (14 new)
4.2 KiB
Fix Verdict Parsing and State Machine Alignment
Goal
Fix the critical verdict-parsing bug that causes PASS verdicts to be falsely classified as BLOCKED, and align the dashboard's determine_task_state() with the orchestrator's state machine specification.
Requirements
R1. Use structured status-line parsing instead of substring search
automaton/dashboard/core/task.py:159-169 currently uses substring search for FAIL/NEEDS_REVIEW/PASS. This means a PASS verdict that mentions a previous failure (which referee.md explicitly requires when comparing bug finder outputs) gets misclassified as BLOCKED.
Fix: Parse the actual status line (## Status: PASS, **Status**: FAIL, etc.) extracted from the verdict content, falling back to substring search only when no structured status line is found.
R2. Fix verdict check ordering
The current code checks FAIL/NEEDS_REVIEW substrings before PASS. A correctly parsed status line makes this irrelevant for structured verdicts — only fall back to substring search for unstructured verdicts, using the same check order (check FAIL/NEEDS_REVIEW first, then PASS) but document the limitation.
R3. Use same parsing in parse_sub_tasks
task.py:215-220 has the same substring-search issue for sub-task verdicts. Apply the same fix.
R4. Align determine_task_state() with orchestrate.md state machine
Four concrete divergences between orchestrate.md:266-282 and task.py:135-198:
| Orchestrator says | Dashboard does | Fix |
|---|---|---|
IMPLEMENTATION.md → Bug Find |
IMPLEMENTATION.md → Implement |
Match orchestrator: show Bug Find when IMPLEMENTATION.md exists but no BUG_REPORT.md or ADVERSARIAL_BUG_REPORT.md |
BUG_REPORT.md + SPEC.md (no ADV) → Adversarial Bug Find |
BUG_REPORT.md alone → Bug Find |
Match orchestrator: BUG_REPORT.md → Bug Find, ADVERSARIAL_BUG_REPORT.md alone → Adversarial Bug Find. When both exist, advance to Doc Review or Referee. |
ADVERSARIAL_BUG_REPORT.md alone → not specified |
ADVERSARIAL_BUG_REPORT.md alone → ADV_BUG_FIND |
Follow orchestrator's intent: a lone ADVERSARIAL_BUG_REPORT without BUG_REPORT technically doesn't reach Adversarial Bug Find per spec. Treat ADV alone same as BUG alone for the dashboard (Bug Find). |
SPEC.md alone → Design or Implement |
SPEC.md alone → Research |
Keep dashboard behavior. The orchestrator spec says "Design or Implement" meaning those are the next steps the orchestrator would drive. The dashboard should show the task in its current state (Research). No change needed. |
R5. Update parse_sub_tasks to match the same logic
Sub-task state determination uses the same function, so these fixes propagate automatically. Verify that sub-tasks with only PARENT_SPEC.md or VRAM_CONFIG.md correctly show as BACKLOG.
R6. Document the minimal verdict schema
Add a note in prompts/referee.md requiring that VERDICT.md include ## Status: PASS / ## Status: FAIL / ## Status: NEEDS_REVIEW as a structured machine-parseable field. The dashboard relies on this for correct classification.
Acceptance Criteria
## Status: PASSverdict mentioning the word "FAIL" in findings → DONE (not BLOCKED)## Status: PASSverdict mentioning "NEEDS_REVIEW" in body → DONE (not BLOCKED)## Status: FAILverdict → BLOCKED## Status: NEEDS_REVIEWverdict → BLOCKEDIMPLEMENTATION.mdalone (no BUG_REPORT, no ADVERSARIAL_BUG_REPORT) → BUG_FIND (not IMPLEMENT)BUG_REPORT.md+SPEC.md(no ADVERSARIAL_BUG_REPORT) → BUG_FINDADVERSARIAL_BUG_REPORT.md+BUG_REPORT.md+SPEC.md→ ADV_BUG_FIND (or higher if DOC_REVIEW/VERDICT present)SPEC.mdalone → RESEARCH (unchanged, confirmed as correct)- Existing tests in
tests/test_task.pystill pass - New tests cover: PASS-verdict-mentions-FAIL, unstructured-verdict-fallback, implement-to-bug-find transition
parse_sub_taskscorrectly parses structured sub-task verdicts
Non-Goals
- Not removing substring fallback entirely (backward compat for unstructured verdicts)
- Not changing orchestrator.md (that spec is the authority)
- Not modifying
ui/app.pyverdict display logic