Files
automaton/tasks/phase-scoped-prompts/IMPLEMENTATION.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

56 lines
3.2 KiB
Markdown

# Implementation: Phase-Scoped Prompts with Forbidden Actions
## Changes Made
### 1. ALLOWED/FORBIDDEN sections in all phase prompts
Each phase prompt now includes:
- **ALLOWED ACTIONS** — explicit list of what the agent can do
- **FORBIDDEN ACTIONS** — explicit list of what the agent cannot do, including "Do NOT" instructions and handling user overrides
Phase-specific definitions:
- research.md: ALLOWED read/ask questions/write SPEC.md; FORBIDDEN edit code, create IMPLEMENTATION.md, skip to implementation
- decompose.md: ALLOWED read SPEC/ask questions/write DECOMPOSITION.md; FORBIDDEN edit code, modify SPEC.md, create sub-task folders
- design.md: ALLOWED read SPEC/ask questions/write DESIGN.md; FORBIDDEN edit code, create IMPLEMENTATION.md, skip to implementation
- test_design.md: ALLOWED read SPEC+DESIGN/ask questions/write TEST_PLAN.md; FORBIDDEN edit code, write test implementations
- implement.md: ALLOWED edit code/write tests/create IMPLEMENTATION.md; FORBIDDEN create new tasks, modify SPEC/DESIGN
- bug_finder.md: ALLOWED read code/SPEC/IMPLEMENTATION/write BUG_REPORT.md; FORBIDDEN edit code, fix bugs
- adversarial_bug_find.md: ALLOWED read code/SPEC/BUG_REPORT/write ADVERSARIAL_BUG_REPORT.md; FORBIDDEN edit code, fix bugs
- doc_review.md: ALLOWED read DESIGN/code/docs/write DOC_REVIEW.md/update docs; FORBIDDEN edit non-doc code, modify SPEC/DESIGN
- referee.md: ALLOWED read all artifacts/write VERDICT.md; FORBIDDEN edit code, modify any artifact other than VERDICT.md
- orchestrate.md: ALLOWED read .state/transition state/create tasks/delegate; FORBIDDEN edit code directly, skip phases
### 2. User override resistance
Each prompt includes a "Handling User Overrides" section telling agents to refuse forbidden actions and suggest the correct phase.
### 3. `.state` precondition check
Every phase prompt includes `.state` as the first file to read, with instructions to STOP if the phase doesn't match.
### 4. Pre-Work Validation (MANDATORY)
Every phase prompt requires running `python ~/.automaton/scripts/status.py --validate-folder --task {task-name}` before starting work.
### 5. Approval gates
- research.md, decompose.md, design.md, test_design.md: include Approval Gate section with `--transition {phase}:awaiting_approval`, `--approve`, and `--transition {next-phase}`
- implement.md, bug_finder.md, adversarial_bug_find.md, doc_review.md, referee.md: include "No Approval Gate" section with direct `--transition`
### 6. Orchestrate.md restructuring
- Reduced from 493 to 143 lines
- State determination logic referenced from workflow.md
- Sub-task management extracted to subtask_management.md
- Gate-check loop with `--validate-folder` and approval pauses
## Files Modified
- `prompts/research.md` (updated)
- `prompts/design.md` (updated)
- `prompts/decompose.md` (updated)
- `prompts/test_design.md` (updated)
- `prompts/implement.md` (updated)
- `prompts/bug_finder.md` (updated)
- `prompts/adversarial_bug_find.md` (updated)
- `prompts/doc_review.md` (updated)
- `prompts/referee.md` (updated)
- `prompts/orchestrate.md` (rewritten, 143 lines)
- `prompts/subtask_management.md` (new, extracted)
## Test Results
- All prompt self-consistency tests passing
- onboarding.md excluded from stop-condition test