Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

3.2 KiB

Implementation: Phase-Scoped Prompts with Forbidden Actions

Changes Made

1. ALLOWED/FORBIDDEN sections in all phase prompts

Each phase prompt now includes:

  • ALLOWED ACTIONS — explicit list of what the agent can do
  • FORBIDDEN ACTIONS — explicit list of what the agent cannot do, including "Do NOT" instructions and handling user overrides

Phase-specific definitions:

  • research.md: ALLOWED read/ask questions/write SPEC.md; FORBIDDEN edit code, create IMPLEMENTATION.md, skip to implementation
  • decompose.md: ALLOWED read SPEC/ask questions/write DECOMPOSITION.md; FORBIDDEN edit code, modify SPEC.md, create sub-task folders
  • design.md: ALLOWED read SPEC/ask questions/write DESIGN.md; FORBIDDEN edit code, create IMPLEMENTATION.md, skip to implementation
  • test_design.md: ALLOWED read SPEC+DESIGN/ask questions/write TEST_PLAN.md; FORBIDDEN edit code, write test implementations
  • implement.md: ALLOWED edit code/write tests/create IMPLEMENTATION.md; FORBIDDEN create new tasks, modify SPEC/DESIGN
  • bug_finder.md: ALLOWED read code/SPEC/IMPLEMENTATION/write BUG_REPORT.md; FORBIDDEN edit code, fix bugs
  • adversarial_bug_find.md: ALLOWED read code/SPEC/BUG_REPORT/write ADVERSARIAL_BUG_REPORT.md; FORBIDDEN edit code, fix bugs
  • doc_review.md: ALLOWED read DESIGN/code/docs/write DOC_REVIEW.md/update docs; FORBIDDEN edit non-doc code, modify SPEC/DESIGN
  • referee.md: ALLOWED read all artifacts/write VERDICT.md; FORBIDDEN edit code, modify any artifact other than VERDICT.md
  • orchestrate.md: ALLOWED read .state/transition state/create tasks/delegate; FORBIDDEN edit code directly, skip phases

2. User override resistance

Each prompt includes a "Handling User Overrides" section telling agents to refuse forbidden actions and suggest the correct phase.

3. .state precondition check

Every phase prompt includes .state as the first file to read, with instructions to STOP if the phase doesn't match.

4. Pre-Work Validation (MANDATORY)

Every phase prompt requires running python ~/.automaton/scripts/status.py --validate-folder --task {task-name} before starting work.

5. Approval gates

  • research.md, decompose.md, design.md, test_design.md: include Approval Gate section with --transition {phase}:awaiting_approval, --approve, and --transition {next-phase}
  • implement.md, bug_finder.md, adversarial_bug_find.md, doc_review.md, referee.md: include "No Approval Gate" section with direct --transition

6. Orchestrate.md restructuring

  • Reduced from 493 to 143 lines
  • State determination logic referenced from workflow.md
  • Sub-task management extracted to subtask_management.md
  • Gate-check loop with --validate-folder and approval pauses

Files Modified

  • prompts/research.md (updated)
  • prompts/design.md (updated)
  • prompts/decompose.md (updated)
  • prompts/test_design.md (updated)
  • prompts/implement.md (updated)
  • prompts/bug_finder.md (updated)
  • prompts/adversarial_bug_find.md (updated)
  • prompts/doc_review.md (updated)
  • prompts/referee.md (updated)
  • prompts/orchestrate.md (rewritten, 143 lines)
  • prompts/subtask_management.md (new, extracted)

Test Results

  • All prompt self-consistency tests passing
  • onboarding.md excluded from stop-condition test