Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

4.7 KiB

IMPLEMENTATION: add-loop-templates-onboarding

Summary

Implemented prompt-file token substitution in the loop runner, created three loop role prompts, filled in both loop templates, and added onboarding documentation.

Changes

R1 -- _resolve_prompt in scripts/loop-runner.py

Added _resolve_prompt(prompt_ref, extras, loop_path, tick_num, role) at line ~248:

  • Searches <loop_path>/<prompt_ref> then ~/.automaton/prompts/<prompt_ref> for the prompt file
  • Reads the file content and substitutes content-level tokens: {task_brief}, {acceptance_criteria}, {next_hint}, {current_task}, {current_phase}, {verdict}, {artifact_content}
  • {artifact_content} reads the file at extras["artifact"] path; empty string if missing
  • Writes substituted content to <loop_path>/outputs/tickN-<role>-prompt.md
  • Returns the temp file path
  • Falls back to raw prompt_ref if file not found (backward compat)

Modified _invoke_harness signature to add loop_path: Optional[Path] = None, tick_num: int = 0. When loop_path is provided, calls _resolve_prompt on the prompt_path before building the harness command.

Updated all three _invoke_harness call sites in cmd_tick (implement ~L626, verify ~L638, orchestrate ~L669) to pass loop_path=loop_path and tick_num=tick_num where tick_num = state.get('iteration_count', 0) + 1.

R2 -- prompts/loop-implement.md

Created the Implement role prompt with:

  • {task_brief}, {acceptance_criteria}, {next_hint}, {current_task}, {current_phase} tokens
  • ALLOWED/FORBIDDEN sections (no --transition, no --approve, no file edits outside cwd)
  • Instructions to read SPEC.md, implement code, run py_compile and pytest

R3 -- prompts/loop-verifier.md

Created the Verify role prompt with:

  • {artifact_content}, {task_brief}, {acceptance_criteria}, {next_hint}, {current_task} tokens
  • Strict JSON output format: {"pass": bool, "score": float, "reasons": [...], "next_hint": "..."}
  • Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress
  • Empty artifact handling: returns {"pass": false, "score": 0.0, ...}

R4 -- prompts/loop-orchestrate.md

Created the Orchestrate role prompt with:

  • {verdict}, {current_task}, {current_phase} tokens
  • Phase transition logic (implement -> code_review -> ... -> complete)
  • FORBIDDEN: no file edits, no --approve --loop (human-only, D4), no auto-approve

R5 -- templates/loops/ci-triage/loop.json

Updated roles from null values to prompt refs:

"roles": {
  "implement": {"prompt": "loop-implement.md"},
  "verify": {"prompt": "loop-verifier.md"},
  "orchestrate": {"prompt": "loop-orchestrate.md"}
}

R6 -- templates/loops/self-improvement/loop.json

Created new template with:

  • work_source: {"kind": "audit", "project": "~/.automaton/"}
  • blast_radius: {"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}
  • brakes: {"max_iterations": 10, "score_plateau_window": 3}
  • Same role prompt refs as ci-triage

R7 -- Onboarding documentation

Updated README.md with a "Loop Engineering" section covering quick start, tick cycle, configuration, monitoring, and halt/resume.

R8 -- tests/test_loop_templates.py

18 tests covering R1-R6:

  • TestResolvePrompt (4 tests): token substitution, artifact content reading, fallback, search order
  • TestPromptFiles (7 tests): prompt file content validation (tokens, JSON instructions, score rubric, FORBIDDEN sections)
  • TestCiTriageTemplate (1 test): template has prompt refs
  • TestSelfImprovementTemplate (5 tests): template exists, audit work source, file scope, prompt refs, brakes
  • TestTickPromptSubstitution (1 test): end-to-end tick with prompt substitution

R9 -- Doc updates

  • CHANGELOG.md: added task 6 entry under [unreleased]
  • design/loops/technical.md section 8: documented prompt resolution and substitution

Test infrastructure updates

Updated tests/test_loop_runner.py, tests/test_blast_radius.py, tests/test_goal_mode.py to use non-existent prompt refs (test-impl.md, test-verify.md, test-orch.md) instead of real prompt file names. This prevents _resolve_prompt from activating in those tests, preserving backward compat behavior.

Updated tests/test_framework_self_consistency.py to exclude loop role prompts from the stop-condition check (they are role prompts, not delivery prompts).

Verification

  • python3 -m py_compile scripts/loop-runner.py -- OK
  • python3 -m pytest tests/test_loop_templates.py -v -- 18 passed
  • python3 -m pytest tests/ -q -- 393 passed (369 existing + 18 new + 6 from self-consistency recount)
  • bash -n scripts/*.sh -- OK (no shell changes)