- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
4.7 KiB
IMPLEMENTATION: add-loop-templates-onboarding
Summary
Implemented prompt-file token substitution in the loop runner, created three loop role prompts, filled in both loop templates, and added onboarding documentation.
Changes
R1 -- _resolve_prompt in scripts/loop-runner.py
Added _resolve_prompt(prompt_ref, extras, loop_path, tick_num, role) at line ~248:
- Searches
<loop_path>/<prompt_ref>then~/.automaton/prompts/<prompt_ref>for the prompt file - Reads the file content and substitutes content-level tokens:
{task_brief},{acceptance_criteria},{next_hint},{current_task},{current_phase},{verdict},{artifact_content} {artifact_content}reads the file atextras["artifact"]path; empty string if missing- Writes substituted content to
<loop_path>/outputs/tickN-<role>-prompt.md - Returns the temp file path
- Falls back to raw
prompt_refif file not found (backward compat)
Modified _invoke_harness signature to add loop_path: Optional[Path] = None, tick_num: int = 0. When loop_path is provided, calls _resolve_prompt on the prompt_path before building the harness command.
Updated all three _invoke_harness call sites in cmd_tick (implement ~L626, verify ~L638, orchestrate ~L669) to pass loop_path=loop_path and tick_num=tick_num where tick_num = state.get('iteration_count', 0) + 1.
R2 -- prompts/loop-implement.md
Created the Implement role prompt with:
{task_brief},{acceptance_criteria},{next_hint},{current_task},{current_phase}tokens- ALLOWED/FORBIDDEN sections (no
--transition, no--approve, no file edits outside cwd) - Instructions to read SPEC.md, implement code, run py_compile and pytest
R3 -- prompts/loop-verifier.md
Created the Verify role prompt with:
{artifact_content},{task_brief},{acceptance_criteria},{next_hint},{current_task}tokens- Strict JSON output format:
{"pass": bool, "score": float, "reasons": [...], "next_hint": "..."} - Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress
- Empty artifact handling: returns
{"pass": false, "score": 0.0, ...}
R4 -- prompts/loop-orchestrate.md
Created the Orchestrate role prompt with:
{verdict},{current_task},{current_phase}tokens- Phase transition logic (implement -> code_review -> ... -> complete)
- FORBIDDEN: no file edits, no
--approve --loop(human-only, D4), no auto-approve
R5 -- templates/loops/ci-triage/loop.json
Updated roles from null values to prompt refs:
"roles": {
"implement": {"prompt": "loop-implement.md"},
"verify": {"prompt": "loop-verifier.md"},
"orchestrate": {"prompt": "loop-orchestrate.md"}
}
R6 -- templates/loops/self-improvement/loop.json
Created new template with:
work_source:{"kind": "audit", "project": "~/.automaton/"}blast_radius:{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}brakes:{"max_iterations": 10, "score_plateau_window": 3}- Same role prompt refs as ci-triage
R7 -- Onboarding documentation
Updated README.md with a "Loop Engineering" section covering quick start, tick cycle, configuration, monitoring, and halt/resume.
R8 -- tests/test_loop_templates.py
18 tests covering R1-R6:
TestResolvePrompt(4 tests): token substitution, artifact content reading, fallback, search orderTestPromptFiles(7 tests): prompt file content validation (tokens, JSON instructions, score rubric, FORBIDDEN sections)TestCiTriageTemplate(1 test): template has prompt refsTestSelfImprovementTemplate(5 tests): template exists, audit work source, file scope, prompt refs, brakesTestTickPromptSubstitution(1 test): end-to-end tick with prompt substitution
R9 -- Doc updates
CHANGELOG.md: added task 6 entry under[unreleased]design/loops/technical.mdsection 8: documented prompt resolution and substitution
Test infrastructure updates
Updated tests/test_loop_runner.py, tests/test_blast_radius.py, tests/test_goal_mode.py to use non-existent prompt refs (test-impl.md, test-verify.md, test-orch.md) instead of real prompt file names. This prevents _resolve_prompt from activating in those tests, preserving backward compat behavior.
Updated tests/test_framework_self_consistency.py to exclude loop role prompts from the stop-condition check (they are role prompts, not delivery prompts).
Verification
python3 -m py_compile scripts/loop-runner.py-- OKpython3 -m pytest tests/test_loop_templates.py -v-- 18 passedpython3 -m pytest tests/ -q-- 393 passed (369 existing + 18 new + 6 from self-consistency recount)bash -n scripts/*.sh-- OK (no shell changes)