Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

6.7 KiB

SPEC: add-loop-templates-onboarding

Context

Tasks 2-5 shipped the brakes layer, runner, goal-mode work sources, and worktree creation. But the loop templates have roles: {implement: null, verify: null, orchestrate: null} -- no prompt references. And no loop prompt files exist in prompts/. This task creates the three loop role prompts, fills in both templates, and adds the critical missing piece: prompt-file token substitution in the runner so that {task_brief}, {acceptance_criteria}, etc. are resolved in the prompt content before the harness sees it.

Non-Goals (deferred)

  • tier budget enforcement in the runner -> v1.1 (the tier field in role config is documented but not enforced; the 16k context floor is the only hard gate).
  • harness.prompt_var / cwd_var / output_var -> v1.1 (the runner uses fixed token names; these config fields are documentation-only).
  • Prompt tuning / iteration -> ongoing (the prompts are v1 starters; real tuning happens when the self-improvement loop runs).
  • Onboarding wizard / interactive setup -> v1.1 (v1 ships docs only).

Requirements

R1 -- Prompt-file token substitution in loop-runner.py

  • New function _resolve_prompt(prompt_ref, extras, loop_path) -> str that:
    1. Resolves prompt_ref (e.g. "loop-implement.md") to a full path: check <loop_path>/<prompt_ref> first, then ~/.automaton/prompts/<prompt_ref>. If neither exists, return prompt_ref as-is (let the harness handle it).
    2. Reads the prompt file content.
    3. Substitutes content-level tokens in the prompt text: {task_brief}, {acceptance_criteria}, {next_hint}, {current_task}, {current_phase}, {verdict}, {artifact_content}.
    4. {artifact_content} is special: it reads the file at extras["artifact"] (the implement output path) and substitutes its content. If the file doesn't exist, substitutes empty string.
    5. Writes the substituted content to a temp file in <loop_path>/outputs/ (e.g. outputs/tickN-<role>-prompt.md).
    6. Returns the temp file path.
  • _invoke_harness is modified to call _resolve_prompt on the prompt_path before building the command. The returned temp file path replaces {prompt} in the command template.
  • If the prompt file doesn't exist (prompt_ref is None or file not found), the runner passes the raw prompt_ref as {prompt} (same as today -- backward compat).
  • Tests: test_resolve_prompt_substitutes_tokens, test_resolve_prompt_reads_artifact_content, test_resolve_prompt_fallback_when_file_missing, test_resolve_prompt_searches_loop_dir_then_framework.

R2 -- prompts/loop-implement.md

  • The Implement role prompt. Instructs the LLM to:
    • Read the task brief ({task_brief}), acceptance criteria ({acceptance_criteria}), and the previous tick's hint ({next_hint}).
    • Implement changes in the current working directory ({cwd}).
    • Write the artifact/implementation per the task's SPEC.
    • The current task is {current_task} in phase {current_phase}.
  • Follows the framework's prompt conventions (ALLOWED/FORBIDDEN sections, no auto-approve, status.py for transitions).
  • Tests: test_loop_implement_prompt_has_tokens, test_loop_implement_prompt_has_forbidden_section.

R3 -- prompts/loop-verifier.md

  • The Verify role prompt. Based on technical.md section 5. Instructs the LLM to:
    • Grade the artifact at {artifact_content} against {acceptance_criteria}.
    • Consider {task_brief} and {next_hint}.
    • Output strict JSON: {"pass": bool, "score": 0.0-1.0, "reasons": [...], "next_hint": "..."}.
    • Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress.
  • Tests: test_loop_verifier_prompt_has_json_instruction, test_loop_verifier_prompt_has_score_rubric, test_loop_verifier_prompt_has_tokens.

R4 -- prompts/loop-orchestrate.md

  • The Orchestrate role prompt. Instructs the LLM to:
    • Read the verdict ({verdict}).
    • Call exactly one status.py operation: --transition (if pass=true and task not complete), --approve (if in an approval-gated phase), or escalate to human_intervention (if pass=false or score is low).
    • No file edits. No auto-approve (D4).
    • The current task is {current_task} in phase {current_phase}.
  • Tests: test_loop_orchestrate_prompt_has_verdict_token, test_loop_orchestrate_prompt_has_no_edit_rule.

R5 -- Update templates/loops/ci-triage/loop.json

  • Fill in roles with prompt references:
    "roles": {
      "implement": {"prompt": "loop-implement.md"},
      "verify": {"prompt": "loop-verifier.md"},
      "orchestrate": {"prompt": "loop-orchestrate.md"}
    }
    
  • Keep all other fields unchanged.
  • Tests: test_ci_triage_template_has_prompt_refs.

R6 -- Create templates/loops/self-improvement/loop.json

  • Per technical.md section 9. Key fields:
    • name: "self-improvement"
    • work_source: {"kind": "audit", "project": "~/.automaton/"}
    • roles: same prompt refs as ci-triage
    • brakes: max_iterations: 10, score_plateau_window: 3
    • blast_radius: {"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}
    • acceptance_criteria: from technical.md section 9
    • schedule: {"interval_seconds": 3600}
  • Use "use_worktree" (not "worktree") for consistency with the code.
  • Tests: test_self_improvement_template_exists, test_self_improvement_template_has_audit_work_source, test_self_improvement_template_has_file_scope.

R7 -- Onboarding documentation

  • Add a "Loop Engineering" section to README.md (or update existing) with:
    • Quick start: status.py --create-loop <name> --from-template ci-triage -> --install-schedule <name>
    • How loops work: one-tick cycle diagram (gate -> find work -> worktree -> implement -> verify -> orchestrate -> state write)
    • How to configure: loop.json fields reference
    • How to monitor: --loop-list, --audit, .state.log
    • How to halt/resume: --approve --loop, --pause-loop, --resume-loop
  • Tests: none (doc-only).

R8 -- New test file tests/test_loop_templates.py

  • Covers R1-R6 as itemized above; target 12-16 tests.
  • Prompt-file substitution tests use tmp_path to create fake prompt files and verify the temp file output.
  • Template tests read the actual template files from templates/loops/.
  • Tests: self-referential.

R9 -- CHANGELOG and doc updates

  • CHANGELOG.md under [unreleased].
  • design/loops/technical.md section 8: note that the runner now resolves and substitutes prompt files.
  • Tests: none (doc-only).

Verification

  • python3 -m py_compile scripts/loop-runner.py
  • python3 -m pytest tests/test_loop_templates.py -v
  • python3 -m pytest tests/ -q -- full suite must remain green; expected total approx 385 (369 + 12-16 new).
  • bash -n scripts/*.sh (no shell changes; safety check).