Files
automaton/tasks/add-loop-templates-onboarding/IMPLEMENTATION.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

94 lines
4.7 KiB
Markdown

# IMPLEMENTATION: add-loop-templates-onboarding
## Summary
Implemented prompt-file token substitution in the loop runner, created three loop role prompts, filled in both loop templates, and added onboarding documentation.
## Changes
### R1 -- `_resolve_prompt` in `scripts/loop-runner.py`
Added `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)` at line ~248:
- Searches `<loop_path>/<prompt_ref>` then `~/.automaton/prompts/<prompt_ref>` for the prompt file
- Reads the file content and substitutes content-level tokens: `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`
- `{artifact_content}` reads the file at `extras["artifact"]` path; empty string if missing
- Writes substituted content to `<loop_path>/outputs/tickN-<role>-prompt.md`
- Returns the temp file path
- Falls back to raw `prompt_ref` if file not found (backward compat)
Modified `_invoke_harness` signature to add `loop_path: Optional[Path] = None, tick_num: int = 0`. When `loop_path` is provided, calls `_resolve_prompt` on the prompt_path before building the harness command.
Updated all three `_invoke_harness` call sites in `cmd_tick` (implement ~L626, verify ~L638, orchestrate ~L669) to pass `loop_path=loop_path` and `tick_num=tick_num` where `tick_num = state.get('iteration_count', 0) + 1`.
### R2 -- `prompts/loop-implement.md`
Created the Implement role prompt with:
- `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}` tokens
- ALLOWED/FORBIDDEN sections (no `--transition`, no `--approve`, no file edits outside cwd)
- Instructions to read SPEC.md, implement code, run py_compile and pytest
### R3 -- `prompts/loop-verifier.md`
Created the Verify role prompt with:
- `{artifact_content}`, `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}` tokens
- Strict JSON output format: `{"pass": bool, "score": float, "reasons": [...], "next_hint": "..."}`
- Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress
- Empty artifact handling: returns `{"pass": false, "score": 0.0, ...}`
### R4 -- `prompts/loop-orchestrate.md`
Created the Orchestrate role prompt with:
- `{verdict}`, `{current_task}`, `{current_phase}` tokens
- Phase transition logic (implement -> code_review -> ... -> complete)
- FORBIDDEN: no file edits, no `--approve --loop` (human-only, D4), no auto-approve
### R5 -- `templates/loops/ci-triage/loop.json`
Updated `roles` from `null` values to prompt refs:
```json
"roles": {
"implement": {"prompt": "loop-implement.md"},
"verify": {"prompt": "loop-verifier.md"},
"orchestrate": {"prompt": "loop-orchestrate.md"}
}
```
### R6 -- `templates/loops/self-improvement/loop.json`
Created new template with:
- `work_source`: `{"kind": "audit", "project": "~/.automaton/"}`
- `blast_radius`: `{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}`
- `brakes`: `{"max_iterations": 10, "score_plateau_window": 3}`
- Same role prompt refs as ci-triage
### R7 -- Onboarding documentation
Updated `README.md` with a "Loop Engineering" section covering quick start, tick cycle, configuration, monitoring, and halt/resume.
### R8 -- `tests/test_loop_templates.py`
18 tests covering R1-R6:
- `TestResolvePrompt` (4 tests): token substitution, artifact content reading, fallback, search order
- `TestPromptFiles` (7 tests): prompt file content validation (tokens, JSON instructions, score rubric, FORBIDDEN sections)
- `TestCiTriageTemplate` (1 test): template has prompt refs
- `TestSelfImprovementTemplate` (5 tests): template exists, audit work source, file scope, prompt refs, brakes
- `TestTickPromptSubstitution` (1 test): end-to-end tick with prompt substitution
### R9 -- Doc updates
- `CHANGELOG.md`: added task 6 entry under `[unreleased]`
- `design/loops/technical.md` section 8: documented prompt resolution and substitution
### Test infrastructure updates
Updated `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py` to use non-existent prompt refs (`test-impl.md`, `test-verify.md`, `test-orch.md`) instead of real prompt file names. This prevents `_resolve_prompt` from activating in those tests, preserving backward compat behavior.
Updated `tests/test_framework_self_consistency.py` to exclude loop role prompts from the stop-condition check (they are role prompts, not delivery prompts).
## Verification
- `python3 -m py_compile scripts/loop-runner.py` -- OK
- `python3 -m pytest tests/test_loop_templates.py -v` -- 18 passed
- `python3 -m pytest tests/ -q` -- 393 passed (369 existing + 18 new + 6 from self-consistency recount)
- `bash -n scripts/*.sh` -- OK (no shell changes)