Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test

- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
This commit is contained in:
Lap Tran
2026-06-26 10:05:18 -04:00
parent fe43b9e1fc
commit bc7daf8590
666 changed files with 15994 additions and 69 deletions
@@ -0,0 +1 @@
complete
@@ -0,0 +1,2 @@
research:approved|2026-06-23T12:53:11.686771+00:00|user
code_review:approved|2026-06-23T13:01:35.363128+00:00|user
@@ -0,0 +1,55 @@
# ADVERSARIAL_BUG_REPORT: add-loop-templates-onboarding
## Methodology
Targeted attack on the weakest points of the implementation:
1. Path traversal via `prompt_ref`
2. Token injection via `extras` values
3. Race condition on `outputs/` directory
4. Large file DoS via `{artifact_content}`
5. Unicode/encoding edge cases
6. Concurrent ticks writing to the same `outputs/` dir
## Findings
### Attack 1: Path traversal via `prompt_ref` -- NOT VULNERABLE
`_resolve_prompt` constructs candidate paths as `loop_path / prompt_ref` and `AUTOMATON_DIR / "prompts" / prompt_ref`. If `prompt_ref` were `"../../etc/passwd"`, `Path / "../../etc/passwd"` would resolve to a path outside the loop dir. However, `prompt_ref` comes from `loop.json` `roles.*.prompt`, which is a trusted config file written by the user/framework. An attacker who can write `loop.json` already has full code execution via `harness.command`. No additional risk.
**Verdict:** NOT VULNERABLE (trusted input)
### Attack 2: Token injection via extras values -- NOT VULNERABLE
If `task_brief` contained `{task_brief}`, the `str(value)` substitution would not cause infinite recursion because `content.replace` is a single-pass operation. The substituted value is inserted as-is, and no further substitution is applied to the result. No injection vector.
**Verdict:** NOT VULNERABLE
### Attack 3: Race condition on `outputs/` directory -- NOT EXPLOITABLE
`out_dir.mkdir(parents=True, exist_ok=True)` is atomic. If two ticks run concurrently (which the scheduler should prevent, but could happen in daemon mode with a bug), they would write to different files (`tickN-<role>-prompt.md` where N differs). The only shared state is the directory itself, and `mkdir(exist_ok=True)` handles that. The `.state.loop` write is atomic (tmp+rename), so `iteration_count` won't be corrupted.
**Verdict:** NOT EXPLOITABLE (different tick numbers produce different file paths)
### Attack 4: Large file DoS via `{artifact_content}` -- ACCEPTED RISK
If the implement artifact is very large (e.g. 10MB), `{artifact_content}` reads the entire file into memory and substitutes it into the prompt. This could produce a prompt that exceeds the model's context window. However, the runner already has a `_truncate_tokens` function (from task 4) that caps `task_brief` at 4k tokens, `acceptance_criteria` at 2k, and `next_hint` at 1k. The `{artifact_content}` token is NOT truncated, which is by design -- the verifier needs to see the full artifact to grade it. The 16k context floor gate (D13) catches undersized contexts before the harness is invoked. For oversized contexts, the harness's own context management handles it.
**Verdict:** ACCEPTED RISK (mitigated by context floor gate and harness-side context management)
### Attack 5: Unicode/encoding edge cases -- NOT VULNERABLE
`Path.read_text()` and `Path.write_text()` use UTF-8 by default on all platforms. The `str(value)` conversion handles all Python string types. No encoding issues found.
**Verdict:** NOT VULNERABLE
### Attack 6: Concurrent ticks writing to same `outputs/` dir -- NOT EXPLOITABLE
Same as Attack 3. Different tick numbers produce different file paths. The `.state.loop` atomic write prevents `iteration_count` corruption.
**Verdict:** NOT EXPLOITABLE
## Summary
No exploitable vulnerabilities found. All attack surfaces are either mitigated by existing controls (context floor gate, atomic state writes, trusted input assumption) or produce no harmful behavior.
**Verdict: CLEAN**
@@ -0,0 +1,34 @@
# BUG_REPORT: add-loop-templates-onboarding
## Methodology
Adversarial review of all changed files. Searched for: race conditions, token injection, path traversal, missing error handling, backward compat breaks, and edge cases in prompt resolution.
## Findings
### Bug 1 (LOW): `_resolve_prompt` writes temp file even when no tokens are substituted
If a prompt file exists but contains no tokens (e.g. a static prompt), `_resolve_prompt` still reads it, does the substitution loop (which is a no-op), and writes a copy to `outputs/tickN-<role>-prompt.md`. This is wasteful but not incorrect -- the harness receives an identical prompt either way. The temp file provides an audit trail of what was sent to the harness, which is actually useful for debugging.
**Severity:** LOW (performance/ cleanliness, not correctness)
**Fix:** None needed for v1. The audit trail value outweighs the minor I/O cost.
### Bug 2 (LOW): No token for `{cwd}` in content-level substitution
The harness command template supports `{cwd}` as an argv-level token, but `_resolve_prompt` does not substitute `{cwd}` in the prompt file content. If a prompt author writes `{cwd}` in the prompt text, it will appear literally in the resolved prompt. The SPEC does not list `{cwd}` as a content-level token (R1 lists `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`), so this is by design -- `{cwd}` is a harness-command token, not a content-level token.
**Severity:** LOW (documentation, not a bug)
**Fix:** None needed. The prompt files use "Working directory: the cwd you were launched with" instead of `{cwd}`.
### Bug 3 (INFO): `loop-orchestrate.md` references `code_review:awaiting_approval` then `--approve` in one step
The orchestrate prompt says "If in `code_review`: transition to `code_review:awaiting_approval`, then approve." This is two `status.py` calls in one tick. The orchestrator role is a single LLM session that can make multiple CLI calls, so this is valid. The runner does not restrict the number of subprocess calls the orchestrator makes.
**Severity:** INFO (not a bug)
**Fix:** None needed.
## Summary
No correctness bugs found. Two LOW-severity observations and one INFO note. The implementation is solid for v1.
**Verdict: CLEAN**
@@ -0,0 +1,78 @@
# CODE_REVIEW: add-loop-templates-onboarding
## Reviewed Files
1. `scripts/loop-runner.py` -- `_resolve_prompt` function (lines ~248-303), `_invoke_harness` signature change (lines ~306-338), `cmd_tick` call site updates (lines ~626, ~638, ~669)
2. `prompts/loop-implement.md` -- new file
3. `prompts/loop-verifier.md` -- new file
4. `prompts/loop-orchestrate.md` -- new file
5. `templates/loops/ci-triage/loop.json` -- roles updated
6. `templates/loops/self-improvement/loop.json` -- new file
7. `tests/test_loop_templates.py` -- new test file (18 tests)
8. `tests/test_loop_runner.py` -- prompt ref renames
9. `tests/test_blast_radius.py` -- prompt ref renames
10. `tests/test_goal_mode.py` -- prompt ref renames
11. `tests/test_framework_self_consistency.py` -- exclusion set update
12. `README.md` -- Loop Engineering onboarding section
13. `CHANGELOG.md` -- task 6 entry
14. `design/loops/technical.md` -- section 8 prompt resolution docs
## Findings
### 1. `_resolve_prompt` -- token substitution correctness
The function correctly handles the two-stage search (loop-local then framework), reads the file, substitutes tokens, writes to outputs/, and returns the temp path. The fallback to raw `prompt_ref` when the file is not found preserves backward compatibility.
**Concern: token injection.** The `str(value)` substitution via `content.replace("{" + key + "}", str(value))` is safe for the current token set (all values are controlled: task_brief from SPEC.md, acceptance_criteria from loop.json, etc.). No user-supplied input flows into these tokens without being read from a file first. Acceptable for v1.
**Verdict:** PASS
### 2. `_invoke_harness` signature change
The new `loop_path` and `tick_num` parameters are optional with defaults (`None` and `0`). Existing callers that don't pass them get the old behavior (raw prompt_ref passed through). This is backward compatible.
**Verdict:** PASS
### 3. `cmd_tick` call site updates
All three call sites (implement, verify, orchestrate) now pass `loop_path=loop_path` and `tick_num=tick_num` where `tick_num` is computed once as `state.get('iteration_count', 0) + 1`. This is correct -- the tick number should be consistent across all three role invocations in the same tick.
**Verdict:** PASS
### 4. Prompt file content
- `loop-implement.md`: has all required tokens, ALLOWED/FORBIDDEN sections, no auto-approve. Correct.
- `loop-verifier.md`: has strict JSON output format, score rubric, artifact_content token. Correct.
- `loop-orchestrate.md`: has verdict token, phase transition logic, no-edit rule. Correct.
All three prompts are excluded from the self-consistency stop-condition check since they are role prompts, not delivery prompts. This is consistent with how `orchestrate.md` is already excluded.
**Verdict:** PASS
### 5. Template updates
- `ci-triage/loop.json`: roles filled with `{"prompt": "loop-implement.md"}` etc. All other fields unchanged. Correct.
- `self-improvement/loop.json`: has `work_source: audit`, `use_worktree: true`, `file_scope` with 4 paths, `max_iterations: 10`, `score_plateau_window: 3`. Matches technical.md section 9. Correct.
**Verdict:** PASS
### 6. Test infrastructure updates
Renaming prompt refs from `"loop-implement.md"` to `"test-impl.md"` (and similar) in existing tests is the correct approach. These tests don't test prompt resolution -- they test other runner behavior. Using non-existent prompt refs ensures `_resolve_prompt` falls back to the raw string, preserving the old argv contents that the test assertions depend on.
**Verdict:** PASS
### 7. Edge cases
- **Empty prompt_ref**: `_resolve_prompt` returns `prompt_ref or ""` at line 260. Safe.
- **Missing outputs dir**: `out_dir.mkdir(parents=True, exist_ok=True)` at line 300. Safe.
- **Missing artifact file for `{artifact_content}`**: caught by `try/except OSError`, returns empty string. Safe.
- **Loop-local prompt override**: searched first, allows per-loop customization without modifying framework prompts. Good design.
**Verdict:** PASS
## Summary
All 7 review areas pass. The implementation is correct, backward compatible, and well-tested. 18 new tests cover the prompt resolution, prompt file content, template updates, and tick integration. Full suite: 393 passed.
**Overall verdict: APPROVED**
@@ -0,0 +1,56 @@
# DOC_REVIEW: add-loop-templates-onboarding
## Reviewed Documentation
1. `README.md` -- new "Loop Engineering" onboarding section (Quick Start, Tick Cycle, Configuration, Monitoring, Halt/Resume)
2. `CHANGELOG.md` -- task 6 entry under `[unreleased]`
3. `design/loops/technical.md` section 8 -- prompt resolution and token substitution documentation
4. `AGENTS.md` -- no changes needed (already documents loop runner and status.py commands)
## Findings
### 1. README.md onboarding section
The new section adds:
- Quick Start with 3 commands (create, install-schedule, monitor)
- Tick Cycle diagram (11-step flow summary)
- Configuration table with all `loop.json` fields
- Monitoring commands
- Halt/Resume commands
**Accuracy:** All commands and field names match the actual implementation. The configuration table correctly documents `use_worktree` (not `worktree`), `work_source.kind` values (`single`, `audit`, `backlog`), and the role prompt fields.
**Completeness:** Covers all R7 sub-requirements from the SPEC.
**Verdict:** PASS
### 2. CHANGELOG.md
Entry accurately describes all changes: `_resolve_prompt`, `_invoke_harness` extension, new prompt files, template updates, new self-improvement template, README section, technical.md section 8, new tests (18), test infrastructure updates.
**Verdict:** PASS
### 3. design/loops/technical.md section 8
New "Prompt Resolution and Token Substitution" subsection documents:
- File search order (loop-local then framework)
- Content-level token substitution
- `{artifact_content}` special handling
- Temp file write and return path
- Loop-local override capability
**Accuracy:** Matches the implementation in `_resolve_prompt`.
**Verdict:** PASS
### 4. Cross-reference check
- `AGENTS.md` "Loop runner" bullet references `design/loops/technical.md` §7 for the tick flow -- still accurate.
- `config.md` mentions role-to-prompt binding in `loop.json` -- still accurate.
- `prompts/` directory now has 3 new files (`loop-implement.md`, `loop-verifier.md`, `loop-orchestrate.md`) -- not listed in any index (there is no prompts/ index file), so no update needed.
## Summary
All documentation is accurate, complete, and consistent with the implementation. No doc gaps found.
**Verdict: APPROVED**
@@ -0,0 +1,93 @@
# IMPLEMENTATION: add-loop-templates-onboarding
## Summary
Implemented prompt-file token substitution in the loop runner, created three loop role prompts, filled in both loop templates, and added onboarding documentation.
## Changes
### R1 -- `_resolve_prompt` in `scripts/loop-runner.py`
Added `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)` at line ~248:
- Searches `<loop_path>/<prompt_ref>` then `~/.automaton/prompts/<prompt_ref>` for the prompt file
- Reads the file content and substitutes content-level tokens: `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`
- `{artifact_content}` reads the file at `extras["artifact"]` path; empty string if missing
- Writes substituted content to `<loop_path>/outputs/tickN-<role>-prompt.md`
- Returns the temp file path
- Falls back to raw `prompt_ref` if file not found (backward compat)
Modified `_invoke_harness` signature to add `loop_path: Optional[Path] = None, tick_num: int = 0`. When `loop_path` is provided, calls `_resolve_prompt` on the prompt_path before building the harness command.
Updated all three `_invoke_harness` call sites in `cmd_tick` (implement ~L626, verify ~L638, orchestrate ~L669) to pass `loop_path=loop_path` and `tick_num=tick_num` where `tick_num = state.get('iteration_count', 0) + 1`.
### R2 -- `prompts/loop-implement.md`
Created the Implement role prompt with:
- `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}` tokens
- ALLOWED/FORBIDDEN sections (no `--transition`, no `--approve`, no file edits outside cwd)
- Instructions to read SPEC.md, implement code, run py_compile and pytest
### R3 -- `prompts/loop-verifier.md`
Created the Verify role prompt with:
- `{artifact_content}`, `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}` tokens
- Strict JSON output format: `{"pass": bool, "score": float, "reasons": [...], "next_hint": "..."}`
- Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress
- Empty artifact handling: returns `{"pass": false, "score": 0.0, ...}`
### R4 -- `prompts/loop-orchestrate.md`
Created the Orchestrate role prompt with:
- `{verdict}`, `{current_task}`, `{current_phase}` tokens
- Phase transition logic (implement -> code_review -> ... -> complete)
- FORBIDDEN: no file edits, no `--approve --loop` (human-only, D4), no auto-approve
### R5 -- `templates/loops/ci-triage/loop.json`
Updated `roles` from `null` values to prompt refs:
```json
"roles": {
"implement": {"prompt": "loop-implement.md"},
"verify": {"prompt": "loop-verifier.md"},
"orchestrate": {"prompt": "loop-orchestrate.md"}
}
```
### R6 -- `templates/loops/self-improvement/loop.json`
Created new template with:
- `work_source`: `{"kind": "audit", "project": "~/.automaton/"}`
- `blast_radius`: `{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}`
- `brakes`: `{"max_iterations": 10, "score_plateau_window": 3}`
- Same role prompt refs as ci-triage
### R7 -- Onboarding documentation
Updated `README.md` with a "Loop Engineering" section covering quick start, tick cycle, configuration, monitoring, and halt/resume.
### R8 -- `tests/test_loop_templates.py`
18 tests covering R1-R6:
- `TestResolvePrompt` (4 tests): token substitution, artifact content reading, fallback, search order
- `TestPromptFiles` (7 tests): prompt file content validation (tokens, JSON instructions, score rubric, FORBIDDEN sections)
- `TestCiTriageTemplate` (1 test): template has prompt refs
- `TestSelfImprovementTemplate` (5 tests): template exists, audit work source, file scope, prompt refs, brakes
- `TestTickPromptSubstitution` (1 test): end-to-end tick with prompt substitution
### R9 -- Doc updates
- `CHANGELOG.md`: added task 6 entry under `[unreleased]`
- `design/loops/technical.md` section 8: documented prompt resolution and substitution
### Test infrastructure updates
Updated `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py` to use non-existent prompt refs (`test-impl.md`, `test-verify.md`, `test-orch.md`) instead of real prompt file names. This prevents `_resolve_prompt` from activating in those tests, preserving backward compat behavior.
Updated `tests/test_framework_self_consistency.py` to exclude loop role prompts from the stop-condition check (they are role prompts, not delivery prompts).
## Verification
- `python3 -m py_compile scripts/loop-runner.py` -- OK
- `python3 -m pytest tests/test_loop_templates.py -v` -- 18 passed
- `python3 -m pytest tests/ -q` -- 393 passed (369 existing + 18 new + 6 from self-consistency recount)
- `bash -n scripts/*.sh` -- OK (no shell changes)
+102
View File
@@ -0,0 +1,102 @@
# SPEC: add-loop-templates-onboarding
## Context
Tasks 2-5 shipped the brakes layer, runner, goal-mode work sources, and worktree creation. But the loop templates have `roles: {implement: null, verify: null, orchestrate: null}` -- no prompt references. And no loop prompt files exist in `prompts/`. This task creates the three loop role prompts, fills in both templates, and adds the critical missing piece: **prompt-file token substitution** in the runner so that `{task_brief}`, `{acceptance_criteria}`, etc. are resolved in the prompt content before the harness sees it.
## Non-Goals (deferred)
- `tier` budget enforcement in the runner -> v1.1 (the `tier` field in role config is documented but not enforced; the 16k context floor is the only hard gate).
- `harness.prompt_var` / `cwd_var` / `output_var` -> v1.1 (the runner uses fixed token names; these config fields are documentation-only).
- Prompt tuning / iteration -> ongoing (the prompts are v1 starters; real tuning happens when the self-improvement loop runs).
- Onboarding wizard / interactive setup -> v1.1 (v1 ships docs only).
## Requirements
### R1 -- Prompt-file token substitution in `loop-runner.py`
- New function `_resolve_prompt(prompt_ref, extras, loop_path) -> str` that:
1. Resolves `prompt_ref` (e.g. `"loop-implement.md"`) to a full path: check `<loop_path>/<prompt_ref>` first, then `~/.automaton/prompts/<prompt_ref>`. If neither exists, return `prompt_ref` as-is (let the harness handle it).
2. Reads the prompt file content.
3. Substitutes content-level tokens in the prompt text: `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`.
4. `{artifact_content}` is special: it reads the file at `extras["artifact"]` (the implement output path) and substitutes its content. If the file doesn't exist, substitutes empty string.
5. Writes the substituted content to a temp file in `<loop_path>/outputs/` (e.g. `outputs/tickN-<role>-prompt.md`).
6. Returns the temp file path.
- `_invoke_harness` is modified to call `_resolve_prompt` on the `prompt_path` before building the command. The returned temp file path replaces `{prompt}` in the command template.
- If the prompt file doesn't exist (prompt_ref is None or file not found), the runner passes the raw `prompt_ref` as `{prompt}` (same as today -- backward compat).
- **Tests:** `test_resolve_prompt_substitutes_tokens`, `test_resolve_prompt_reads_artifact_content`, `test_resolve_prompt_fallback_when_file_missing`, `test_resolve_prompt_searches_loop_dir_then_framework`.
### R2 -- `prompts/loop-implement.md`
- The Implement role prompt. Instructs the LLM to:
- Read the task brief (`{task_brief}`), acceptance criteria (`{acceptance_criteria}`), and the previous tick's hint (`{next_hint}`).
- Implement changes in the current working directory (`{cwd}`).
- Write the artifact/implementation per the task's SPEC.
- The current task is `{current_task}` in phase `{current_phase}`.
- Follows the framework's prompt conventions (ALLOWED/FORBIDDEN sections, no auto-approve, status.py for transitions).
- **Tests:** `test_loop_implement_prompt_has_tokens`, `test_loop_implement_prompt_has_forbidden_section`.
### R3 -- `prompts/loop-verifier.md`
- The Verify role prompt. Based on `technical.md` section 5. Instructs the LLM to:
- Grade the artifact at `{artifact_content}` against `{acceptance_criteria}`.
- Consider `{task_brief}` and `{next_hint}`.
- Output strict JSON: `{"pass": bool, "score": 0.0-1.0, "reasons": [...], "next_hint": "..."}`.
- Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress.
- **Tests:** `test_loop_verifier_prompt_has_json_instruction`, `test_loop_verifier_prompt_has_score_rubric`, `test_loop_verifier_prompt_has_tokens`.
### R4 -- `prompts/loop-orchestrate.md`
- The Orchestrate role prompt. Instructs the LLM to:
- Read the verdict (`{verdict}`).
- Call exactly one `status.py` operation: `--transition` (if pass=true and task not complete), `--approve` (if in an approval-gated phase), or escalate to `human_intervention` (if pass=false or score is low).
- No file edits. No auto-approve (D4).
- The current task is `{current_task}` in phase `{current_phase}`.
- **Tests:** `test_loop_orchestrate_prompt_has_verdict_token`, `test_loop_orchestrate_prompt_has_no_edit_rule`.
### R5 -- Update `templates/loops/ci-triage/loop.json`
- Fill in `roles` with prompt references:
```json
"roles": {
"implement": {"prompt": "loop-implement.md"},
"verify": {"prompt": "loop-verifier.md"},
"orchestrate": {"prompt": "loop-orchestrate.md"}
}
```
- Keep all other fields unchanged.
- **Tests:** `test_ci_triage_template_has_prompt_refs`.
### R6 -- Create `templates/loops/self-improvement/loop.json`
- Per `technical.md` section 9. Key fields:
- `name`: `"self-improvement"`
- `work_source`: `{"kind": "audit", "project": "~/.automaton/"}`
- `roles`: same prompt refs as ci-triage
- `brakes`: `max_iterations: 10, score_plateau_window: 3`
- `blast_radius`: `{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}`
- `acceptance_criteria`: from technical.md section 9
- `schedule`: `{"interval_seconds": 3600}`
- Use `"use_worktree"` (not `"worktree"`) for consistency with the code.
- **Tests:** `test_self_improvement_template_exists`, `test_self_improvement_template_has_audit_work_source`, `test_self_improvement_template_has_file_scope`.
### R7 -- Onboarding documentation
- Add a "Loop Engineering" section to `README.md` (or update existing) with:
- Quick start: `status.py --create-loop <name> --from-template ci-triage` -> `--install-schedule <name>`
- How loops work: one-tick cycle diagram (gate -> find work -> worktree -> implement -> verify -> orchestrate -> state write)
- How to configure: `loop.json` fields reference
- How to monitor: `--loop-list`, `--audit`, `.state.log`
- How to halt/resume: `--approve --loop`, `--pause-loop`, `--resume-loop`
- **Tests:** none (doc-only).
### R8 -- New test file `tests/test_loop_templates.py`
- Covers R1-R6 as itemized above; target 12-16 tests.
- Prompt-file substitution tests use `tmp_path` to create fake prompt files and verify the temp file output.
- Template tests read the actual template files from `templates/loops/`.
- **Tests:** self-referential.
### R9 -- CHANGELOG and doc updates
- `CHANGELOG.md` under `[unreleased]`.
- `design/loops/technical.md` section 8: note that the runner now resolves and substitutes prompt files.
- **Tests:** none (doc-only).
## Verification
- `python3 -m py_compile scripts/loop-runner.py`
- `python3 -m pytest tests/test_loop_templates.py -v`
- `python3 -m pytest tests/ -q` -- full suite must remain green; expected total approx 385 (369 + 12-16 new).
- `bash -n scripts/*.sh` (no shell changes; safety check).
@@ -0,0 +1,42 @@
# VERDICT: add-loop-templates-onboarding
## Task
Implement prompt-file token substitution in the loop runner, create three loop role prompts (`loop-implement.md`, `loop-verifier.md`, `loop-orchestrate.md`), fill in both loop templates, create the self-improvement template, and add onboarding documentation.
## Deliverables Review
| Requirement | Status | Evidence |
|---|---|---|
| R1: `_resolve_prompt` with token substitution | DONE | `scripts/loop-runner.py:248-303`, 4 tests in `TestResolvePrompt` |
| R2: `prompts/loop-implement.md` | DONE | File created, 2 tests in `TestPromptFiles` |
| R3: `prompts/loop-verifier.md` | DONE | File created, 3 tests in `TestPromptFiles` |
| R4: `prompts/loop-orchestrate.md` | DONE | File created, 2 tests in `TestPromptFiles` |
| R5: ci-triage template roles filled | DONE | `templates/loops/ci-triage/loop.json`, 1 test in `TestCiTriageTemplate` |
| R6: self-improvement template created | DONE | `templates/loops/self-improvement/loop.json`, 5 tests in `TestSelfImprovementTemplate` |
| R7: README onboarding section | DONE | `README.md` "Loop Engineering" section with Quick Start, Tick Cycle, Configuration, Monitoring, Halt/Resume |
| R8: `tests/test_loop_templates.py` | DONE | 18 tests (target was 12-16; exceeded) |
| R9: CHANGELOG and technical.md | DONE | `CHANGELOG.md` task 6 entry, `design/loops/technical.md` section 8 updated |
## Quality Assessment
- **Test coverage:** 18 new tests, all passing. Full suite 393 passed (was 369). No regressions.
- **Backward compatibility:** `_invoke_harness` new params are optional. Existing tests updated to use non-existent prompt refs so `_resolve_prompt` fallback path is exercised. No breaking changes.
- **Code quality:** `_resolve_prompt` is clean, well-structured, handles all edge cases (missing file, missing artifact, empty prompt_ref, missing outputs dir). Follows existing code conventions.
- **Documentation:** README onboarding section is comprehensive. technical.md section 8 documents the prompt resolution flow. CHANGELOG is detailed.
- **Security:** Adversarial review found no exploitable vulnerabilities. Path traversal is mitigated by trusted input. Token injection is not possible (single-pass substitution). Large artifact DoS is mitigated by context floor gate.
## Pipeline Artifacts
- SPEC.md -- written and approved
- IMPLEMENTATION.md -- written
- CODE_REVIEW.md -- written, approved
- BUG_REPORT.md -- written (CLEAN, 2 LOW + 1 INFO)
- ADVERSARIAL_BUG_REPORT.md -- written (CLEAN, no exploitable vulnerabilities)
- DOC_REVIEW.md -- written (APPROVED)
## Verdict
**APPROVED -- ready for complete.**
All 9 requirements (R1-R9) are fully implemented, tested, and documented. The task delivers the critical missing piece of loop engineering v1: prompt-file token substitution that closes the feedback loop between ticks. The three loop role prompts provide the LLM instructions for the Implement/Verify/Orchestrate cycle. The self-improvement template enables the framework to improve itself via audit-driven loops.