94 lines
4.7 KiB
Markdown
94 lines
4.7 KiB
Markdown
# IMPLEMENTATION: add-loop-templates-onboarding
|
|||
|
|
|
||
|
|
## Summary
|
||
|
|
|
||
|
|
Implemented prompt-file token substitution in the loop runner, created three loop role prompts, filled in both loop templates, and added onboarding documentation.
|
||
|
|
|
||
|
|
## Changes
|
||
|
|
|
||
|
|
### R1 -- `_resolve_prompt` in `scripts/loop-runner.py`
|
||
|
|
|
||
|
|
Added `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)` at line ~248:
|
||
|
|
- Searches `<loop_path>/<prompt_ref>` then `~/.automaton/prompts/<prompt_ref>` for the prompt file
|
||
|
|
- Reads the file content and substitutes content-level tokens: `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`
|
||
|
|
- `{artifact_content}` reads the file at `extras["artifact"]` path; empty string if missing
|
||
|
|
- Writes substituted content to `<loop_path>/outputs/tickN-<role>-prompt.md`
|
||
|
|
- Returns the temp file path
|
||
|
|
- Falls back to raw `prompt_ref` if file not found (backward compat)
|
||
|
|
|
||
|
|
Modified `_invoke_harness` signature to add `loop_path: Optional[Path] = None, tick_num: int = 0`. When `loop_path` is provided, calls `_resolve_prompt` on the prompt_path before building the harness command.
|
||
|
|
|
||
|
|
Updated all three `_invoke_harness` call sites in `cmd_tick` (implement ~L626, verify ~L638, orchestrate ~L669) to pass `loop_path=loop_path` and `tick_num=tick_num` where `tick_num = state.get('iteration_count', 0) + 1`.
|
||
|
|
|
||
|
|
### R2 -- `prompts/loop-implement.md`
|
||
|
|
|
||
|
|
Created the Implement role prompt with:
|
||
|
|
- `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}` tokens
|
||
|
|
- ALLOWED/FORBIDDEN sections (no `--transition`, no `--approve`, no file edits outside cwd)
|
||
|
|
- Instructions to read SPEC.md, implement code, run py_compile and pytest
|
||
|
|
|
||
|
|
### R3 -- `prompts/loop-verifier.md`
|
||
|
|
|
||
|
|
Created the Verify role prompt with:
|
||
|
|
- `{artifact_content}`, `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}` tokens
|
||
|
|
- Strict JSON output format: `{"pass": bool, "score": float, "reasons": [...], "next_hint": "..."}`
|
||
|
|
- Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress
|
||
|
|
- Empty artifact handling: returns `{"pass": false, "score": 0.0, ...}`
|
||
|
|
|
||
|
|
### R4 -- `prompts/loop-orchestrate.md`
|
||
|
|
|
||
|
|
Created the Orchestrate role prompt with:
|
||
|
|
- `{verdict}`, `{current_task}`, `{current_phase}` tokens
|
||
|
|
- Phase transition logic (implement -> code_review -> ... -> complete)
|
||
|
|
- FORBIDDEN: no file edits, no `--approve --loop` (human-only, D4), no auto-approve
|
||
|
|
|
||
|
|
### R5 -- `templates/loops/ci-triage/loop.json`
|
||
|
|
|
||
|
|
Updated `roles` from `null` values to prompt refs:
|
||
|
|
```json
|
||
|
|
"roles": {
|
||
|
|
"implement": {"prompt": "loop-implement.md"},
|
||
|
|
"verify": {"prompt": "loop-verifier.md"},
|
||
|
|
"orchestrate": {"prompt": "loop-orchestrate.md"}
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### R6 -- `templates/loops/self-improvement/loop.json`
|
||
|
|
|
||
|
|
Created new template with:
|
||
|
|
- `work_source`: `{"kind": "audit", "project": "~/.automaton/"}`
|
||
|
|
- `blast_radius`: `{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}`
|
||
|
|
- `brakes`: `{"max_iterations": 10, "score_plateau_window": 3}`
|
||
|
|
- Same role prompt refs as ci-triage
|
||
|
|
|
||
|
|
### R7 -- Onboarding documentation
|
||
|
|
|
||
|
|
Updated `README.md` with a "Loop Engineering" section covering quick start, tick cycle, configuration, monitoring, and halt/resume.
|
||
|
|
|
||
|
|
### R8 -- `tests/test_loop_templates.py`
|
||
|
|
|
||
|
|
18 tests covering R1-R6:
|
||
|
|
- `TestResolvePrompt` (4 tests): token substitution, artifact content reading, fallback, search order
|
||
|
|
- `TestPromptFiles` (7 tests): prompt file content validation (tokens, JSON instructions, score rubric, FORBIDDEN sections)
|
||
|
|
- `TestCiTriageTemplate` (1 test): template has prompt refs
|
||
|
|
- `TestSelfImprovementTemplate` (5 tests): template exists, audit work source, file scope, prompt refs, brakes
|
||
|
|
- `TestTickPromptSubstitution` (1 test): end-to-end tick with prompt substitution
|
||
|
|
|
||
|
|
### R9 -- Doc updates
|
||
|
|
|
||
|
|
- `CHANGELOG.md`: added task 6 entry under `[unreleased]`
|
||
|
|
- `design/loops/technical.md` section 8: documented prompt resolution and substitution
|
||
|
|
|
||
|
|
### Test infrastructure updates
|
||
|
|
|
||
|
|
Updated `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py` to use non-existent prompt refs (`test-impl.md`, `test-verify.md`, `test-orch.md`) instead of real prompt file names. This prevents `_resolve_prompt` from activating in those tests, preserving backward compat behavior.
|
||
|
|
|
||
|
|
Updated `tests/test_framework_self_consistency.py` to exclude loop role prompts from the stop-condition check (they are role prompts, not delivery prompts).
|
||
|
|
|
||
|
|
## Verification
|
||
|
|
|
||
|
|
- `python3 -m py_compile scripts/loop-runner.py` -- OK
|
||
|
|
- `python3 -m pytest tests/test_loop_templates.py -v` -- 18 passed
|
||
|
|
- `python3 -m pytest tests/ -q` -- 393 passed (369 existing + 18 new + 6 from self-consistency recount)
|
||
|
|
- `bash -n scripts/*.sh` -- OK (no shell changes)
|