CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
6.7 KiB
6.7 KiB
SPEC: add-loop-templates-onboarding
Context
Tasks 2-5 shipped the brakes layer, runner, goal-mode work sources, and worktree creation. But the loop templates have roles: {implement: null, verify: null, orchestrate: null} -- no prompt references. And no loop prompt files exist in prompts/. This task creates the three loop role prompts, fills in both templates, and adds the critical missing piece: prompt-file token substitution in the runner so that {task_brief}, {acceptance_criteria}, etc. are resolved in the prompt content before the harness sees it.
Non-Goals (deferred)
tierbudget enforcement in the runner -> v1.1 (thetierfield in role config is documented but not enforced; the 16k context floor is the only hard gate).harness.prompt_var/cwd_var/output_var-> v1.1 (the runner uses fixed token names; these config fields are documentation-only).- Prompt tuning / iteration -> ongoing (the prompts are v1 starters; real tuning happens when the self-improvement loop runs).
- Onboarding wizard / interactive setup -> v1.1 (v1 ships docs only).
Requirements
R1 -- Prompt-file token substitution in loop-runner.py
- New function
_resolve_prompt(prompt_ref, extras, loop_path) -> strthat:- Resolves
prompt_ref(e.g."loop-implement.md") to a full path: check<loop_path>/<prompt_ref>first, then~/.automaton/prompts/<prompt_ref>. If neither exists, returnprompt_refas-is (let the harness handle it). - Reads the prompt file content.
- Substitutes content-level tokens in the prompt text:
{task_brief},{acceptance_criteria},{next_hint},{current_task},{current_phase},{verdict},{artifact_content}. {artifact_content}is special: it reads the file atextras["artifact"](the implement output path) and substitutes its content. If the file doesn't exist, substitutes empty string.- Writes the substituted content to a temp file in
<loop_path>/outputs/(e.g.outputs/tickN-<role>-prompt.md). - Returns the temp file path.
- Resolves
_invoke_harnessis modified to call_resolve_prompton theprompt_pathbefore building the command. The returned temp file path replaces{prompt}in the command template.- If the prompt file doesn't exist (prompt_ref is None or file not found), the runner passes the raw
prompt_refas{prompt}(same as today -- backward compat). - Tests:
test_resolve_prompt_substitutes_tokens,test_resolve_prompt_reads_artifact_content,test_resolve_prompt_fallback_when_file_missing,test_resolve_prompt_searches_loop_dir_then_framework.
R2 -- prompts/loop-implement.md
- The Implement role prompt. Instructs the LLM to:
- Read the task brief (
{task_brief}), acceptance criteria ({acceptance_criteria}), and the previous tick's hint ({next_hint}). - Implement changes in the current working directory (
{cwd}). - Write the artifact/implementation per the task's SPEC.
- The current task is
{current_task}in phase{current_phase}.
- Read the task brief (
- Follows the framework's prompt conventions (ALLOWED/FORBIDDEN sections, no auto-approve, status.py for transitions).
- Tests:
test_loop_implement_prompt_has_tokens,test_loop_implement_prompt_has_forbidden_section.
R3 -- prompts/loop-verifier.md
- The Verify role prompt. Based on
technical.mdsection 5. Instructs the LLM to:- Grade the artifact at
{artifact_content}against{acceptance_criteria}. - Consider
{task_brief}and{next_hint}. - Output strict JSON:
{"pass": bool, "score": 0.0-1.0, "reasons": [...], "next_hint": "..."}. - Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress.
- Grade the artifact at
- Tests:
test_loop_verifier_prompt_has_json_instruction,test_loop_verifier_prompt_has_score_rubric,test_loop_verifier_prompt_has_tokens.
R4 -- prompts/loop-orchestrate.md
- The Orchestrate role prompt. Instructs the LLM to:
- Read the verdict (
{verdict}). - Call exactly one
status.pyoperation:--transition(if pass=true and task not complete),--approve(if in an approval-gated phase), or escalate tohuman_intervention(if pass=false or score is low). - No file edits. No auto-approve (D4).
- The current task is
{current_task}in phase{current_phase}.
- Read the verdict (
- Tests:
test_loop_orchestrate_prompt_has_verdict_token,test_loop_orchestrate_prompt_has_no_edit_rule.
R5 -- Update templates/loops/ci-triage/loop.json
- Fill in
roleswith prompt references:"roles": { "implement": {"prompt": "loop-implement.md"}, "verify": {"prompt": "loop-verifier.md"}, "orchestrate": {"prompt": "loop-orchestrate.md"} } - Keep all other fields unchanged.
- Tests:
test_ci_triage_template_has_prompt_refs.
R6 -- Create templates/loops/self-improvement/loop.json
- Per
technical.mdsection 9. Key fields:name:"self-improvement"work_source:{"kind": "audit", "project": "~/.automaton/"}roles: same prompt refs as ci-triagebrakes:max_iterations: 10, score_plateau_window: 3blast_radius:{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}acceptance_criteria: from technical.md section 9schedule:{"interval_seconds": 3600}
- Use
"use_worktree"(not"worktree") for consistency with the code. - Tests:
test_self_improvement_template_exists,test_self_improvement_template_has_audit_work_source,test_self_improvement_template_has_file_scope.
R7 -- Onboarding documentation
- Add a "Loop Engineering" section to
README.md(or update existing) with:- Quick start:
status.py --create-loop <name> --from-template ci-triage->--install-schedule <name> - How loops work: one-tick cycle diagram (gate -> find work -> worktree -> implement -> verify -> orchestrate -> state write)
- How to configure:
loop.jsonfields reference - How to monitor:
--loop-list,--audit,.state.log - How to halt/resume:
--approve --loop,--pause-loop,--resume-loop
- Quick start:
- Tests: none (doc-only).
R8 -- New test file tests/test_loop_templates.py
- Covers R1-R6 as itemized above; target 12-16 tests.
- Prompt-file substitution tests use
tmp_pathto create fake prompt files and verify the temp file output. - Template tests read the actual template files from
templates/loops/. - Tests: self-referential.
R9 -- CHANGELOG and doc updates
CHANGELOG.mdunder[unreleased].design/loops/technical.mdsection 8: note that the runner now resolves and substitutes prompt files.- Tests: none (doc-only).
Verification
python3 -m py_compile scripts/loop-runner.pypython3 -m pytest tests/test_loop_templates.py -vpython3 -m pytest tests/ -q-- full suite must remain green; expected total approx 385 (369 + 12-16 new).bash -n scripts/*.sh(no shell changes; safety check).