CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
79 lines
4.2 KiB
Markdown
79 lines
4.2 KiB
Markdown
# CODE_REVIEW: add-loop-templates-onboarding
|
|
|
|
## Reviewed Files
|
|
|
|
1. `scripts/loop-runner.py` -- `_resolve_prompt` function (lines ~248-303), `_invoke_harness` signature change (lines ~306-338), `cmd_tick` call site updates (lines ~626, ~638, ~669)
|
|
2. `prompts/loop-implement.md` -- new file
|
|
3. `prompts/loop-verifier.md` -- new file
|
|
4. `prompts/loop-orchestrate.md` -- new file
|
|
5. `templates/loops/ci-triage/loop.json` -- roles updated
|
|
6. `templates/loops/self-improvement/loop.json` -- new file
|
|
7. `tests/test_loop_templates.py` -- new test file (18 tests)
|
|
8. `tests/test_loop_runner.py` -- prompt ref renames
|
|
9. `tests/test_blast_radius.py` -- prompt ref renames
|
|
10. `tests/test_goal_mode.py` -- prompt ref renames
|
|
11. `tests/test_framework_self_consistency.py` -- exclusion set update
|
|
12. `README.md` -- Loop Engineering onboarding section
|
|
13. `CHANGELOG.md` -- task 6 entry
|
|
14. `design/loops/technical.md` -- section 8 prompt resolution docs
|
|
|
|
## Findings
|
|
|
|
### 1. `_resolve_prompt` -- token substitution correctness
|
|
|
|
The function correctly handles the two-stage search (loop-local then framework), reads the file, substitutes tokens, writes to outputs/, and returns the temp path. The fallback to raw `prompt_ref` when the file is not found preserves backward compatibility.
|
|
|
|
**Concern: token injection.** The `str(value)` substitution via `content.replace("{" + key + "}", str(value))` is safe for the current token set (all values are controlled: task_brief from SPEC.md, acceptance_criteria from loop.json, etc.). No user-supplied input flows into these tokens without being read from a file first. Acceptable for v1.
|
|
|
|
**Verdict:** PASS
|
|
|
|
### 2. `_invoke_harness` signature change
|
|
|
|
The new `loop_path` and `tick_num` parameters are optional with defaults (`None` and `0`). Existing callers that don't pass them get the old behavior (raw prompt_ref passed through). This is backward compatible.
|
|
|
|
**Verdict:** PASS
|
|
|
|
### 3. `cmd_tick` call site updates
|
|
|
|
All three call sites (implement, verify, orchestrate) now pass `loop_path=loop_path` and `tick_num=tick_num` where `tick_num` is computed once as `state.get('iteration_count', 0) + 1`. This is correct -- the tick number should be consistent across all three role invocations in the same tick.
|
|
|
|
**Verdict:** PASS
|
|
|
|
### 4. Prompt file content
|
|
|
|
- `loop-implement.md`: has all required tokens, ALLOWED/FORBIDDEN sections, no auto-approve. Correct.
|
|
- `loop-verifier.md`: has strict JSON output format, score rubric, artifact_content token. Correct.
|
|
- `loop-orchestrate.md`: has verdict token, phase transition logic, no-edit rule. Correct.
|
|
|
|
All three prompts are excluded from the self-consistency stop-condition check since they are role prompts, not delivery prompts. This is consistent with how `orchestrate.md` is already excluded.
|
|
|
|
**Verdict:** PASS
|
|
|
|
### 5. Template updates
|
|
|
|
- `ci-triage/loop.json`: roles filled with `{"prompt": "loop-implement.md"}` etc. All other fields unchanged. Correct.
|
|
- `self-improvement/loop.json`: has `work_source: audit`, `use_worktree: true`, `file_scope` with 4 paths, `max_iterations: 10`, `score_plateau_window: 3`. Matches technical.md section 9. Correct.
|
|
|
|
**Verdict:** PASS
|
|
|
|
### 6. Test infrastructure updates
|
|
|
|
Renaming prompt refs from `"loop-implement.md"` to `"test-impl.md"` (and similar) in existing tests is the correct approach. These tests don't test prompt resolution -- they test other runner behavior. Using non-existent prompt refs ensures `_resolve_prompt` falls back to the raw string, preserving the old argv contents that the test assertions depend on.
|
|
|
|
**Verdict:** PASS
|
|
|
|
### 7. Edge cases
|
|
|
|
- **Empty prompt_ref**: `_resolve_prompt` returns `prompt_ref or ""` at line 260. Safe.
|
|
- **Missing outputs dir**: `out_dir.mkdir(parents=True, exist_ok=True)` at line 300. Safe.
|
|
- **Missing artifact file for `{artifact_content}`**: caught by `try/except OSError`, returns empty string. Safe.
|
|
- **Loop-local prompt override**: searched first, allows per-loop customization without modifying framework prompts. Good design.
|
|
|
|
**Verdict:** PASS
|
|
|
|
## Summary
|
|
|
|
All 7 review areas pass. The implementation is correct, backward compatible, and well-tested. 18 new tests cover the prompt resolution, prompt file content, template updates, and tick integration. Full suite: 393 passed.
|
|
|
|
**Overall verdict: APPROVED**
|