CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
8.1 KiB
8.1 KiB
SPEC: add-goal-mode
Context
Task 3 (add-loop-runner) shipped the graded JSON parser, score_history cap, and the score circuit-breaker gate. The "verifier session, graded JSON, score circuit-breaker" framing from the v1 README is therefore already delivered. Task 4 closes the goal-oriented loop on the runner side: gives the runner real work sources beyond current_task, feeds the verifier acceptance criteria + a prior-tick hint, and closes the next_hint feedback path into the next tick's Implement/Verify sessions.
Non-Goals (deferred)
--goalCLI flag → v1.1 (R9 from research; adds CLI surface without serving any v1 design doc requirement).loop-verifier.md/loop-implement.md/loop-orchestrate.mdprompt text → task 6 (this task only wires the substitution tokens; the prompts that consume them land in task 6).backlogintegration with thedesign/context-sizing/workstream → task 7.parse_verdictscore clamp +passstring coercion → v1.1 hardening (already tracked in task-3 BUG_REPORT).outputs.retentioninloop.json→ v1.1.--create-taskauto-creation from audit violations beyond minimal name resolution → v1.1 hardening.
Requirements
R1 -- find_work work_source dispatch
- Replace the inline
single-only block incmd_tickwith a_find_work(state, cfg, project_dir)helper that dispatches oncfg.get("work_source", {}).get("kind", "single"). - Missing
work_sourcefield or missingkind→ fall back to"single"with a.state.logWARNING line (preserves backward compat with the currentci-triage/loop.jsontemplate, which has nowork_sourcefield). singlewith nullcurrent_task→ SKIPno_current_task(unchanged from task 3).- All kinds write the resolved task name into
state["current_task"]before returning so downstream steps see it. - Unknown
kind→ WARNING + fallback to"single". - Tests:
test_find_work_single,test_find_work_missing_work_source_falls_back_to_single,test_find_work_unknown_kind_warns_and_falls_back.
R2 -- audit work_source
work_source.kind == "audit"→ callstatus.py --audit --json --project <p>via_run_json. Parse the violations list. Pick the highest-severity unresolved violation (severity ordering: high > med > low). Use the violation'staskfield ascurrent_taskwhen present. If the violation has no associated task, callstatus.py --create-task <slug>(slug derived from the violation message) and set the new task ascurrent_task. If no unresolved violations → SKIPno_work(new skip reason; CLEAN scheduler exit; does not incrementiteration_count).work_source.project(optional) overrides the project path passed to--audit; defaults to the loop's own project.- Tests:
test_audit_picks_highest_severity_violation,test_audit_creates_task_when_violation_has_no_task,test_audit_skip_when_no_violations,test_audit_uses_work_source_project.
R3 -- backlog work_source
work_source.kind == "backlog"→ read<framework>/design/<area>/BACKLOG.mdwhereareacomes fromwork_source.area(default"loops"). Parse the topmost- [ ]checkbox line. Map to a task name by slugifying the item's bold heading (e.g.**design-update-loop-template**→design-update-loop-template). Set ascurrent_task. If no[ ]items remain → SKIPno_work.work_source.areaoverrides the area path underdesign/.- Tests:
test_backlog_picks_top_unchecked_item,test_backlog_skip_when_empty,test_backlog_uses_area_path.
R4 -- Verifier-prompt token plumbing
- Extend the substitution map in
_invoke_harness()/_substitute()to recognize three new tokens (in addition to the existing seven:{prompt},{cwd},{output},{artifact},{verdict},{current_task},{current_phase}):{task_brief}-- read from<task_dir>/RESEARCH.mdif present, else<task_dir>/DESIGN.md, else<task_dir>/SPEC.md, else empty string. Capped at 4k tokens via R5.{acceptance_criteria}-- read fromloop.jsonacceptance_criteria(string OR list; list joined with newlines). Capped at 2k tokens.{next_hint}-- read fromstate.get("last_verdict", {}).get("next_hint", "")(empty on first tick or after--approve). Capped at 1k tokens.
- Tokens absent from the prompt stay literal (same rule as today -- a prompt that omits
{task_brief}is unaffected). - Tests:
test_task_brief_substituted_from_research,test_acceptance_criteria_substituted_from_loop_json_list,test_next_hint_substituted_from_last_verdict,test_missing_tokens_leave_prompt_intact.
R5 -- _truncate_tokens(text, max_tokens) helper
- Stdlib-only approximate token cap. No tokenizer dependency. Heuristic:
max_tokens * 4chars (4-chars-per-token approximation). When the input exceeds the char budget, truncate and append a trailing…[truncated]marker. Used fortask_brief(4000),acceptance_criteria(2000),next_hint(1000). - Inputs at or under the cap are returned unchanged.
- Tests:
test_truncate_short_text_unchanged,test_truncate_long_text_capped_with_marker,test_truncate_returns_empty_for_empty_input.
R6 -- next_hint feedback loop closure
- The Implement and Verify harness invocations receive
{next_hint}fromstate["last_verdict"]["next_hint"]via R4. This closes the loop: tick N's verifier hint becomes tick N+1's Implement context. - A tick with no prior verdict (first tick, or after
--approvecleared state) passes an empty{next_hint}string (no KeyError, no spurious substitution). last_verdictis cleared on--approve --loop(already happens today via the resume path -- verify and assert in tests).- Tests:
test_next_hint_fed_into_next_tick_implement,test_first_tick_has_empty_next_hint.
R7 -- loop.json schema additions
- Document
work_sourceandacceptance_criteriafields indesign/loops/technical.mdschema section (§2) and the self-improvement example (§9). - Update
templates/loops/ci-triage/loop.jsonto include:"work_source": {"kind": "single"}(explicit; current template omits the field entirely)."acceptance_criteria": ["All R-numbers from SPEC.md are implemented.", "Tests pass with no regressions.", "Pipeline driven to complete."](self-documenting placeholder;nullroles stay -- task 6 fills them with prompt text).
--create-loopdoes NOT strictly validatework_sourceshape; missingwork_sourcecontinues to fall back to"single"(R1).acceptance_criteriais an optional free-form field (string OR list of strings).- Tests:
test_ci_triage_template_has_work_source,test_ci_triage_template_has_acceptance_criteria,test_create_loop_preserves_acceptance_criteria.
R8 -- status.py --audit --json mode
- Add
--jsonsupport tocmd_audit. When--jsonis set, emit a single JSON line on stdout (machine-readable, pickable by_run_json):{"violations": [...], "loops": [...], "total_tasks": <int>, "untracked_tasks": <int>}- Each violation:
{"category": <int 1-6>, "severity": "high"|"med"|"low", "task": <str|null>, "message": <str>, "resolved": false}.
- Existing human-readable
--auditoutput (no--json) is unchanged. - This is the data source the
auditwork consumes (R2). - Tests:
test_audit_json_emits_violations_array,test_audit_json_includes_loops_block,test_audit_json_pickable_by_runner_run_json.
R9 -- New test file tests/test_goal_mode.py
- Mirrors
test_loop_runner.py's stubbing pattern (monkeypatch.setattr(subprocess, "run", fake_run)) andtest_status_brakes.py's--audit --jsonassertions. - Covers R1-R8 as itemized above; target 12-16 tests.
- Add one regression test:
test_existing_single_work_source_loop_ticks_unchanged-- a loop withcurrent_taskset and nowork_sourcefield still ticks exactly as before (backward compat with all task-3 fixtures). - All subprocess calls stubbed; no live LLM in CI.
Verification
python3 -m py_compile scripts/loop-runner.py scripts/status.pypython3 -m pytest tests/test_goal_mode.py -vpython3 -m pytest tests/ -q-- full suite must remain green; expected total ≈ 340 (328 + 12-16 new).bash -n scripts/*.sh(no shell changes; safety check).