CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
4.3 KiB
4.3 KiB
Implementation: add-goal-mode
Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in scripts/loop-runner.py, scripts/status.py, templates/loops/ci-triage/loop.json, design/loops/technical.md, and tests/test_goal_mode.py.
Files changed
scripts/loop-runner.py--_find_workdispatch,_find_work_audit,_find_work_backlog,_truncate_tokens,_read_task_brief,_acceptance_criteria_text,_next_hint_text, new substitution tokens incmd_tick.scripts/status.py----audit --jsonmode incmd_audit.templates/loops/ci-triage/loop.json-- explicitwork_sourceandacceptance_criteriafields.design/loops/technical.md-- schema section updated withwork_sourceandacceptance_criteria.tests/test_goal_mode.py-- 26 tests covering R1-R8 + regression.
R-by-R coverage
| Req | Code |
|---|---|
| R1 find_work dispatch | _find_work(state, cfg, loop_path, project_dir) dispatches on cfg["work_source"]["kind"]; missing/unknown falls back to "single" with WARNING log |
| R2 audit work_source | _find_work_audit calls status.py --audit --json, sorts by severity (high>med>low), uses violation task or creates one via --create-task |
| R3 backlog work_source | _find_work_backlog reads design/<area>/BACKLOG.md, picks topmost - [ ] line, slugifies the **bold** heading |
| R4 verifier tokens | {task_brief}, {acceptance_criteria}, {next_hint} added to extras dict in cmd_tick implement/verify invocations; substituted via _substitute |
| R5 truncate_tokens | _truncate_tokens(text, max_tokens) -- 4 chars/token heuristic, appends ...[truncated] marker; task_brief=4000, acceptance=2000, next_hint=1000 |
| R6 next_hint loop | _next_hint_text(state) reads state["last_verdict"]["next_hint"]; empty on first tick / after approve; fed into both implement and verify |
| R7 loop.json schema | ci-triage template updated; technical.md schema section updated |
| R8 audit --json | cmd_audit in status.py: when --json, emits {"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N} as single JSON line |
Key design decisions
_find_workreturns(task, skip_reason)tuple;skip_reasonisNonewhen work found,"no_current_task"for single-with-null,"no_work"for audit/backlog with no items._find_work_auditcreates tasks viastatus.py --create-taskwhen a violation has no associated task; slug derived from_slugify(message)._find_work_backlogmaps**bold-name**in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.- Token substitution only applies when the harness command template contains the placeholder; prompts that omit
{task_brief}etc. are unaffected. --audit --jsonoutput is a single JSON line on stdout, parseable by_run_json(which takes the last line).
Tests (tests/test_goal_mode.py)
26 tests across 8 classes; all subprocess.run calls stubbed via monkeypatch.
TestFindWorkDispatch(3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.TestAuditWorkSource(4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.TestBacklogWorkSource(3): picks top unchecked item; skips when empty; uses area path.TestVerifierTokens(4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.TestTruncateTokens(3): short text unchanged; long text capped with marker; empty returns empty.TestNextHintFeedback(2): hint fed into next tick; first tick has empty hint.TestLoopJsonSchemaAdditions(3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.TestAuditJson(3): emits violations array; includes loops block; pickable by runner _run_json.TestRegressionBackwardCompat(1): existing single loop with no work_source/acceptance_criteria ticks unchanged.
Verification
python3 -m py_compile scripts/loop-runner.py scripts/status.py-- PASSpython3 -m pytest tests/test_goal_mode.py -v-- 26 passedpython3 -m pytest tests/ -q-- 354 passed (328 baseline + 26 new)bash -n scripts/*.sh-- no shell changes