Files
automaton/tasks/complete/add-goal-mode/IMPLEMENTATION.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

4.3 KiB

Implementation: add-goal-mode

Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in scripts/loop-runner.py, scripts/status.py, templates/loops/ci-triage/loop.json, design/loops/technical.md, and tests/test_goal_mode.py.

Files changed

  • scripts/loop-runner.py -- _find_work dispatch, _find_work_audit, _find_work_backlog, _truncate_tokens, _read_task_brief, _acceptance_criteria_text, _next_hint_text, new substitution tokens in cmd_tick.
  • scripts/status.py -- --audit --json mode in cmd_audit.
  • templates/loops/ci-triage/loop.json -- explicit work_source and acceptance_criteria fields.
  • design/loops/technical.md -- schema section updated with work_source and acceptance_criteria.
  • tests/test_goal_mode.py -- 26 tests covering R1-R8 + regression.

R-by-R coverage

Req Code
R1 find_work dispatch _find_work(state, cfg, loop_path, project_dir) dispatches on cfg["work_source"]["kind"]; missing/unknown falls back to "single" with WARNING log
R2 audit work_source _find_work_audit calls status.py --audit --json, sorts by severity (high>med>low), uses violation task or creates one via --create-task
R3 backlog work_source _find_work_backlog reads design/<area>/BACKLOG.md, picks topmost - [ ] line, slugifies the **bold** heading
R4 verifier tokens {task_brief}, {acceptance_criteria}, {next_hint} added to extras dict in cmd_tick implement/verify invocations; substituted via _substitute
R5 truncate_tokens _truncate_tokens(text, max_tokens) -- 4 chars/token heuristic, appends ...[truncated] marker; task_brief=4000, acceptance=2000, next_hint=1000
R6 next_hint loop _next_hint_text(state) reads state["last_verdict"]["next_hint"]; empty on first tick / after approve; fed into both implement and verify
R7 loop.json schema ci-triage template updated; technical.md schema section updated
R8 audit --json cmd_audit in status.py: when --json, emits {"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N} as single JSON line

Key design decisions

  • _find_work returns (task, skip_reason) tuple; skip_reason is None when work found, "no_current_task" for single-with-null, "no_work" for audit/backlog with no items.
  • _find_work_audit creates tasks via status.py --create-task when a violation has no associated task; slug derived from _slugify(message).
  • _find_work_backlog maps **bold-name** in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.
  • Token substitution only applies when the harness command template contains the placeholder; prompts that omit {task_brief} etc. are unaffected.
  • --audit --json output is a single JSON line on stdout, parseable by _run_json (which takes the last line).

Tests (tests/test_goal_mode.py)

26 tests across 8 classes; all subprocess.run calls stubbed via monkeypatch.

  • TestFindWorkDispatch (3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.
  • TestAuditWorkSource (4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.
  • TestBacklogWorkSource (3): picks top unchecked item; skips when empty; uses area path.
  • TestVerifierTokens (4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.
  • TestTruncateTokens (3): short text unchanged; long text capped with marker; empty returns empty.
  • TestNextHintFeedback (2): hint fed into next tick; first tick has empty hint.
  • TestLoopJsonSchemaAdditions (3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.
  • TestAuditJson (3): emits violations array; includes loops block; pickable by runner _run_json.
  • TestRegressionBackwardCompat (1): existing single loop with no work_source/acceptance_criteria ticks unchanged.

Verification

  • python3 -m py_compile scripts/loop-runner.py scripts/status.py -- PASS
  • python3 -m pytest tests/test_goal_mode.py -v -- 26 passed
  • python3 -m pytest tests/ -q -- 354 passed (328 baseline + 26 new)
  • bash -n scripts/*.sh -- no shell changes