Files
automaton/tasks/complete/add-goal-mode/SPEC.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

8.1 KiB

SPEC: add-goal-mode

Context

Task 3 (add-loop-runner) shipped the graded JSON parser, score_history cap, and the score circuit-breaker gate. The "verifier session, graded JSON, score circuit-breaker" framing from the v1 README is therefore already delivered. Task 4 closes the goal-oriented loop on the runner side: gives the runner real work sources beyond current_task, feeds the verifier acceptance criteria + a prior-tick hint, and closes the next_hint feedback path into the next tick's Implement/Verify sessions.

Non-Goals (deferred)

  • --goal CLI flag → v1.1 (R9 from research; adds CLI surface without serving any v1 design doc requirement).
  • loop-verifier.md / loop-implement.md / loop-orchestrate.md prompt text → task 6 (this task only wires the substitution tokens; the prompts that consume them land in task 6).
  • backlog integration with the design/context-sizing/ workstream → task 7.
  • parse_verdict score clamp + pass string coercion → v1.1 hardening (already tracked in task-3 BUG_REPORT).
  • outputs.retention in loop.json → v1.1.
  • --create-task auto-creation from audit violations beyond minimal name resolution → v1.1 hardening.

Requirements

R1 -- find_work work_source dispatch

  • Replace the inline single-only block in cmd_tick with a _find_work(state, cfg, project_dir) helper that dispatches on cfg.get("work_source", {}).get("kind", "single").
  • Missing work_source field or missing kind → fall back to "single" with a .state.log WARNING line (preserves backward compat with the current ci-triage/loop.json template, which has no work_source field).
  • single with null current_task → SKIP no_current_task (unchanged from task 3).
  • All kinds write the resolved task name into state["current_task"] before returning so downstream steps see it.
  • Unknown kind → WARNING + fallback to "single".
  • Tests: test_find_work_single, test_find_work_missing_work_source_falls_back_to_single, test_find_work_unknown_kind_warns_and_falls_back.

R2 -- audit work_source

  • work_source.kind == "audit" → call status.py --audit --json --project <p> via _run_json. Parse the violations list. Pick the highest-severity unresolved violation (severity ordering: high > med > low). Use the violation's task field as current_task when present. If the violation has no associated task, call status.py --create-task <slug> (slug derived from the violation message) and set the new task as current_task. If no unresolved violations → SKIP no_work (new skip reason; CLEAN scheduler exit; does not increment iteration_count).
  • work_source.project (optional) overrides the project path passed to --audit; defaults to the loop's own project.
  • Tests: test_audit_picks_highest_severity_violation, test_audit_creates_task_when_violation_has_no_task, test_audit_skip_when_no_violations, test_audit_uses_work_source_project.

R3 -- backlog work_source

  • work_source.kind == "backlog" → read <framework>/design/<area>/BACKLOG.md where area comes from work_source.area (default "loops"). Parse the topmost - [ ] checkbox line. Map to a task name by slugifying the item's bold heading (e.g. **design-update-loop-template** → design-update-loop-template). Set as current_task. If no [ ] items remain → SKIP no_work.
  • work_source.area overrides the area path under design/.
  • Tests: test_backlog_picks_top_unchecked_item, test_backlog_skip_when_empty, test_backlog_uses_area_path.

R4 -- Verifier-prompt token plumbing

  • Extend the substitution map in _invoke_harness() / _substitute() to recognize three new tokens (in addition to the existing seven: {prompt}, {cwd}, {output}, {artifact}, {verdict}, {current_task}, {current_phase}):
    • {task_brief} -- read from <task_dir>/RESEARCH.md if present, else <task_dir>/DESIGN.md, else <task_dir>/SPEC.md, else empty string. Capped at 4k tokens via R5.
    • {acceptance_criteria} -- read from loop.json acceptance_criteria (string OR list; list joined with newlines). Capped at 2k tokens.
    • {next_hint} -- read from state.get("last_verdict", {}).get("next_hint", "") (empty on first tick or after --approve). Capped at 1k tokens.
  • Tokens absent from the prompt stay literal (same rule as today -- a prompt that omits {task_brief} is unaffected).
  • Tests: test_task_brief_substituted_from_research, test_acceptance_criteria_substituted_from_loop_json_list, test_next_hint_substituted_from_last_verdict, test_missing_tokens_leave_prompt_intact.

R5 -- _truncate_tokens(text, max_tokens) helper

  • Stdlib-only approximate token cap. No tokenizer dependency. Heuristic: max_tokens * 4 chars (4-chars-per-token approximation). When the input exceeds the char budget, truncate and append a trailing …[truncated] marker. Used for task_brief (4000), acceptance_criteria (2000), next_hint (1000).
  • Inputs at or under the cap are returned unchanged.
  • Tests: test_truncate_short_text_unchanged, test_truncate_long_text_capped_with_marker, test_truncate_returns_empty_for_empty_input.

R6 -- next_hint feedback loop closure

  • The Implement and Verify harness invocations receive {next_hint} from state["last_verdict"]["next_hint"] via R4. This closes the loop: tick N's verifier hint becomes tick N+1's Implement context.
  • A tick with no prior verdict (first tick, or after --approve cleared state) passes an empty {next_hint} string (no KeyError, no spurious substitution).
  • last_verdict is cleared on --approve --loop (already happens today via the resume path -- verify and assert in tests).
  • Tests: test_next_hint_fed_into_next_tick_implement, test_first_tick_has_empty_next_hint.

R7 -- loop.json schema additions

  • Document work_source and acceptance_criteria fields in design/loops/technical.md schema section (§2) and the self-improvement example (§9).
  • Update templates/loops/ci-triage/loop.json to include:
    • "work_source": {"kind": "single"} (explicit; current template omits the field entirely).
    • "acceptance_criteria": ["All R-numbers from SPEC.md are implemented.", "Tests pass with no regressions.", "Pipeline driven to complete."] (self-documenting placeholder; null roles stay -- task 6 fills them with prompt text).
  • --create-loop does NOT strictly validate work_source shape; missing work_source continues to fall back to "single" (R1). acceptance_criteria is an optional free-form field (string OR list of strings).
  • Tests: test_ci_triage_template_has_work_source, test_ci_triage_template_has_acceptance_criteria, test_create_loop_preserves_acceptance_criteria.

R8 -- status.py --audit --json mode

  • Add --json support to cmd_audit. When --json is set, emit a single JSON line on stdout (machine-readable, pickable by _run_json):
    • {"violations": [...], "loops": [...], "total_tasks": <int>, "untracked_tasks": <int>}
    • Each violation: {"category": <int 1-6>, "severity": "high"|"med"|"low", "task": <str|null>, "message": <str>, "resolved": false}.
  • Existing human-readable --audit output (no --json) is unchanged.
  • This is the data source the audit work consumes (R2).
  • Tests: test_audit_json_emits_violations_array, test_audit_json_includes_loops_block, test_audit_json_pickable_by_runner_run_json.

R9 -- New test file tests/test_goal_mode.py

  • Mirrors test_loop_runner.py's stubbing pattern (monkeypatch.setattr(subprocess, "run", fake_run)) and test_status_brakes.py's --audit --json assertions.
  • Covers R1-R8 as itemized above; target 12-16 tests.
  • Add one regression test: test_existing_single_work_source_loop_ticks_unchanged -- a loop with current_task set and no work_source field still ticks exactly as before (backward compat with all task-3 fixtures).
  • All subprocess calls stubbed; no live LLM in CI.

Verification

  • python3 -m py_compile scripts/loop-runner.py scripts/status.py
  • python3 -m pytest tests/test_goal_mode.py -v
  • python3 -m pytest tests/ -q -- full suite must remain green; expected total ≈ 340 (328 + 12-16 new).
  • bash -n scripts/*.sh (no shell changes; safety check).