Files
automaton/tasks/complete/add-goal-mode/IMPLEMENTATION.md
T

54 lines
4.3 KiB
Markdown
Raw Normal View History

# Implementation: add-goal-mode
Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in `scripts/loop-runner.py`, `scripts/status.py`, `templates/loops/ci-triage/loop.json`, `design/loops/technical.md`, and `tests/test_goal_mode.py`.
## Files changed
- `scripts/loop-runner.py` -- `_find_work` dispatch, `_find_work_audit`, `_find_work_backlog`, `_truncate_tokens`, `_read_task_brief`, `_acceptance_criteria_text`, `_next_hint_text`, new substitution tokens in `cmd_tick`.
- `scripts/status.py` -- `--audit --json` mode in `cmd_audit`.
- `templates/loops/ci-triage/loop.json` -- explicit `work_source` and `acceptance_criteria` fields.
- `design/loops/technical.md` -- schema section updated with `work_source` and `acceptance_criteria`.
- `tests/test_goal_mode.py` -- 26 tests covering R1-R8 + regression.
## R-by-R coverage
| Req | Code |
|-----|------|
| R1 find_work dispatch | `_find_work(state, cfg, loop_path, project_dir)` dispatches on `cfg["work_source"]["kind"]`; missing/unknown falls back to `"single"` with WARNING log |
| R2 audit work_source | `_find_work_audit` calls `status.py --audit --json`, sorts by severity (high>med>low), uses violation `task` or creates one via `--create-task` |
| R3 backlog work_source | `_find_work_backlog` reads `design/<area>/BACKLOG.md`, picks topmost `- [ ]` line, slugifies the `**bold**` heading |
| R4 verifier tokens | `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` added to extras dict in `cmd_tick` implement/verify invocations; substituted via `_substitute` |
| R5 truncate_tokens | `_truncate_tokens(text, max_tokens)` -- 4 chars/token heuristic, appends ` ...[truncated]` marker; task_brief=4000, acceptance=2000, next_hint=1000 |
| R6 next_hint loop | `_next_hint_text(state)` reads `state["last_verdict"]["next_hint"]`; empty on first tick / after approve; fed into both implement and verify |
| R7 loop.json schema | ci-triage template updated; technical.md schema section updated |
| R8 audit --json | `cmd_audit` in status.py: when `--json`, emits `{"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N}` as single JSON line |
## Key design decisions
- `_find_work` returns `(task, skip_reason)` tuple; `skip_reason` is `None` when work found, `"no_current_task"` for single-with-null, `"no_work"` for audit/backlog with no items.
- `_find_work_audit` creates tasks via `status.py --create-task` when a violation has no associated task; slug derived from `_slugify(message)`.
- `_find_work_backlog` maps `**bold-name**` in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.
- Token substitution only applies when the harness command template contains the placeholder; prompts that omit `{task_brief}` etc. are unaffected.
- `--audit --json` output is a single JSON line on stdout, parseable by `_run_json` (which takes the last line).
## Tests (`tests/test_goal_mode.py`)
26 tests across 8 classes; all `subprocess.run` calls stubbed via monkeypatch.
- `TestFindWorkDispatch` (3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.
- `TestAuditWorkSource` (4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.
- `TestBacklogWorkSource` (3): picks top unchecked item; skips when empty; uses area path.
- `TestVerifierTokens` (4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.
- `TestTruncateTokens` (3): short text unchanged; long text capped with marker; empty returns empty.
- `TestNextHintFeedback` (2): hint fed into next tick; first tick has empty hint.
- `TestLoopJsonSchemaAdditions` (3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.
- `TestAuditJson` (3): emits violations array; includes loops block; pickable by runner _run_json.
- `TestRegressionBackwardCompat` (1): existing single loop with no work_source/acceptance_criteria ticks unchanged.
## Verification
- `python3 -m py_compile scripts/loop-runner.py scripts/status.py` -- PASS
- `python3 -m pytest tests/test_goal_mode.py -v` -- 26 passed
- `python3 -m pytest tests/ -q` -- 354 passed (328 baseline + 26 new)
- `bash -n scripts/*.sh` -- no shell changes