CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
82 lines
8.1 KiB
Markdown
82 lines
8.1 KiB
Markdown
# SPEC: add-goal-mode
|
|
|
|
## Context
|
|
|
|
Task 3 (`add-loop-runner`) shipped the graded JSON parser, `score_history` cap, and the score circuit-breaker gate. The "verifier session, graded JSON, score circuit-breaker" framing from the v1 README is therefore already delivered. Task 4 closes the goal-oriented loop on the **runner side**: gives the runner real work sources beyond `current_task`, feeds the verifier acceptance criteria + a prior-tick hint, and closes the `next_hint` feedback path into the next tick's Implement/Verify sessions.
|
|
|
|
## Non-Goals (deferred)
|
|
|
|
- `--goal` CLI flag → v1.1 (R9 from research; adds CLI surface without serving any v1 design doc requirement).
|
|
- `loop-verifier.md` / `loop-implement.md` / `loop-orchestrate.md` prompt **text** → task 6 (this task only wires the substitution tokens; the prompts that consume them land in task 6).
|
|
- `backlog` integration with the `design/context-sizing/` workstream → task 7.
|
|
- `parse_verdict` score clamp + `pass` string coercion → v1.1 hardening (already tracked in task-3 BUG_REPORT).
|
|
- `outputs.retention` in `loop.json` → v1.1.
|
|
- `--create-task` auto-creation from audit violations beyond minimal name resolution → v1.1 hardening.
|
|
|
|
## Requirements
|
|
|
|
### R1 -- `find_work` work_source dispatch
|
|
- Replace the inline `single`-only block in `cmd_tick` with a `_find_work(state, cfg, project_dir)` helper that dispatches on `cfg.get("work_source", {}).get("kind", "single")`.
|
|
- Missing `work_source` field or missing `kind` → fall back to `"single"` with a `.state.log` WARNING line (preserves backward compat with the current `ci-triage/loop.json` template, which has no `work_source` field).
|
|
- `single` with null `current_task` → SKIP `no_current_task` (unchanged from task 3).
|
|
- All kinds write the resolved task name into `state["current_task"]` before returning so downstream steps see it.
|
|
- Unknown `kind` → WARNING + fallback to `"single"`.
|
|
- **Tests:** `test_find_work_single`, `test_find_work_missing_work_source_falls_back_to_single`, `test_find_work_unknown_kind_warns_and_falls_back`.
|
|
|
|
### R2 -- `audit` work_source
|
|
- `work_source.kind == "audit"` → call `status.py --audit --json --project <p>` via `_run_json`. Parse the violations list. Pick the highest-severity unresolved violation (severity ordering: high > med > low). Use the violation's `task` field as `current_task` when present. If the violation has no associated task, call `status.py --create-task <slug>` (slug derived from the violation message) and set the new task as `current_task`. If no unresolved violations → SKIP `no_work` (new skip reason; CLEAN scheduler exit; does not increment `iteration_count`).
|
|
- `work_source.project` (optional) overrides the project path passed to `--audit`; defaults to the loop's own project.
|
|
- **Tests:** `test_audit_picks_highest_severity_violation`, `test_audit_creates_task_when_violation_has_no_task`, `test_audit_skip_when_no_violations`, `test_audit_uses_work_source_project`.
|
|
|
|
### R3 -- `backlog` work_source
|
|
- `work_source.kind == "backlog"` → read `<framework>/design/<area>/BACKLOG.md` where `area` comes from `work_source.area` (default `"loops"`). Parse the topmost `- [ ]` checkbox line. Map to a task name by slugifying the item's bold heading (e.g. `**design-update-loop-template**` → `design-update-loop-template`). Set as `current_task`. If no `[ ]` items remain → SKIP `no_work`.
|
|
- `work_source.area` overrides the area path under `design/`.
|
|
- **Tests:** `test_backlog_picks_top_unchecked_item`, `test_backlog_skip_when_empty`, `test_backlog_uses_area_path`.
|
|
|
|
### R4 -- Verifier-prompt token plumbing
|
|
- Extend the substitution map in `_invoke_harness()` / `_substitute()` to recognize three new tokens (in addition to the existing seven: `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`):
|
|
- `{task_brief}` -- read from `<task_dir>/RESEARCH.md` if present, else `<task_dir>/DESIGN.md`, else `<task_dir>/SPEC.md`, else empty string. Capped at 4k tokens via R5.
|
|
- `{acceptance_criteria}` -- read from `loop.json` `acceptance_criteria` (string OR list; list joined with newlines). Capped at 2k tokens.
|
|
- `{next_hint}` -- read from `state.get("last_verdict", {}).get("next_hint", "")` (empty on first tick or after `--approve`). Capped at 1k tokens.
|
|
- Tokens absent from the prompt stay literal (same rule as today -- a prompt that omits `{task_brief}` is unaffected).
|
|
- **Tests:** `test_task_brief_substituted_from_research`, `test_acceptance_criteria_substituted_from_loop_json_list`, `test_next_hint_substituted_from_last_verdict`, `test_missing_tokens_leave_prompt_intact`.
|
|
|
|
### R5 -- `_truncate_tokens(text, max_tokens)` helper
|
|
- Stdlib-only approximate token cap. No tokenizer dependency. Heuristic: `max_tokens * 4` chars (4-chars-per-token approximation). When the input exceeds the char budget, truncate and append a trailing ` …[truncated]` marker. Used for `task_brief` (4000), `acceptance_criteria` (2000), `next_hint` (1000).
|
|
- Inputs at or under the cap are returned unchanged.
|
|
- **Tests:** `test_truncate_short_text_unchanged`, `test_truncate_long_text_capped_with_marker`, `test_truncate_returns_empty_for_empty_input`.
|
|
|
|
### R6 -- `next_hint` feedback loop closure
|
|
- The Implement and Verify harness invocations receive `{next_hint}` from `state["last_verdict"]["next_hint"]` via R4. This closes the loop: tick N's verifier hint becomes tick N+1's Implement context.
|
|
- A tick with no prior verdict (first tick, or after `--approve` cleared state) passes an empty `{next_hint}` string (no KeyError, no spurious substitution).
|
|
- `last_verdict` is cleared on `--approve --loop` (already happens today via the resume path -- verify and assert in tests).
|
|
- **Tests:** `test_next_hint_fed_into_next_tick_implement`, `test_first_tick_has_empty_next_hint`.
|
|
|
|
### R7 -- `loop.json` schema additions
|
|
- Document `work_source` and `acceptance_criteria` fields in `design/loops/technical.md` schema section (§2) and the self-improvement example (§9).
|
|
- Update `templates/loops/ci-triage/loop.json` to include:
|
|
- `"work_source": {"kind": "single"}` (explicit; current template omits the field entirely).
|
|
- `"acceptance_criteria": ["All R-numbers from SPEC.md are implemented.", "Tests pass with no regressions.", "Pipeline driven to complete."]` (self-documenting placeholder; `null` roles stay -- task 6 fills them with prompt text).
|
|
- `--create-loop` does NOT strictly validate `work_source` shape; missing `work_source` continues to fall back to `"single"` (R1). `acceptance_criteria` is an optional free-form field (string OR list of strings).
|
|
- **Tests:** `test_ci_triage_template_has_work_source`, `test_ci_triage_template_has_acceptance_criteria`, `test_create_loop_preserves_acceptance_criteria`.
|
|
|
|
### R8 -- `status.py --audit --json` mode
|
|
- Add `--json` support to `cmd_audit`. When `--json` is set, emit a single JSON line on stdout (machine-readable, pickable by `_run_json`):
|
|
- `{"violations": [...], "loops": [...], "total_tasks": <int>, "untracked_tasks": <int>}`
|
|
- Each violation: `{"category": <int 1-6>, "severity": "high"|"med"|"low", "task": <str|null>, "message": <str>, "resolved": false}`.
|
|
- Existing human-readable `--audit` output (no `--json`) is **unchanged**.
|
|
- This is the data source the `audit` work consumes (R2).
|
|
- **Tests:** `test_audit_json_emits_violations_array`, `test_audit_json_includes_loops_block`, `test_audit_json_pickable_by_runner_run_json`.
|
|
|
|
### R9 -- New test file `tests/test_goal_mode.py`
|
|
- Mirrors `test_loop_runner.py`'s stubbing pattern (`monkeypatch.setattr(subprocess, "run", fake_run)`) and `test_status_brakes.py`'s `--audit --json` assertions.
|
|
- Covers R1-R8 as itemized above; target 12-16 tests.
|
|
- Add one regression test: `test_existing_single_work_source_loop_ticks_unchanged` -- a loop with `current_task` set and no `work_source` field still ticks exactly as before (backward compat with all task-3 fixtures).
|
|
- All subprocess calls stubbed; no live LLM in CI.
|
|
|
|
## Verification
|
|
|
|
- `python3 -m py_compile scripts/loop-runner.py scripts/status.py`
|
|
- `python3 -m pytest tests/test_goal_mode.py -v`
|
|
- `python3 -m pytest tests/ -q` -- full suite must remain green; expected total ≈ 340 (328 + 12-16 new).
|
|
- `bash -n scripts/*.sh` (no shell changes; safety check). |