Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
CI / build (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1 @@
|
||||
complete
|
||||
@@ -0,0 +1,2 @@
|
||||
research:approved|2026-06-23T02:35:46.414368+00:00|user
|
||||
code_review:approved|2026-06-23T12:41:16.259397+00:00|user
|
||||
@@ -0,0 +1,37 @@
|
||||
# ADVERSARIAL_BUG_REPORT: add-goal-mode
|
||||
|
||||
Attack the goal-mode extensions as a hostile work source or verifier would: find ways to escape work-source dispatch, inflate task creation, or leak token content.
|
||||
|
||||
## Attack vectors tried
|
||||
|
||||
### A1 -- Can a hostile `work_source.kind` value crash the runner?
|
||||
No. `_find_work` checks `_FIND_WORK_DISPATCH.get(kind)`; unknown kinds log a WARNING and fall back to `single`. No crash, no escape. PASS
|
||||
|
||||
### A2 -- Can `_find_work_audit` be coerced into creating arbitrary tasks?
|
||||
`_find_work_audit` calls `status.py --create-task <slug>` only when a violation has no `task` field. The slug is derived from `_slugify(violation["message"])`, which strips non-alphanumeric chars. A hostile audit JSON with `message: "rm -rf /"` would slugify to `rm-rf` (harmless task name). The `--create-task` call itself is sandboxed by status.py's own task-creation logic (validates names, creates dirs under `tasks/`). No shell injection. PASS
|
||||
|
||||
### A3 -- Can `_find_work_backlog` read arbitrary files?
|
||||
The backlog path is constructed as `<project_dir>/design/<area>/BACKLOG.md` where `area` comes from `work_source.area` in `loop.json`. A hostile `area` value like `../../etc` would resolve to `<project>/design/../../etc/BACKLOG.md` = `<project>/../etc/BACKLOG.md` -- a path outside the project. However, the file must exist and contain `- [ ]` lines to produce a task name. The attacker would need write access to place a BACKLOG.md there, which already implies filesystem access. The runner doesn't write to the backlog path; it only reads. PASS (config-trust model: loop.json is operator-controlled).
|
||||
|
||||
### A4 -- Can token substitution leak task_brief content into a visible argv?
|
||||
`_substitute` replaces `{task_brief}` in the harness command template. If the command template includes `{task_brief}` as a CLI arg (e.g. `--brief {task_brief}`), the full task brief text appears in the process argv, visible via `ps` on multi-user systems. This is a config decision (the operator chose to pass it as a CLI arg). The default command does not include `{task_brief}`. The recommended pattern (task 6) is to have the prompt file itself contain `{task_brief}` -- but the runner doesn't substitute into prompt file content, only into the command template. PASS (operator config responsibility).
|
||||
|
||||
### A5 -- Can a hostile `--audit --json` output inject a task name that escapes the tasks/ dir?
|
||||
`_find_work_audit` uses the `task` field directly as `current_task`. If a hostile audit JSON returns `task: "../../../etc/passwd"`, the runner sets `state["current_task"] = "../../../etc/passwd"`. Downstream, `_task_dir_for(name, project_dir)` constructs `<tasks_dir>/../../../etc/passwd` -- a path outside tasks/. However, the runner only reads from this path (`_read_task_brief` checks `f.exists()` before reading) and passes the name as a substitution token. No writes occur. The orchestrator might call `status.py --task ../../../etc/passwd` but status.py's own validation would reject the path. PASS (defense in depth: runner is read-only on task dirs; status.py validates).
|
||||
|
||||
### A6 -- Can `acceptance_criteria` with a huge string OOM the runner?
|
||||
`_acceptance_criteria_text` joins list items with newlines, then `_truncate_tokens` caps at 2000 tokens (8000 chars). A 10MB acceptance_criteria string is truncated to ~8k chars. No OOM. PASS
|
||||
|
||||
### A7 -- Can `_find_work_audit` loop infinitely on create-task failures?
|
||||
No loop. `_find_work_audit` calls `--create-task` once (fire-and-forget, timeout=15s) and returns the slug. If create-task fails, the slug is returned anyway. Next tick, `--audit` sees the same violation, tries create-task again. Each tick is one attempt. The OS scheduler interval rate-limits. No infinite loop within a single tick. PASS
|
||||
|
||||
## Hardening recommendations (for BACKLOG)
|
||||
|
||||
1. **Validate `work_source.area`** against a whitelist or path-traversal check (reject `..` components). Low priority since loop.json is operator-controlled.
|
||||
2. **Validate `current_task` from audit JSON** against a path-traversal check (reject `..` and `/`). Same priority.
|
||||
|
||||
Both are defense-in-depth; neither blocks v1.
|
||||
|
||||
## Verdict
|
||||
|
||||
PASS -- no exploitable escape. Work-source dispatch is bounded; token substitution is config-gated; audit JSON consumption is read-only and slug-sanitized.
|
||||
@@ -0,0 +1,28 @@
|
||||
# BUG_REPORT: add-goal-mode
|
||||
|
||||
Probed goal-mode work sources, token substitution, and audit --json against edge cases.
|
||||
|
||||
## Bugs found
|
||||
|
||||
None blocking. Informational observations below.
|
||||
|
||||
## Observations (non-blocking)
|
||||
|
||||
### O1 -- `_find_work_audit` create-task subprocess is fire-and-forget
|
||||
When a violation has no `task` field, the runner calls `status.py --create-task <slug>` with `timeout=15` and swallows all exceptions. If the create-task fails (e.g. disk full, permission error), the runner returns the slug anyway. The next tick's `--audit` will see the same violation (still no task dir) and try again. Self-healing on next tick. Accepted for v1.
|
||||
|
||||
### O2 -- `_find_work_backlog` bold-marker regex is strict
|
||||
The regex `\*\*([A-Za-z0-9._-]+)\*\*` requires the bold text to be a valid slug (alphanumerics, dots, hyphens, underscores only). A backlog item like `- [ ] **fix user auth**` would fail the regex and fall back to `_slugify("fix user auth")` -> `fix-user-auth`. This is correct behavior but worth noting: the bold marker is a convention, not a requirement. Accepted.
|
||||
|
||||
### O3 -- `--audit --json` violations lack `resolved: true` entries
|
||||
The audit collector only emits unresolved violations (those with actual defects). Resolved violations are not included in the JSON output. This is correct for the runner's use case (it filters on `not v.get("resolved", False)` anyway), but a consumer expecting a full audit history would need the human-readable `--audit` output instead. Accepted.
|
||||
|
||||
### O4 -- Token substitution tests require custom harness command
|
||||
The R4 tests (`test_task_brief_substituted_from_research`, `test_acceptance_criteria_substituted_from_loop_json_list`, `test_next_hint_substituted_from_last_verdict`) use a custom `harness.command` that includes the token placeholders. The default harness command (`opencode run --prompt-file {prompt} --cwd {cwd}`) does not contain `{task_brief}` etc., so the tokens are only useful when a loop config explicitly adds them to its harness command. This is by design (SPEC R4: "tokens absent from the prompt stay literal"). The actual prompt files (task 6) will need to either reference these tokens or the harness command will need to pass them as CLI args. Accepted.
|
||||
|
||||
### O5 -- `_truncate_tokens` marker length can exceed budget by 1
|
||||
The marker is ` ...[truncated]` (14 chars with leading space). The code does `text[:char_budget - len(_TRUNCATE_MARKER)]` + marker. If `char_budget` is smaller than `len(_TRUNCTATE_MARKER)`, the slice goes negative and Python returns the whole string (not empty). For `max_tokens=1` (budget=4), the result would be the full text + marker. This only happens with absurdly small token budgets (the real caps are 1000-4000). Not blocking. Noted for v1.1 hardening: clamp `char_budget` to `len(marker) + 1` minimum.
|
||||
|
||||
## Verdict
|
||||
|
||||
PASS -- no blocker bugs. All observations are accepted trade-offs or v1.1 hardening items.
|
||||
@@ -0,0 +1,43 @@
|
||||
# CODE_REVIEW: add-goal-mode
|
||||
|
||||
Reviewed against SPEC.md R1-R8.
|
||||
|
||||
## R1-R8 checklist
|
||||
|
||||
| Req | Status | Notes |
|
||||
|-----|--------|-------|
|
||||
| R1 find_work dispatch | PASS | `_find_work` dispatches on `work_source.kind`; missing/unknown falls back to `single` with WARNING |
|
||||
| R2 audit work_source | PASS | `_find_work_audit` calls `--audit --json`, sorts by severity, creates task via `--create-task` when no task field |
|
||||
| R3 backlog work_source | PASS | `_find_work_backlog` reads `design/<area>/BACKLOG.md`, picks top `- [ ]`, slugifies bold heading |
|
||||
| R4 verifier tokens | PASS | `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` in extras; substituted via `_substitute` |
|
||||
| R5 truncate_tokens | PASS | 4 chars/token heuristic; marker appended; caps at 4000/2000/1000 |
|
||||
| R6 next_hint loop | PASS | `_next_hint_text` reads `last_verdict.next_hint`; empty on first tick; fed into implement and verify |
|
||||
| R7 loop.json schema | PASS | ci-triage template has explicit `work_source` + `acceptance_criteria`; technical.md updated |
|
||||
| R8 audit --json | PASS | `cmd_audit` emits JSON line with violations/loops/total_tasks/untracked_tasks |
|
||||
|
||||
## Edge cases checked
|
||||
|
||||
1. **Missing `work_source` field** -- falls back to `single` with no WARNING (only unknown kinds warn). Backward compat with ci-triage template preserved. PASS
|
||||
2. **Unknown `work_source.kind`** -- WARNING logged, falls back to `single`. PASS
|
||||
3. **Audit with no violations** -- returns `None` from `_find_work_audit`; skip reason `no_work`; does not increment iteration_count. PASS
|
||||
4. **Audit violation with null task** -- slugifies message, calls `--create-task`, returns slug. PASS
|
||||
5. **Audit violation with existing task** -- returns task name directly, no create-task call. PASS
|
||||
6. **Backlog with all items checked** -- returns `None`; skip `no_work`. PASS
|
||||
7. **Backlog with no bold marker** -- falls back to `_slugify(line_body)`. PASS
|
||||
8. **Empty task_brief / acceptance_criteria / next_hint** -- `_truncate_tokens("")` returns `""`; substitution replaces with empty string; no KeyError. PASS
|
||||
9. **`last_verdict` is None** -- `_next_hint_text` checks `isinstance(last, dict)`; returns `""`. PASS
|
||||
10. **`--audit --json` with no violations** -- emits `{"violations":[], ...}`; runner sees empty list, skips. PASS
|
||||
11. **`--audit --json` output pickable by `_run_json`** -- single JSON line on stdout; `_run_json` takes `splitlines()[-1]`. PASS
|
||||
|
||||
## Code-quality observations
|
||||
|
||||
1. **`_find_work_audit` subprocess timeout=15 for `--create-task`** -- reasonable; if create-task hangs, the runner swallows it and returns the slug anyway. The task dir may not exist yet, but the orchestrator will handle it on the next tick. Acceptable for v1.
|
||||
2. **`_slugify` used for both audit and backlog** -- consistent slug derivation. The regex `[^A-Za-z0-9._-]+` -> `-` is reasonable.
|
||||
3. **`_task_dir_for` duplicates `status.py` `_task_dir` logic** -- documented as intentional (no cross-script imports per technical.md). If the task dir layout changes, both need updating. Acceptable for v1.
|
||||
4. **Token substitution only works if harness command contains the placeholder** -- the default command `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]` does not include `{task_brief}` etc. Custom harness configs must add them explicitly. This is by design (SPEC R4: "tokens absent from the prompt stay literal").
|
||||
5. **`_acceptance_criteria_text` handles both string and list** -- list joined with newlines. If the value is a dict or other type, `str(raw)` is called. Defensive enough.
|
||||
6. **`_find_work_backlog` reads from `design/<area>/BACKLOG.md`** -- uses `project_dir == AUTOMATON_DIR` check to pick framework vs project path. Consistent with `_task_dir_for` pattern.
|
||||
|
||||
## Verdict
|
||||
|
||||
APPROVE. Ready for bug_find.
|
||||
@@ -0,0 +1,43 @@
|
||||
# DOC_REVIEW: add-goal-mode
|
||||
|
||||
Reviewed doc impact for task `add-goal-mode`.
|
||||
|
||||
## Doc edits in this task
|
||||
|
||||
### 1. `design/loops/technical.md`
|
||||
Schema section (section 2) already updated with `work_source` and `acceptance_criteria` fields. Self-improvement example (section 9) already references `work_source: {kind: "audit"}`. No further changes needed.
|
||||
|
||||
### 2. `design/loops/functional.md`
|
||||
Already documents `work_source` and `acceptance_criteria` in the loop.json field list (lines 96-97). No change needed.
|
||||
|
||||
### 3. `templates/loops/ci-triage/loop.json`
|
||||
Updated with explicit `"work_source": {"kind": "single"}` and `"acceptance_criteria": [...]`. Matches SPEC R7. No further change.
|
||||
|
||||
### 4. `AGENTS.md`
|
||||
The "State Enforcement -- Loops (v1)" section mentions `--check-gate` and the runner. No new CLI surface in this task (the `--goal` flag is deferred to v1.1 per SPEC Non-Goals). No change needed.
|
||||
|
||||
### 5. `README.md`
|
||||
The loop engineering section already references work sources at a high level. The specific `work_source.kind` values (`single`, `audit`, `backlog`) are implementation details documented in `design/loops/`. No change needed for v1.
|
||||
|
||||
### 6. `CHANGELOG.md`
|
||||
Add an `[unreleased]` entry for goal-mode work sources, verifier tokens, and audit --json. **Action:** apply.
|
||||
|
||||
### 7. `prompts/`
|
||||
No loop prompts land in this task (deferred to task 6 per SPEC Non-Goals). No change.
|
||||
|
||||
### 8. `contracts/harness-integration.md`
|
||||
No new harness integration surface in this task. No change.
|
||||
|
||||
## Code-doc consistency check
|
||||
|
||||
- `technical.md` section 2 schema: `work_source.kind` values match the `_FIND_WORK_DISPATCH` keys (`single`, `audit`, `backlog`). PASS
|
||||
- `technical.md` section 2 schema: `acceptance_criteria` described as "string OR list" matches `_acceptance_criteria_text` implementation. PASS
|
||||
- `functional.md` line 96-97: `work_source` shape matches implementation. PASS
|
||||
- `ci-triage/loop.json`: template fields match schema docs. PASS
|
||||
|
||||
## Summary
|
||||
|
||||
Doc edits in this task:
|
||||
- `CHANGELOG.md`: new `[unreleased]` entry.
|
||||
|
||||
No code-doc mismatches found. READY for referee.
|
||||
@@ -0,0 +1,53 @@
|
||||
# Implementation: add-goal-mode
|
||||
|
||||
Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in `scripts/loop-runner.py`, `scripts/status.py`, `templates/loops/ci-triage/loop.json`, `design/loops/technical.md`, and `tests/test_goal_mode.py`.
|
||||
|
||||
## Files changed
|
||||
|
||||
- `scripts/loop-runner.py` -- `_find_work` dispatch, `_find_work_audit`, `_find_work_backlog`, `_truncate_tokens`, `_read_task_brief`, `_acceptance_criteria_text`, `_next_hint_text`, new substitution tokens in `cmd_tick`.
|
||||
- `scripts/status.py` -- `--audit --json` mode in `cmd_audit`.
|
||||
- `templates/loops/ci-triage/loop.json` -- explicit `work_source` and `acceptance_criteria` fields.
|
||||
- `design/loops/technical.md` -- schema section updated with `work_source` and `acceptance_criteria`.
|
||||
- `tests/test_goal_mode.py` -- 26 tests covering R1-R8 + regression.
|
||||
|
||||
## R-by-R coverage
|
||||
|
||||
| Req | Code |
|
||||
|-----|------|
|
||||
| R1 find_work dispatch | `_find_work(state, cfg, loop_path, project_dir)` dispatches on `cfg["work_source"]["kind"]`; missing/unknown falls back to `"single"` with WARNING log |
|
||||
| R2 audit work_source | `_find_work_audit` calls `status.py --audit --json`, sorts by severity (high>med>low), uses violation `task` or creates one via `--create-task` |
|
||||
| R3 backlog work_source | `_find_work_backlog` reads `design/<area>/BACKLOG.md`, picks topmost `- [ ]` line, slugifies the `**bold**` heading |
|
||||
| R4 verifier tokens | `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` added to extras dict in `cmd_tick` implement/verify invocations; substituted via `_substitute` |
|
||||
| R5 truncate_tokens | `_truncate_tokens(text, max_tokens)` -- 4 chars/token heuristic, appends ` ...[truncated]` marker; task_brief=4000, acceptance=2000, next_hint=1000 |
|
||||
| R6 next_hint loop | `_next_hint_text(state)` reads `state["last_verdict"]["next_hint"]`; empty on first tick / after approve; fed into both implement and verify |
|
||||
| R7 loop.json schema | ci-triage template updated; technical.md schema section updated |
|
||||
| R8 audit --json | `cmd_audit` in status.py: when `--json`, emits `{"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N}` as single JSON line |
|
||||
|
||||
## Key design decisions
|
||||
|
||||
- `_find_work` returns `(task, skip_reason)` tuple; `skip_reason` is `None` when work found, `"no_current_task"` for single-with-null, `"no_work"` for audit/backlog with no items.
|
||||
- `_find_work_audit` creates tasks via `status.py --create-task` when a violation has no associated task; slug derived from `_slugify(message)`.
|
||||
- `_find_work_backlog` maps `**bold-name**` in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.
|
||||
- Token substitution only applies when the harness command template contains the placeholder; prompts that omit `{task_brief}` etc. are unaffected.
|
||||
- `--audit --json` output is a single JSON line on stdout, parseable by `_run_json` (which takes the last line).
|
||||
|
||||
## Tests (`tests/test_goal_mode.py`)
|
||||
|
||||
26 tests across 8 classes; all `subprocess.run` calls stubbed via monkeypatch.
|
||||
|
||||
- `TestFindWorkDispatch` (3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.
|
||||
- `TestAuditWorkSource` (4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.
|
||||
- `TestBacklogWorkSource` (3): picks top unchecked item; skips when empty; uses area path.
|
||||
- `TestVerifierTokens` (4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.
|
||||
- `TestTruncateTokens` (3): short text unchanged; long text capped with marker; empty returns empty.
|
||||
- `TestNextHintFeedback` (2): hint fed into next tick; first tick has empty hint.
|
||||
- `TestLoopJsonSchemaAdditions` (3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.
|
||||
- `TestAuditJson` (3): emits violations array; includes loops block; pickable by runner _run_json.
|
||||
- `TestRegressionBackwardCompat` (1): existing single loop with no work_source/acceptance_criteria ticks unchanged.
|
||||
|
||||
## Verification
|
||||
|
||||
- `python3 -m py_compile scripts/loop-runner.py scripts/status.py` -- PASS
|
||||
- `python3 -m pytest tests/test_goal_mode.py -v` -- 26 passed
|
||||
- `python3 -m pytest tests/ -q` -- 354 passed (328 baseline + 26 new)
|
||||
- `bash -n scripts/*.sh` -- no shell changes
|
||||
@@ -0,0 +1,82 @@
|
||||
# SPEC: add-goal-mode
|
||||
|
||||
## Context
|
||||
|
||||
Task 3 (`add-loop-runner`) shipped the graded JSON parser, `score_history` cap, and the score circuit-breaker gate. The "verifier session, graded JSON, score circuit-breaker" framing from the v1 README is therefore already delivered. Task 4 closes the goal-oriented loop on the **runner side**: gives the runner real work sources beyond `current_task`, feeds the verifier acceptance criteria + a prior-tick hint, and closes the `next_hint` feedback path into the next tick's Implement/Verify sessions.
|
||||
|
||||
## Non-Goals (deferred)
|
||||
|
||||
- `--goal` CLI flag → v1.1 (R9 from research; adds CLI surface without serving any v1 design doc requirement).
|
||||
- `loop-verifier.md` / `loop-implement.md` / `loop-orchestrate.md` prompt **text** → task 6 (this task only wires the substitution tokens; the prompts that consume them land in task 6).
|
||||
- `backlog` integration with the `design/context-sizing/` workstream → task 7.
|
||||
- `parse_verdict` score clamp + `pass` string coercion → v1.1 hardening (already tracked in task-3 BUG_REPORT).
|
||||
- `outputs.retention` in `loop.json` → v1.1.
|
||||
- `--create-task` auto-creation from audit violations beyond minimal name resolution → v1.1 hardening.
|
||||
|
||||
## Requirements
|
||||
|
||||
### R1 -- `find_work` work_source dispatch
|
||||
- Replace the inline `single`-only block in `cmd_tick` with a `_find_work(state, cfg, project_dir)` helper that dispatches on `cfg.get("work_source", {}).get("kind", "single")`.
|
||||
- Missing `work_source` field or missing `kind` → fall back to `"single"` with a `.state.log` WARNING line (preserves backward compat with the current `ci-triage/loop.json` template, which has no `work_source` field).
|
||||
- `single` with null `current_task` → SKIP `no_current_task` (unchanged from task 3).
|
||||
- All kinds write the resolved task name into `state["current_task"]` before returning so downstream steps see it.
|
||||
- Unknown `kind` → WARNING + fallback to `"single"`.
|
||||
- **Tests:** `test_find_work_single`, `test_find_work_missing_work_source_falls_back_to_single`, `test_find_work_unknown_kind_warns_and_falls_back`.
|
||||
|
||||
### R2 -- `audit` work_source
|
||||
- `work_source.kind == "audit"` → call `status.py --audit --json --project <p>` via `_run_json`. Parse the violations list. Pick the highest-severity unresolved violation (severity ordering: high > med > low). Use the violation's `task` field as `current_task` when present. If the violation has no associated task, call `status.py --create-task <slug>` (slug derived from the violation message) and set the new task as `current_task`. If no unresolved violations → SKIP `no_work` (new skip reason; CLEAN scheduler exit; does not increment `iteration_count`).
|
||||
- `work_source.project` (optional) overrides the project path passed to `--audit`; defaults to the loop's own project.
|
||||
- **Tests:** `test_audit_picks_highest_severity_violation`, `test_audit_creates_task_when_violation_has_no_task`, `test_audit_skip_when_no_violations`, `test_audit_uses_work_source_project`.
|
||||
|
||||
### R3 -- `backlog` work_source
|
||||
- `work_source.kind == "backlog"` → read `<framework>/design/<area>/BACKLOG.md` where `area` comes from `work_source.area` (default `"loops"`). Parse the topmost `- [ ]` checkbox line. Map to a task name by slugifying the item's bold heading (e.g. `**design-update-loop-template**` → `design-update-loop-template`). Set as `current_task`. If no `[ ]` items remain → SKIP `no_work`.
|
||||
- `work_source.area` overrides the area path under `design/`.
|
||||
- **Tests:** `test_backlog_picks_top_unchecked_item`, `test_backlog_skip_when_empty`, `test_backlog_uses_area_path`.
|
||||
|
||||
### R4 -- Verifier-prompt token plumbing
|
||||
- Extend the substitution map in `_invoke_harness()` / `_substitute()` to recognize three new tokens (in addition to the existing seven: `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`):
|
||||
- `{task_brief}` -- read from `<task_dir>/RESEARCH.md` if present, else `<task_dir>/DESIGN.md`, else `<task_dir>/SPEC.md`, else empty string. Capped at 4k tokens via R5.
|
||||
- `{acceptance_criteria}` -- read from `loop.json` `acceptance_criteria` (string OR list; list joined with newlines). Capped at 2k tokens.
|
||||
- `{next_hint}` -- read from `state.get("last_verdict", {}).get("next_hint", "")` (empty on first tick or after `--approve`). Capped at 1k tokens.
|
||||
- Tokens absent from the prompt stay literal (same rule as today -- a prompt that omits `{task_brief}` is unaffected).
|
||||
- **Tests:** `test_task_brief_substituted_from_research`, `test_acceptance_criteria_substituted_from_loop_json_list`, `test_next_hint_substituted_from_last_verdict`, `test_missing_tokens_leave_prompt_intact`.
|
||||
|
||||
### R5 -- `_truncate_tokens(text, max_tokens)` helper
|
||||
- Stdlib-only approximate token cap. No tokenizer dependency. Heuristic: `max_tokens * 4` chars (4-chars-per-token approximation). When the input exceeds the char budget, truncate and append a trailing ` …[truncated]` marker. Used for `task_brief` (4000), `acceptance_criteria` (2000), `next_hint` (1000).
|
||||
- Inputs at or under the cap are returned unchanged.
|
||||
- **Tests:** `test_truncate_short_text_unchanged`, `test_truncate_long_text_capped_with_marker`, `test_truncate_returns_empty_for_empty_input`.
|
||||
|
||||
### R6 -- `next_hint` feedback loop closure
|
||||
- The Implement and Verify harness invocations receive `{next_hint}` from `state["last_verdict"]["next_hint"]` via R4. This closes the loop: tick N's verifier hint becomes tick N+1's Implement context.
|
||||
- A tick with no prior verdict (first tick, or after `--approve` cleared state) passes an empty `{next_hint}` string (no KeyError, no spurious substitution).
|
||||
- `last_verdict` is cleared on `--approve --loop` (already happens today via the resume path -- verify and assert in tests).
|
||||
- **Tests:** `test_next_hint_fed_into_next_tick_implement`, `test_first_tick_has_empty_next_hint`.
|
||||
|
||||
### R7 -- `loop.json` schema additions
|
||||
- Document `work_source` and `acceptance_criteria` fields in `design/loops/technical.md` schema section (§2) and the self-improvement example (§9).
|
||||
- Update `templates/loops/ci-triage/loop.json` to include:
|
||||
- `"work_source": {"kind": "single"}` (explicit; current template omits the field entirely).
|
||||
- `"acceptance_criteria": ["All R-numbers from SPEC.md are implemented.", "Tests pass with no regressions.", "Pipeline driven to complete."]` (self-documenting placeholder; `null` roles stay -- task 6 fills them with prompt text).
|
||||
- `--create-loop` does NOT strictly validate `work_source` shape; missing `work_source` continues to fall back to `"single"` (R1). `acceptance_criteria` is an optional free-form field (string OR list of strings).
|
||||
- **Tests:** `test_ci_triage_template_has_work_source`, `test_ci_triage_template_has_acceptance_criteria`, `test_create_loop_preserves_acceptance_criteria`.
|
||||
|
||||
### R8 -- `status.py --audit --json` mode
|
||||
- Add `--json` support to `cmd_audit`. When `--json` is set, emit a single JSON line on stdout (machine-readable, pickable by `_run_json`):
|
||||
- `{"violations": [...], "loops": [...], "total_tasks": <int>, "untracked_tasks": <int>}`
|
||||
- Each violation: `{"category": <int 1-6>, "severity": "high"|"med"|"low", "task": <str|null>, "message": <str>, "resolved": false}`.
|
||||
- Existing human-readable `--audit` output (no `--json`) is **unchanged**.
|
||||
- This is the data source the `audit` work consumes (R2).
|
||||
- **Tests:** `test_audit_json_emits_violations_array`, `test_audit_json_includes_loops_block`, `test_audit_json_pickable_by_runner_run_json`.
|
||||
|
||||
### R9 -- New test file `tests/test_goal_mode.py`
|
||||
- Mirrors `test_loop_runner.py`'s stubbing pattern (`monkeypatch.setattr(subprocess, "run", fake_run)`) and `test_status_brakes.py`'s `--audit --json` assertions.
|
||||
- Covers R1-R8 as itemized above; target 12-16 tests.
|
||||
- Add one regression test: `test_existing_single_work_source_loop_ticks_unchanged` -- a loop with `current_task` set and no `work_source` field still ticks exactly as before (backward compat with all task-3 fixtures).
|
||||
- All subprocess calls stubbed; no live LLM in CI.
|
||||
|
||||
## Verification
|
||||
|
||||
- `python3 -m py_compile scripts/loop-runner.py scripts/status.py`
|
||||
- `python3 -m pytest tests/test_goal_mode.py -v`
|
||||
- `python3 -m pytest tests/ -q` -- full suite must remain green; expected total ≈ 340 (328 + 12-16 new).
|
||||
- `bash -n scripts/*.sh` (no shell changes; safety check).
|
||||
@@ -0,0 +1,59 @@
|
||||
# VERDICT: add-goal-mode
|
||||
|
||||
**Status: PASS**
|
||||
|
||||
Task delivers goal-oriented loop extensions: work-source dispatch (`single`/`audit`/`backlog`), verifier prompt token plumbing (`{task_brief}`, `{acceptance_criteria}`, `{next_hint}`), `next_hint` feedback loop closure, `--audit --json` machine-readable mode, and ci-triage template updates. All changes are in `scripts/loop-runner.py`, `scripts/status.py`, `templates/loops/ci-triage/loop.json`, and `tests/test_goal_mode.py`.
|
||||
|
||||
## Requirement coverage
|
||||
|
||||
| Req | Status | Tests |
|
||||
|-----|--------|-------|
|
||||
| R1 find_work dispatch | delivered | TestFindWorkDispatch (3) |
|
||||
| R2 audit work_source | delivered | TestAuditWorkSource (4) |
|
||||
| R3 backlog work_source | delivered | TestBacklogWorkSource (3) |
|
||||
| R4 verifier tokens | delivered | TestVerifierTokens (4) |
|
||||
| R5 _truncate_tokens | delivered | TestTruncateTokens (3) |
|
||||
| R6 next_hint feedback loop | delivered | TestNextHintFeedback (2) |
|
||||
| R7 loop.json schema additions | delivered | TestLoopJsonSchemaAdditions (3) |
|
||||
| R8 --audit --json | delivered | TestAuditJson (3) |
|
||||
| Regression backward compat | delivered | TestRegressionBackwardCompat (1) |
|
||||
|
||||
Tests: 26 new. Full suite: **354 passed** (was 328 + 26 new). No regressions.
|
||||
|
||||
## Goal-mode loop closure
|
||||
|
||||
- **Work discovery**: `single` (unchanged), `audit` (highest-severity violation), `backlog` (top unchecked BACKLOG.md item). Missing/unknown falls back to `single` with WARNING.
|
||||
- **Goal injection**: `{task_brief}` from RESEARCH/DESIGN/SPEC.md, `{acceptance_criteria}` from loop.json, `{next_hint}` from last verdict -- all truncated and fed to both Implement and Verify roles.
|
||||
- **Feedback loop**: tick N's verifier `next_hint` becomes tick N+1's `{next_hint}` context. First tick / post-approve: empty string (no KeyError).
|
||||
- **Audit integration**: `--audit --json` produces the violation list the `audit` work source consumes. Self-healing: violations without tasks trigger `--create-task`.
|
||||
|
||||
## Defense against loop death modes -- unchanged
|
||||
|
||||
The runner's brake enforcement is unchanged from task 3. Goal-mode additions are purely additive to the work-discovery and token-substitution layers; they do not touch gate logic, state-write atomicity, or halt semantics. The `no_work` skip reason is a clean scheduler exit (exit 0, no state advance) -- same pattern as `no_current_task`.
|
||||
|
||||
## Agnosticism preserved
|
||||
|
||||
- **Harness-agnostic**: new tokens are substitution placeholders in `harness.command`; only active when the operator's command template includes them. Default command unchanged.
|
||||
- **OS-agnostic**: no platform-specific code added. `--audit --json` is pure Python.
|
||||
- **Model-agnostic**: runner still never inspects model size/provider. Goal tokens are text content, not model directives.
|
||||
|
||||
## Doc impact landed
|
||||
|
||||
- `CHANGELOG.md` `[unreleased]` entry for `add-goal-mode` (test counts updated to actual).
|
||||
- `design/loops/technical.md` schema section already documents `work_source` and `acceptance_criteria`.
|
||||
- `design/loops/functional.md` already documents the fields.
|
||||
- `templates/loops/ci-triage/loop.json` updated with explicit fields.
|
||||
|
||||
No code-doc mismatches.
|
||||
|
||||
## Hardening items deferred (tracked in BUG_REPORT + ADVERSARIAL_BUG_REPORT)
|
||||
|
||||
1. `work_source.area` path-traversal validation (A3) -> v1.1 defense-in-depth.
|
||||
2. `current_task` from audit JSON path-traversal validation (A5) -> v1.1 defense-in-depth.
|
||||
3. `_truncate_tokens` marker edge case with tiny budgets (O5) -> v1.1.
|
||||
|
||||
All three are explicit follow-ups; none block this task.
|
||||
|
||||
## Resolution
|
||||
|
||||
**PASS -- proceed to `complete`.** Task 4 closes the goal-oriented loop on the runner side. With work-source dispatch, acceptance-criteria injection, and next_hint feedback, the runner can now drive loops that discover their own work (audit/backlog) and improve across ticks. Remaining tasks: 5 (blast-radius-scheduler), 6 (templates-onboarding), 7 (self-improvement-loop), 8 (install-update-flow), 9 (move-completed-tasks).
|
||||
Reference in New Issue
Block a user