feat(model-divergence): full enforcement — manifest, transition, claim, audit, loop gates, detect script

Completes all 3 model-divergence enforcement subtasks:

- scripts/detect_models.py: probes opencode.json + localhost endpoints,
  builds models.json with --json/--write/--force
- scripts/status.py: CONFLICT_MATRIX, --model flag, --transition --model,
  --claim --model, --audit Category 6, model-divergence brake gate in
  --check-gate, helpers for manifest loading and conflict checking
- scripts/loop-runner.py: _role_model() helper + {model} passed via extras
  dict to _invoke_harness for implement, verify, orchestrate roles
- tests/test_model_divergence.py: 33 tests covering all enforcement layers
- Single-LLM mode: record model advisory, no conflict check
- Multi-LLM mode (2+ models): conflict matrix enforced at transition, claim,
  and loop brake gate
- Project-level models.json preferred over global ~/.automaton/models.json
This commit is contained in:
Lap Tran
2026-06-26 13:23:17 -04:00
parent bc7daf8590
commit 35e449b03e
648 changed files with 1120 additions and 15492 deletions
+11
View File
@@ -51,6 +51,17 @@ To disable auto-detection and use manual values:
- **Override context window**: 128k
```
## Available Models
Models available for model-divergence enforcement. This file is managed by `scripts/detect_models.py`. In single-LLM mode (0-1 models), no hard blocks are enforced. In multi-LLM mode (2+ models), the conflict matrix enforces role-model separation.
- **Default**: omlx/Ornith-1.0-35B-4bit-mlx # Used when no role-specific binding is set
- **Advised**: true # Recommend a second model in single-LLM mode
No additional models are configured in the manifest. To add models:
1. Run `python3 ~/.automaton/scripts/detect_models.py --write` to auto-detect from opencode.json and localhost endpoints.
2. Or manually create `~/.automaton/models.json` (see `design/framework/technical.md` §2 for schema).
## System Requirements
Requirements for the environment the framework runs in.
+1 -1
View File
@@ -1,2 +1,2 @@
#!/usr/bin/env bash
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/private/var/folders/f5/yv0dzbnx47x3yp8sc_2519gh0000gn/T/pytest-of-laptran/pytest-120/test_uninstall_via_disabled_re0"
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/private/var/folders/f5/yv0dzbnx47x3yp8sc_2519gh0000gn/T/pytest-of-laptran/pytest-126/test_uninstall_via_disabled_re0"
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env python3
"""Probe opencode.json and localhost endpoints to produce a candidate models.json.
Usage:
python3 scripts/detect_models.py [--json] [--write]
Without --json, prints a human-readable report.
With --json, emits the candidate models.json to stdout as the last JSON line.
With --write, writes the candidate to ~/.automaton/models.json (idempotent,
never overwrites an existing file unless --force is also given).
Probing strategy (stdlib only):
1. Parse opencode.json (or opencode.jsonc) for configured provider+model pairs.
2. Probe localhost endpoints to find locally-running LLM servers:
- http://localhost:8080/v1/models (llama.cpp / generic OpenAI-compatible)
- http://localhost:11434/api/tags (Ollama)
- http://localhost:1234/v1/models (LM Studio)
- http://localhost:8000/v1/models (vLLM)
3. Merge results into a candidate models.json.
"""
from __future__ import annotations
import json
import os
import re
import sys
from pathlib import Path
from typing import Optional
AUTOMATON_DIR = Path.home() / ".automaton"
# ---------------------------------------------------------------------------
# opencode.json parsing
# ---------------------------------------------------------------------------
def _find_opencode_json() -> Optional[Path]:
"""Locate the opencode config file (opencode.json or opencode.jsonc)."""
candidates = [
Path.cwd() / "opencode.json",
Path.cwd() / "opencode.jsonc",
AUTOMATON_DIR / "opencode.json",
AUTOMATON_DIR / "opencode.jsonc",
Path.home() / ".opencode.json",
Path.home() / ".config" / "opencode" / "opencode.json",
Path.home() / ".config" / "opencode" / "opencode.jsonc",
]
for p in candidates:
if p.exists():
return p
return None
def _parse_opencode_models(config_path: Path) -> list[dict]:
"""Extract model entries from an opencode.json config.
Expected structure (common patterns):
{
"providers": {
"opencode": { "model": "glm-4.6", ... },
...
}
}
or a flatter:
{
"model": "glm-4.6",
...
}
"""
try:
content = config_path.read_text(encoding="utf-8")
except OSError:
return []
# Strip JSONC comments (// line comments only, sufficient for our use)
content = re.sub(r"//.*", "", content)
try:
data = json.loads(content)
except json.JSONDecodeError:
return []
if not isinstance(data, dict):
return []
models: list[dict] = []
seen: set[str] = set()
# Check top-level "model" field (single-model config)
single = data.get("model")
if isinstance(single, str) and single not in seen:
seen.add(single)
models.append({"name": single, "provider": "opencode", "context_window": None, "location": "remote"})
# Check providers dict
providers = data.get("providers") or {}
for prov_name, prov_cfg in providers.items():
if isinstance(prov_cfg, dict):
model_name = prov_cfg.get("model")
if isinstance(model_name, str) and model_name not in seen:
seen.add(model_name)
models.append({"name": model_name, "provider": prov_name, "context_window": None, "location": "remote"})
# Check "models" list (explicit model roster)
model_list = data.get("models")
if isinstance(model_list, list):
for entry in model_list:
if isinstance(entry, dict):
name = entry.get("name") or entry.get("model")
if isinstance(name, str) and name not in seen:
seen.add(name)
models.append({
"name": name,
"provider": entry.get("provider", "opencode"),
"context_window": entry.get("context_window"),
"location": entry.get("location", "remote"),
})
return models
# ---------------------------------------------------------------------------
# Localhost probing
# ---------------------------------------------------------------------------
def _fetch_json(url: str, timeout: int = 5) -> Optional[dict]:
"""Fetch a JSON response from a URL using urllib (stdlib)."""
import urllib.request
import urllib.error
try:
req = urllib.request.Request(url, method="GET")
with urllib.request.urlopen(req, timeout=timeout) as resp:
body = resp.read().decode("utf-8")
return json.loads(body)
except (OSError, urllib.error.URLError, json.JSONDecodeError, ValueError):
return None
def _probe_ollama() -> list[dict]:
"""Probe Ollama: GET http://localhost:11434/api/tags → models[].name"""
data = _fetch_json("http://localhost:11434/api/tags")
if not data:
return []
models_list = data.get("models") or []
return [
{"name": m.get("name"), "provider": "ollama", "context_window": None, "location": "http://localhost:11434"}
for m in models_list
if isinstance(m, dict) and isinstance(m.get("name"), str)
]
def _probe_openai_compatible(url: str, provider: str) -> list[dict]:
"""Probe an OpenAI-compatible /v1/models endpoint."""
data = _fetch_json(url)
if not data:
return []
model_list = data.get("data") or []
return [
{"name": m.get("id"), "provider": provider, "context_window": None, "location": url}
for m in model_list
if isinstance(m, dict) and isinstance(m.get("id"), str)
]
_ENDPOINTS = [
("http://localhost:8080/v1/models", "llama.cpp"),
("http://localhost:11434/api/tags", "ollama"), # handled separately above
("http://localhost:1234/v1/models", "lm-studio"),
("http://localhost:8000/v1/models", "vllm"),
]
def _probe_localhost() -> list[dict]:
"""Probe all known localhost endpoints and merge results."""
seen_names: set[str] = set()
models: list[dict] = []
for url, provider in _ENDPOINTS:
if provider == "ollama":
result = _probe_ollama()
else:
result = _probe_openai_compatible(url, provider)
for m in result:
n = m.get("name")
if isinstance(n, str) and n not in seen_names:
seen_names.add(n)
models.append(m)
return models
# ---------------------------------------------------------------------------
# Merge & write
# ---------------------------------------------------------------------------
def build_candidate_models(probe_local: bool = True) -> dict:
"""Build a candidate models.json dict.
1. Parse models from opencode.json
2. Optionally probe localhost endpoints
3. Merge: opencode config models come first; local probes fill in gaps.
4. Build result with default, advised, models[].
"""
opencode_path = _find_opencode_json()
config_models: list[dict] = []
if opencode_path:
config_models = _parse_opencode_models(opencode_path)
local_models: list[dict] = []
if probe_local:
local_models = _probe_localhost()
# Merge: key by name, config models take priority (unordered)
merged: dict[str, dict] = {}
for m in config_models:
n = m["name"]
if n not in merged:
merged[n] = m
for m in local_models:
n = m.get("name")
if n and n not in merged:
merged[n] = m
models_list = list(merged.values())
# Determine default: first config model, or first local model, or empty
default_name: Optional[str] = None
if config_models:
default_name = config_models[0].get("name")
elif local_models:
default_name = local_models[0].get("name")
# Determine advised: if only 0-1 models, set advised=true; else false
advised = len(models_list) <= 1
result: dict = {
"schema_version": 1,
"default": default_name,
"advised": advised,
"models": models_list,
}
return result
def write_models_file(candidate: dict, force: bool = False) -> bool:
"""Write candidate models.json to AUTOMATON_DIR.
Never overwrites an existing file unless force=True.
Returns True if written, False if skipped.
"""
target = AUTOMATON_DIR / "models.json"
if target.exists() and not force:
return False
target.write_text(json.dumps(candidate, indent=2) + "\n")
return True
def format_report(candidate: dict) -> str:
"""Human-readable report of the candidate models."""
lines = []
lines.append("=== Model Detection Report ===")
lines.append("")
source = "No opencode.json found" if not _find_opencode_json() else f"Config: {_find_opencode_json()}"
lines.append(f"Source: {source}")
lines.append("")
models = candidate.get("models", [])
if not models:
lines.append("No models detected.")
else:
lines.append(f"Detected {len(models)} model(s):")
for m in models:
loc = m.get("location", "unknown")
prov = m.get("provider", "?")
ctx = m.get("context_window")
ctx_str = f", context: {ctx}" if ctx else ""
lines.append(f" - {m['name']} ({prov}, {loc}{ctx_str})")
lines.append("")
lines.append(f"Default: {candidate.get('default', 'none')}")
lines.append(f"Advised: {candidate.get('advised', False)}")
lines.append(f"Mode: {'multi-LLM' if len(models) >= 2 else 'single-LLM'}")
lines.append("")
target = AUTOMATON_DIR / "models.json"
if target.exists():
lines.append(f"models.json already exists at {target} (use --force to overwrite)")
else:
lines.append(f"Ready to write to {target} (use --write to create)")
return "\n".join(lines)
def main() -> int:
import argparse
parser = argparse.ArgumentParser(description="Detect available LLM models and write models.json")
parser.add_argument("--json", action="store_true", help="Output candidate JSON on last line")
parser.add_argument("--write", action="store_true", help="Write candidate models.json to ~/.automaton/ (idempotent)")
parser.add_argument("--force", action="store_true", help="Overwrite existing models.json")
parser.add_argument("--no-probe", action="store_true", help="Skip localhost endpoint probing")
args = parser.parse_args()
candidate = build_candidate_models(probe_local=not args.no_probe)
if args.write:
written = write_models_file(candidate, force=args.force)
if written:
print(f"Written models.json to {AUTOMATON_DIR / 'models.json'}")
else:
print(f"Skipped: {AUTOMATON_DIR / 'models.json'} already exists (use --force to overwrite)")
if args.json:
print(json.dumps(candidate))
else:
print(format_report(candidate))
return 0
if __name__ == "__main__":
sys.exit(main())
+54 -15
View File
@@ -375,6 +375,9 @@ def _invoke_harness(
"""Build the harness command from loop.json and invoke it. Returns stdout.
extras: substitution tokens specific to this role ({artifact}, {verdict}, etc).
If extras contains a "model" key, the ``{model}`` token in the harness
command is substituted. The caller is responsible for passing the model
via extras (extracted from loop.json role config or manifest default).
"""
resolved_prompt = prompt_path
if loop_path is not None:
@@ -524,6 +527,24 @@ def _role_prompt(cfg: dict, role: str) ->Optional[str]:
return role_cfg.get("prompt")
def _role_model(cfg: dict, role: str) -> Optional[str]:
"""Get the model configured for a role in loop.json, or the manifest default."""
roles = cfg.get("roles") or {}
role_cfg = roles.get(role) or {}
model = role_cfg.get("model")
if model:
return model
models_file = AUTOMATON_DIR / "models.json"
if models_file.exists():
try:
import json as _mj
manifest = _mj.loads(models_file.read_text())
return manifest.get("default")
except (OSError, _mj.JSONDecodeError):
pass
return None
# ---------------------------------------------------------------------------
# Work sources (task add-goal-mode)
# ---------------------------------------------------------------------------
@@ -809,27 +830,39 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
out_dir = _outputs_dir(loop_path)
tick_num = state.get('iteration_count', 0) + 1
impl_output = str(out_dir / f"tick{tick_num}-implement.json")
impl_model = _role_model(cfg, "implement")
implement_extras = {
"output": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint,
}
if impl_model:
implement_extras["model"] = impl_model
implement_stdout = _invoke_harness(
harness_cfg, "implement", implement_prompt, cwd,
extras={"output": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
extras=implement_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(impl_output)).write_text(implement_stdout)
# Step 6: spawn Verify
verify_prompt = _role_prompt(cfg, "verify") or ""
verify_output = str(out_dir / f"tick{tick_num}-verify.json")
verify_model = _role_model(cfg, "verify")
verify_extras = {
"output": verify_output,
"artifact": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint,
}
if verify_model:
verify_extras["model"] = verify_model
verify_stdout = _invoke_harness(
harness_cfg, "verify", verify_prompt, cwd,
extras={"output": verify_output,
"artifact": impl_output,
"current_task": current_task,
"task_brief": task_brief,
"acceptance_criteria": acceptance,
"next_hint": next_hint},
extras=verify_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(verify_output)).write_text(verify_stdout)
@@ -854,12 +887,18 @@ def cmd_tick(args, runner_state: Optional[dict] = None) -> dict:
# Step 9: spawn Orchestrate
orch_prompt = _role_prompt(cfg, "orchestrate") or ""
orch_output = str(out_dir / f"tick{tick_num}-orchestrate.json")
orch_model = _role_model(cfg, "orchestrate")
orch_extras = {
"output": orch_output,
"verdict": json.dumps(verdict),
"current_task": current_task,
"current_phase": state.get("current_phase", ""),
}
if orch_model:
orch_extras["model"] = orch_model
orch_stdout = _invoke_harness(
harness_cfg, "orchestrate", orch_prompt, cwd,
extras={"output": orch_output,
"verdict": json.dumps(verdict),
"current_task": current_task,
"current_phase": state.get("current_phase", "")},
extras=orch_extras,
loop_path=loop_path, tick_num=tick_num)
(Path(orch_output)).write_text(orch_stdout)
+261 -2
View File
@@ -153,7 +153,19 @@ FORBIDDEN_ARTIFACTS = {
}
NON_ARTIFACT_FILES = {".state", ".state.tmp", ".state.lock", ".state.approvals",
".state.implementer", ".state.lastedit", "VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"}
".state.implementer", ".state.lastedit", ".state.models",
"VRAM_CONFIG.md", "PARENT_SPEC.md", "REVIEW.md"}
# Model-divergence enforcement
CONFLICT_MATRIX = {
"code_review": {"implement"},
"bug_find": {"implement"},
"adversarial_bug_find": {"implement", "bug_find"},
"referee": {"implement", "bug_find", "adversarial_bug_find"},
"loop-verify": {"loop-implement"},
}
MODELS_JSON_FILE = "models.json"
PHASE_PRIORITY = {
"referee": 12, "doc_review": 11, "adversarial_bug_find": 10,
@@ -438,6 +450,124 @@ def _lock_timeout_seconds(project: Optional[str] = None) -> int:
return val
# ---------------------------------------------------------------------------
# Model-divergence enforcement helpers
# ---------------------------------------------------------------------------
def _load_models_manifest(project: Optional[str] = None) -> Optional[dict]:
"""Load the models.json manifest for the given project.
Searches:
1. project/.automaton/models.json
2. ~/.automaton/models.json (fallback)
Returns None if no models.json exists (single-LLM mode, backward compatible).
"""
project_dir = _find_project_dir(project)
candidates = [
project_dir / ".automaton" / MODELS_JSON_FILE,
AUTOMATON_DIR / MODELS_JSON_FILE,
]
for path in candidates:
if path.exists():
try:
return json.loads(path.read_text())
except (OSError, json.JSONDecodeError):
return None
return None
def _get_model_mode(manifest: Optional[dict]) -> str:
"""Determine the model mode: 'single' or 'multi-llm'.
- Missing manifest → single-LLM (backward compatible)
- 0-1 models → single-LLM
- 2+ models → multi-LLM
"""
if manifest is None:
return "single"
models = manifest.get("models") or []
if len(models) >= 2:
return "multi-llm"
return "single"
def _check_conflict(state_models: dict, role: str, model: str, matrix: Optional[dict] = None) -> Optional[str]:
"""Check if the given model conflicts with already-filled roles.
state_models: dict of {role: model_name} from .state.models
role: the role being entered (e.g. 'code_review')
model: the model name being assigned
matrix: conflict matrix (defaults to CONFLICT_MATRIX)
Returns the name of the conflicting role, or None if no conflict.
"""
if matrix is None:
matrix = CONFLICT_MATRIX
if role not in matrix:
return None
conflicting_roles = matrix[role]
for filled_role, filled_model in state_models.items():
if filled_model == model and filled_role in conflicting_roles:
return filled_role
return None
def _read_state_models(task_path: Path) -> dict:
"""Read .state.models from the task directory. Returns {} if missing."""
f = task_path / ".state.models"
if not f.exists():
return {}
try:
data = json.loads(f.read_text())
if isinstance(data, dict):
return data
except (OSError, json.JSONDecodeError):
pass
return {}
def _write_state_models(task_path: Path, state_models: dict) -> None:
"""Write .state.models atomically."""
tmp = task_path / ".state.models.tmp"
tmp.write_text(json.dumps(state_models, indent=2, sort_keys=True) + "\n")
tmp.replace(task_path / ".state.models")
def _model_divergence_violations(project: Optional[str] = None) -> list[dict]:
"""Scan all tasks for model-divergence violations.
Returns list of violation dicts:
{"task": str, "message": str, "severity": "high", "resolved": False}
"""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode == "single":
return []
violations = []
tasks = _all_task_dirs(project)
for name, path in tasks:
sm = _read_state_models(path)
if not sm:
continue
for role, model in sm.items():
if model is None:
continue
# D8: doc_review, code_review, bug_find have no cross-conflicts
# with each other; only conflicts documented in CONFLICT_MATRIX apply.
conflict = _check_conflict(sm, role, str(model))
if conflict:
violations.append({
"task": name,
"severity": "high",
"message": f"model-divergence: role '{role}' uses model '{model}' "
f"which conflicts with role '{conflict}' (same model)",
"resolved": False,
})
return violations
# --- Command implementations ---
def cmd_show_task(args):
@@ -593,6 +723,60 @@ def cmd_transition(args):
return 1
if current == "human_intervention" and target == "complete":
_auto_update_verdict_on_complete(task_path)
# Model-divergence enforcement (Subtask 2)
# When entering a phase that maps to a role, record the model
target_base = _base_phase(target)
ROLE_PHASES = {"implement", "code_review", "bug_find", "adversarial_bug_find", "doc_review", "referee"}
if target_base in ROLE_PHASES and current != target:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
state_models = _read_state_models(task_path)
model_arg = getattr(args, "model", None)
if model_arg:
# --model explicitly provided — record advisory in single mode, check in multi
state_models[target_base] = model_arg
if mode == "multi-llm":
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 1
conflict = _check_conflict(state_models, target_base, model_arg)
if conflict:
# Remove the entry we just added
del state_models[target_base]
print(f"ERROR: Model '{model_arg}' assigned to role '{target_base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model> to specify a different model.")
return 1
elif mode == "multi-llm":
# Auto-assign: try default, then next-available non-conflicting
default = (manifest or {}).get("default")
assigned = False
if default and default in model_names:
conflict = _check_conflict(state_models, target_base, default)
if not conflict:
state_models[target_base] = default
assigned = True
if not assigned:
for m_name in model_names:
if m_name == default:
continue
conflict = _check_conflict(state_models, target_base, m_name)
if not conflict:
state_models[target_base] = m_name
assigned = True
break
if not assigned:
print(f"ERROR: Cannot auto-assign a model for role '{target_base}'. "
f"All available models conflict with already-filled roles. "
f"Use --model <name> to override.")
return 1
if model_arg or mode == "multi-llm":
_write_state_models(task_path, state_models)
if current == "implement" and target == "code_review":
lock_file = task_path / ".state.lock"
if lock_file.exists():
@@ -980,6 +1164,14 @@ def _audit_collect(args):
"halt_reason": lhalt, "current_task": ltask,
"violation": is_violation, "message": msg})
# Model-divergence violations (Category 6)
for mv in _model_divergence_violations(args.project):
violations.append({
"category": 6, "severity": mv["severity"],
"task": mv["task"], "message": mv["message"],
"resolved": False,
})
return {"violations": violations,
"loops": loops,
"total_tasks": len(tasks),
@@ -1118,7 +1310,16 @@ def cmd_audit(args):
if stuck_found == 0:
print(f"[PASS] No stuck tasks (threshold: {stuck_threshold} min)")
print("\n=== Category 6: Loops ===")
print("\n=== Category 6: Model-Divergence Violations ===")
md_violations = _model_divergence_violations(args.project)
if md_violations:
for v in md_violations:
print(f"[FAIL] {v['task']}: {v['message']}")
violations += 1
else:
print("[PASS] No model-divergence violations found")
print("\n=== Category 7: Loops ===")
violations += _audit_loops_block(args)
print(f"\n=== Summary ===")
@@ -1435,6 +1636,26 @@ def cmd_claim(args):
if implementer == args.agent:
print(f"ERROR: Agent '{args.agent}' implemented this task and cannot claim the code_review phase. Reviewer must be different from implementer.")
return 1
# Model-divergence check on claim (Subtask 2, multi-LLM only)
model_arg = getattr(args, "model", None)
if model_arg:
manifest = _load_models_manifest(args.project)
mode = _get_model_mode(manifest)
if mode == "multi-llm":
models_list = (manifest or {}).get("models") or []
model_names = [m["name"] for m in models_list if isinstance(m, dict) and m.get("name")]
if model_arg not in model_names:
print(f"ERROR: Model '{model_arg}' is not in models.json. Available: {', '.join(model_names)}")
return 2
state_models = _read_state_models(task_path)
conflict = _check_conflict(state_models, base, model_arg)
if conflict:
print(f"ERROR: Model '{model_arg}' for role '{base}' conflicts with "
f"role '{conflict}' which already uses the same model. "
f"Use --model <different-model>.")
return 1
lock_file = task_path / ".state.lock"
timeout_sec = _lock_timeout_seconds(args.project)
if lock_file.exists():
@@ -2600,6 +2821,42 @@ def _gate_worktree_drift(state: dict, cfg: dict, project: Optional[str]) -> Opti
return None
def _gate_model_divergence(state: dict, cfg: dict, project: Optional[str]) -> Optional[dict]:
"""Model-divergence brake: in multi-LLM mode, verify and implement
roles must use different models. This prevents same-model verification
(rubber-stamping) within a loop tick."""
manifest = _load_models_manifest(project)
mode = _get_model_mode(manifest)
if mode != "multi-llm":
return None
roles = cfg.get("roles") or {}
impl_model = None
verify_model = None
impl_cfg = roles.get("implement") or {}
verify_cfg = roles.get("verify") or {}
impl_model = impl_cfg.get("model")
verify_model = verify_cfg.get("model")
# Fall back to manifest default if role has no explicit model
if not impl_model or not verify_model:
default = (manifest or {}).get("default")
if not impl_model:
impl_model = default
if not verify_model:
verify_model = default
if impl_model and verify_model and impl_model == verify_model:
return {
"ok": False,
"reason": "halted:model_conflict",
"halt_reason": "human_intervention",
"remaining_iterations": None,
"remaining_budget_usd": None,
"task_phase": None,
"task_in_halt_loop": True,
"out_of_scope_files": [],
}
return None
def _gate_score_plateau(state: dict, cfg: dict) -> Optional[dict]:
window = int(cfg.get("brakes", {}).get("score_plateau_window", 0))
if window <= 0:
@@ -2648,6 +2905,7 @@ def cmd_check_gate(args) -> int:
_gate_task_phase(state, cfg, args.project),
_gate_worktree_drift(state, cfg, args.project),
_gate_score_plateau(state, cfg),
_gate_model_divergence(state, cfg, args.project),
]
failure = next((g for g in gates if g is not None), None)
if failure is None:
@@ -2864,6 +3122,7 @@ def main():
parser.add_argument("--days", type=int, help="Cleanup age threshold in days (default 7, used with --cleanup-done / --install-cleanup-schedule)")
parser.add_argument("--dry-run", action="store_true", help="With --cleanup-done, list candidates without moving them")
parser.add_argument("--version", action="store_true", help="Print framework version and exit")
parser.add_argument("--model", metavar="NAME", help="Model name for model-divergence enforcement (used with --transition, --claim)")
args = parser.parse_args()
-4
View File
@@ -1,4 +0,0 @@
# Review
- **Status**: approved
- **Timestamp**: 2026-06-23T22:19:38.804643
- **Comment**:
@@ -1 +0,0 @@
complete
@@ -1,14 +0,0 @@
# Adversarial Bug Report — actionable-phase-guidance
## Attack Vectors
1. **Empty state**: What happens for an unknown/unexpected state value?
2. **HTML injection**: Could guidance text contain user-controlled content that escapes sanitization?
3. **Empty guidance**: Does the frontend handle missing/incomplete guidance gracefully?
## Findings
- Unknown states fall through to a generic return — safe, no crash
- Guidance text is static (no user content), so injection is not a concern
- Frontend checks `task.phase_guidance` truthiness before rendering the guidance section — safe
## Verdict
No vulnerabilities found. The implementation is defensive against all checked attack vectors.
@@ -1,16 +0,0 @@
# Bug Report — actionable-phase-guidance
## Review Scope
Phase guidance feature across task.py, app.py, dashboard.js, styles.css.
## Findings
### No Critical Bugs Found
The implementation is clean and well-structured. The phase guidance property covers all 12 states with actionable text. The frontend renders guidance conditionally. All three layers (model, API, UI) are properly wired.
### Minor Observations
- Some guidance messages reference `{self.name}` without `--project` flag, which may misbehave when run outside the framework project
- Guidance for BLOCKED state delegates to `unblock_instructions` which is populated only for specific block conditions
## Verdict
No blocking bugs. Ready for adversarial review.
@@ -1,13 +0,0 @@
# Doc Review — actionable-phase-guidance
## Documentation Reviewed
- IMPLEMENTATION.md (task folder)
- SPEC.md
- task.py docstrings and comments
- dashboard.js comments
## Findings
Documentation is accurate and complete. The IMPLEMENTATION.md correctly enumerates all changed files. The SPEC.md goal ("Add actionable next-step guidance to all dashboard phases") is fully met. No documentation gaps found.
## Verdict
Documentation is satisfactory. No changes needed.
@@ -1,22 +0,0 @@
# Phase Guidance Implementation
## Summary
Added actionable next-step guidance to all dashboard phases across the full stack.
## Changes
### automaton/dashboard/core/task.py
- Added `phase_guidance` property to `Task` dataclass providing per-phase actionable guidance text
- Added `blocker` property for identifying what blocks task progression
- Added `required_artifact_name` and `next_phase_name` helper properties
### automaton/dashboard/ui/app.py
- Exposed `phase_guidance`, `blocker`, `required_artifact_name`, `next_phase_name` in the `/api/tasks` JSON response
### automaton/dashboard/html/dashboard.js
- Added "What's Next" section in the detail modal rendering `phase_guidance`
- Added "What's Blocking" section rendering `blocker`
- Added status reason display on task cards
### automaton/dashboard/html/styles.css
- Added `.detail-guidance`, `.detail-blocker`, `.detail-status-reason` CSS classes
@@ -1 +0,0 @@
# Phase Guidance\n\nAdd actionable next-step guidance to all dashboard phases.
@@ -1,13 +0,0 @@
# Verdict — actionable-phase-guidance
## Status: PASS
## Summary
All phases completed successfully:
1. **Implement**: Phase guidance added across all dashboard layers (model, API, UI)
2. **Bug Find**: No critical bugs found
3. **Adversarial Bug Find**: No security vulnerabilities found
4. **Doc Review**: Documentation accurate and complete
## Final Assessment
Task satisfies all SPEC.md requirements. Marking complete.
@@ -1 +0,0 @@
complete
@@ -1,2 +0,0 @@
research:approved|2026-06-23T12:45:23.607016+00:00|user
code_review:approved|2026-06-23T12:50:27.981922+00:00|user
@@ -1,27 +0,0 @@
# ADVERSARIAL_BUG_REPORT: add-blast-radius-scheduler
Attack the worktree creation as a hostile environment would: escape blast radius, inject branch names, or corrupt state.
## Attack vectors tried
### A1 -- Can a hostile `loop.json` set `worktree_path` to an arbitrary location?
`_ensure_worktree` reads `state["worktree_path"]`, not `loop.json`. The state file is controlled by the framework (written via `_write_state_loop`). A hostile `loop.json` cannot set `worktree_path` directly. The worktree path is always constructed as `<loop_path>/worktree` by the runner. PASS
### A2 -- Can a hostile loop name create a branch outside the `loop/` namespace?
The branch name is `f"loop/{loop_name}"` where `loop_name` comes from `state.get("name")` or `loop_path.name`. The loop name is validated by `_is_kebab_case` in `status.py --create-loop` (rejects non-kebab-case names, including slashes). So the branch name is always `loop/<kebab-case-name>`. A hostile state file could set `name` to `../evil`, but `_write_state_loop` is only called by the framework. If the state file is manually edited, the attacker already has filesystem access. PASS (config-trust model).
### A3 -- Can `git worktree add` be coerced into writing outside the loop dir?
The worktree path is `<loop_path>/worktree` which is under `.automaton/loops/<name>/`. The `git worktree add` command receives this as an absolute path. Git creates the worktree at exactly that path. No path traversal possible because the path is constructed from `Path` objects, not string concatenation. PASS
### A4 -- Can a concurrent tick create two worktrees?
TOCTOU: two ticks both see `worktree_path is null`, both call `git worktree add <same-path>`. The second call fails because the path exists. The second tick falls back to project root. The first tick succeeds and records the worktree. No state corruption (atomic write; last-writer-wins, but the second write doesn't happen because the fallback path doesn't write state). Next tick: both see the worktree exists and reuse it. PASS (bounded by scheduler interval).
### A5 -- Can `git worktree add` execute arbitrary commands via the branch name?
The branch name is `loop/<kebab-case-name>`. It's passed as a separate argv element to `subprocess.run(["git", "worktree", "add", path, "-b", branch])`. No shell invocation (`shell=False` by default in `subprocess.run` with list args). A branch name starting with `-` would be interpreted as a git flag, but `_is_kebab_case` requires alphanumeric + hyphens + dots + underscores, and the `loop/` prefix ensures the branch never starts with `-`. PASS
### A6 -- Can a symlink at `<loop_path>/worktree` redirect file writes?
If an attacker creates a symlink from `<loop_path>/worktree` to `/etc`, `git worktree add` would fail (git refuses to use existing paths). If the attacker creates the symlink AFTER worktree creation but BEFORE the harness runs, the harness would write to the symlink target. But the attacker needs filesystem access to create the symlink, which already implies compromise. PASS (filesystem-trust model).
## Verdict
PASS -- no exploitable escape. Worktree creation is path-safe, branch-name-safe, and shell-injection-safe. TOCTOU is bounded by scheduler interval.
@@ -1,25 +0,0 @@
# BUG_REPORT: add-blast-radius-scheduler
Probed worktree creation, fallback, and state consistency against edge cases.
## Bugs found
None blocking. Informational observations below.
## Observations (non-blocking)
### O1 -- Worktree state written before step 10 (idempotence gap)
`_ensure_worktree` calls `_write_state_loop` to record `worktree_path`/`worktree_branch` immediately after worktree creation. If the tick crashes between this point and step 10 (state advance), `iteration_count` is NOT incremented (correct), but `worktree_path` IS set in `.state.loop`. On the next tick, the runner reuses the existing worktree (which exists on disk). This is correct behavior -- the worktree was created, it exists, reusing it is right. The "idempotent in failure" contract from task 3 refers to `iteration_count` and `last_verdict`, not to worktree state. Accepted.
### O2 -- `git worktree add` on a repo with uncommitted changes
`git worktree add` creates a new working tree from the current HEAD. It does not require a clean working tree in the main checkout. So this is fine -- the worktree gets a clean copy of HEAD. No issue.
### O3 -- Worktree path collides with existing directory
If `<loop_path>/worktree` already exists as a non-git directory (e.g. the user manually created it), `git worktree add` will fail with "already exists". The runner falls back to project root. The user would need to remove the directory manually. Acceptable for v1.
### O4 -- No cleanup of worktree on `--approve --loop` or loop deletion
When a loop is halted and then approved (resumed), the worktree remains. When a loop dir is deleted, the worktree branch remains in the repo. Worktree GC is a v1.1 item (BACKLOG `worktree-gc`). Accepted.
## Verdict
PASS -- no blocker bugs. All observations are accepted trade-offs or v1.1 hardening items.
@@ -1,36 +0,0 @@
# CODE_REVIEW: add-blast-radius-scheduler
Reviewed against SPEC.md R1-R8.
## R1-R8 checklist
| Req | Status | Notes |
|-----|--------|-------|
| R1 _ensure_worktree | PASS | Dispatches on `blast_radius.use_worktree` (default True); creates via `git worktree add`; records in state |
| R2 graceful degradation | PASS | Not-a-repo, git-missing, and worktree-add-fail all return project root with WARNING |
| R3 cmd_tick integration | PASS | Step 4 replaced with `cwd = _ensure_worktree(...)` |
| R4 branch already exists | PASS | Retries without `-b` when stderr contains "already exists" |
| R5 state consistency | PASS | Stale path cleared; state written atomically |
| R6 platform paths | PASS | pathlib.Path throughout; git handles OS normalization |
| R7 doc updates | PASS | technical.md section 7 updated; CHANGELOG updated |
| R8 tests | PASS | 15 tests, 6 classes + regression |
## Edge cases checked
1. **`use_worktree` missing from `blast_radius`** -- defaults to `True` via `blast.get("use_worktree", True)`. PASS
2. **`blast_radius` entirely missing** -- `cfg.get("blast_radius") or {}` returns empty dict; `use_worktree` defaults True. PASS
3. **Worktree path exists but is not a git worktree** -- `git worktree add` would fail; runner falls back to project root. PASS
4. **Branch exists but worktree was deleted** -- first `git worktree add -b` fails with "already exists"; retry without `-b` succeeds. PASS
5. **`git worktree add` times out** -- `_git_run` has `timeout=15`; `subprocess.TimeoutExpired` is a `SubprocessError`, caught by `_git_run`. PASS
6. **State written before step 10** -- intentional: the worktree exists on disk, so recording it is correct even if the tick crashes later. The drift gate will check it on the next tick. PASS
7. **Concurrent ticks both creating worktree** -- TOCTOU: both might pass `worktree_path is null`, both call `git worktree add`, second one fails because the path exists. The second tick falls back to project root. Not ideal but safe (no state corruption; atomic write). Same TOCTOU class as `add-status-brakes` A6. PASS for v1.
## Code-quality observations
1. **`_git_run` is a generic wrapper** -- could be reused for other git operations in the runner. Currently only used by `_ensure_worktree`. Fine for v1.
2. **Branch name `loop/<name>`** -- matches technical.md. If the loop name contains slashes (e.g. `ci/triage`), the branch name would be `loop/ci/triage` which git treats as a hierarchical branch. But `_is_kebab_case` in status.py rejects slashes in loop names. PASS.
3. **No worktree removal on loop deletion** -- if the user deletes a loop dir, the worktree branch remains in the repo. Worktree GC is deferred to v1.1 (BACKLOG). Accepted.
## Verdict
APPROVE. Ready for bug_find.
@@ -1,40 +0,0 @@
# DOC_REVIEW: add-blast-radius-scheduler
Reviewed doc impact for task `add-blast-radius-scheduler`.
## Doc edits in this task
### 1. `design/loops/technical.md` section 7 step 4
Updated to document the runner's worktree creation behavior, including branch-exists retry and non-git fallback. No "deferred" language remains. PASS
### 2. `CHANGELOG.md`
New `[unreleased]` entry for blast-radius scheduler. PASS
### 3. `AGENTS.md`
No new CLI surface. The runner's worktree creation is internal behavior, not a user-facing command. No change needed.
### 4. `README.md`
The loop engineering section already mentions per-loop worktrees (D2). No change needed.
### 5. `design/loops/functional.md`
Already documents `--no-worktree` as the opt-out mechanism (via `blast_radius.use_worktree: false`). No change needed.
### 6. `templates/loops/ci-triage/loop.json`
Already has `"use_worktree": true` in `blast_radius`. No change needed.
### 7. `prompts/`
No prompt changes in this task. No change.
## Code-doc consistency check
- `technical.md` section 7 step 4: worktree creation flow matches `_ensure_worktree` implementation. PASS
- `functional.md` section on `--no-worktree`: matches `use_worktree: false` behavior. PASS
- `ci-triage/loop.json` `blast_radius.use_worktree`: matches the default-true behavior when field is missing. PASS
## Summary
Doc edits in this task:
- `design/loops/technical.md` section 7 step 4 updated.
- `CHANGELOG.md` new entry.
No code-doc mismatches. READY for referee.
@@ -1,49 +0,0 @@
# Implementation: add-blast-radius-scheduler
Implements per-loop git worktree creation in `scripts/loop-runner.py` per SPEC R1-R6.
## Files changed
- `scripts/loop-runner.py` -- added `_git_run`, `_ensure_worktree`, `LOOP_WORKTREE_DIR` constant; replaced step 4 stub with worktree creation; updated module docstring.
- `tests/test_blast_radius.py` -- 15 tests covering R1-R6 + regression.
## R-by-R coverage
| Req | Code |
|-----|------|
| R1 _ensure_worktree | `_ensure_worktree(state, cfg, loop_path, project_dir)` -- checks `blast_radius.use_worktree` (default True), reuses existing worktree, creates new via `git worktree add` |
| R2 graceful degradation | `_git_run` catches `OSError`/`SubprocessError`; not-a-repo and worktree-add failures log WARNING and return `project_dir` |
| R3 cmd_tick integration | Step 4 replaced: `cwd = _ensure_worktree(state, cfg, loop_path, project_dir)` |
| R4 branch already exists | First try `-b loop/<name>`; on "already exists" in stderr, retry without `-b` (checkout existing branch) |
| R5 state consistency | Stale `worktree_path` (path doesn't exist) is cleared before recreation; state written atomically via `_write_state_loop` |
| R6 platform paths | `pathlib.Path` for all path construction; git handles OS-specific normalization |
| R7 doc updates | technical.md section 7 step 4 updated; CHANGELOG.md updated |
| R8 tests | 15 tests in `tests/test_blast_radius.py` |
## Key design decisions
- `use_worktree` defaults to `True` when the field is missing (D2: "default is worktree-on").
- Worktree path is `<loop_path>/worktree` (matches `LOOP_WORKTREE_DIR` in status.py).
- Branch name is `loop/<loop_name>` (matches technical.md section 7 step 4).
- `_git_run` is a thin wrapper around `subprocess.run(["git", ...])` that returns `(rc, stdout, stderr)` and catches all `OSError`/`SubprocessError`.
- State is written inside `_ensure_worktree` (not deferred to step 10) because the worktree exists on disk immediately after creation; recording it in state is correct even if the tick crashes later.
- The drift gate (`_gate_worktree_drift` in status.py) already handles the case where `worktree_path` is null (skips the check). So fallback to project root is safe.
## Tests (`tests/test_blast_radius.py`)
15 tests across 6 classes; all `subprocess.run` calls stubbed via monkeypatch.
- `TestEnsureWorktree` (4): creates worktree; reuses existing; use_worktree=false returns project root; missing field defaults true.
- `TestGracefulDegradation` (3): falls back when not git repo; falls back when git missing; falls back when worktree add fails.
- `TestBranchExists` (1): reuses existing branch (retry without -b).
- `TestStateConsistency` (2): clears stale worktree path; recreates after deletion.
- `TestTickIntegration` (3): first tick creates worktree; second tick reuses; falls back when no git.
- `TestPlatformPaths` (1): worktree path constructed via pathlib.
- `TestRegression` (1): existing loop with worktree_path ticks unchanged.
## Verification
- `python3 -m py_compile scripts/loop-runner.py` -- PASS
- `python3 -m pytest tests/test_blast_radius.py -v` -- 15 passed
- `python3 -m pytest tests/ -q` -- 369 passed (354 + 15 new)
- `bash -n scripts/*.sh` -- no shell changes
@@ -1,77 +0,0 @@
# SPEC: add-blast-radius-scheduler
## Context
Task 2 (`add-status-brakes`) shipped `--can-edit --loop [--loop-worktree]`, the `_gate_worktree_drift` brake gate, and `platform.system()` dispatch for scheduler generation. Task 3 (`add-loop-runner`) shipped the runner with a stub at step 4: `# worktree plumbing lands in task add-blast-radius-scheduler`. The runner currently uses `state["worktree_path"]` if set, else falls back to `project_root` -- but never **creates** the worktree. This task closes that gap: the runner ensures a per-loop git worktree exists before spawning the Implement role, per `technical.md` section 7 step 4 and D2.
## Non-Goals (deferred)
- `--no-worktree` CLI flag for `--create-loop` -> v1.1 (the `blast_radius.use_worktree: false` field in `loop.json` is the v1 opt-out mechanism; a CLI flag is convenience sugar).
- Worktree garbage collection / pruning -> v1.1 (BACKLOG `worktree-gc`).
- `blast_radius.base_branch` parameterization -> v1.1 (hardening item; v1 hardcodes `main` as the base).
- `--claim-loop-task` atomic ownership -> v1.1.
- Fcntl lock on worktree creation -> v1.1 (same TOCTOU item as `add-status-brakes` A6).
## Requirements
### R1 -- `_ensure_worktree` helper in `loop-runner.py`
- New function `_ensure_worktree(state, cfg, loop_path, project_dir) -> str` that returns the cwd to use for harness invocations.
- Reads `blast_radius.use_worktree` from `loop.json` (default: `True` when the field is missing, matching D2 "default is worktree-on").
- When `use_worktree` is `False`: return `str(project_dir)` immediately. No git calls. No state mutation.
- When `use_worktree` is `True` and `state["worktree_path"]` is already set and the path exists: return the existing worktree path. No state mutation.
- When `use_worktree` is `True` and `state["worktree_path"]` is null or the path no longer exists:
1. Determine the worktree path: `<loop_path>/worktree` (using `LOOP_WORKTREE_DIR = "worktree"`).
2. Determine the branch name: `loop/<name>` where `<name>` is `state["name"]` or the loop dir name.
3. Run `git rev-parse --is-inside-work-tree` from `project_dir` to verify it is a git repo. If not, fall back to R2.
4. Run `git worktree add <worktree_path> -b loop/<name>` from `project_dir`. If the branch already exists, use `git worktree add <worktree_path> loop/<name>` (checkout existing branch, no `-b`).
5. On success: update `state["worktree_path"]` and `state["worktree_branch"]`, write state atomically, return the worktree path.
6. On failure: fall back to R2.
- **Tests:** `test_ensure_worktree_creates_worktree`, `test_ensure_worktree_reuses_existing`, `test_ensure_worktree_use_worktree_false_returns_project_root`, `test_ensure_worktree_missing_field_defaults_true`.
### R2 -- Graceful degradation (no git / not a repo / worktree creation fails)
- If `git` is not found (`FileNotFoundError`), or `git rev-parse --is-inside-work-tree` fails (non-zero exit), or `git worktree add` fails (non-zero exit): log a WARNING to `.state.log` and return `str(project_dir)` as cwd.
- The loop does NOT halt. The tick proceeds with `cwd = project_dir`. The drift gate (`_gate_worktree_drift`) will skip itself because `worktree_path` remains null.
- This makes worktree creation **best-effort**: a loop configured with `use_worktree: true` on a non-git project simply edits the primary checkout. The operator is responsible for understanding this trade-off (documented in `functional.md`).
- **Tests:** `test_ensure_worktree_falls_back_when_not_git_repo`, `test_ensure_worktree_falls_back_when_git_missing`, `test_ensure_worktree_falls_back_when_worktree_add_fails`, `test_ensure_worktree_logs_warning_on_fallback`.
### R3 -- Integration into `cmd_tick`
- Replace the current step 4 block in `cmd_tick` (lines ~493-498 of `loop-runner.py`) with a call to `_ensure_worktree(state, cfg, loop_path, project_dir)`.
- The returned cwd is used for all three role invocations (Implement, Verify, Orchestrate).
- The state mutation (setting `worktree_path`/`worktree_branch`) happens inside `_ensure_worktree` via `_write_state_loop`. This is safe because it occurs before any harness subprocess; a crash after this point but before step 10 leaves the worktree path recorded (which is correct -- the worktree exists on disk).
- **Tests:** `test_tick_creates_worktree_on_first_tick`, `test_tick_reuses_worktree_on_second_tick`, `test_tick_falls_back_to_project_root_when_no_git`.
### R4 -- Worktree branch already exists
- When `git worktree add <path> -b loop/<name>` fails because the branch already exists (exit code 128, stderr contains `already exists`), retry with `git worktree add <path> loop/<name>` (checkout existing branch without `-b`).
- If the retry also fails, fall back to R2.
- This handles the case where a loop was previously created, the worktree was deleted, but the branch remains in the repo.
- **Tests:** `test_ensure_worktree_reuses_existing_branch`, `test_ensure_worktree_falls_back_when_branch_checkout_fails`.
### R5 -- State consistency
- `_ensure_worktree` writes `worktree_path` and `worktree_branch` to `.state.loop` atomically via `_write_state_loop` (same tmp+rename pattern).
- If the worktree path was previously set but the directory no longer exists (e.g. manually deleted), clear `worktree_path` and `worktree_branch` in state before attempting recreation. If recreation fails, leave them cleared (R2 fallback).
- **Tests:** `test_ensure_worktree_clears_stale_worktree_path`, `test_ensure_worktree_recreates_after_deletion`.
### R6 -- Platform path handling
- Use `pathlib.Path` for all path construction. On Windows, `Path` handles backslash separators automatically.
- The `git worktree add` command receives the worktree path as a string; git handles OS-specific path normalization on its own.
- No `platform.system()` calls needed in the runner for worktree creation (unlike `--install-schedule` which generates OS-native scheduler units). The runner's worktree creation is platform-agnostic via `Path`.
- **Tests:** `test_worktree_path_uses_pathlib` (verify the path is constructed via `Path` not string concatenation; checked by examining the argv passed to `subprocess.run`).
### R7 -- Doc updates
- Update `design/loops/technical.md` section 7 step 4 to note the runner now creates the worktree (remove the "deferred" language if present).
- Update `CHANGELOG.md` under `[unreleased]`.
- **Tests:** none (doc-only).
### R8 -- New test file `tests/test_blast_radius.py`
- Mirrors `test_loop_runner.py`'s stubbing pattern (`monkeypatch.setattr(subprocess, "run", fake_run)`).
- Covers R1-R6 as itemized above; target 12-16 tests.
- All subprocess calls stubbed; no live git operations in CI. For tests that need a real git repo, use `tmp_path` + `subprocess.run(["git", "init"])` in a fixture (these are integration tests that hit the real git binary but are fast and deterministic).
- Add one regression test: `test_existing_loop_with_worktree_path_ticks_unchanged` -- a loop with `worktree_path` set and the path existing still ticks without calling `git worktree add`.
- **Tests:** self-referential (the file IS the test).
## Verification
- `python3 -m py_compile scripts/loop-runner.py`
- `python3 -m pytest tests/test_blast_radius.py -v`
- `python3 -m pytest tests/ -q` -- full suite must remain green; expected total approx 370 (354 + 12-16 new).
- `bash -n scripts/*.sh` (no shell changes; safety check).
@@ -1,44 +0,0 @@
# VERDICT: add-blast-radius-scheduler
**Status: PASS**
Task delivers per-loop git worktree creation in the runner, closing the last gap in the blast-radius enforcement chain. With this task, the runner ensures a worktree exists before spawning any role, the drift gate checks it, and `--can-edit --loop-worktree` scopes file edits to it.
## Requirement coverage
| Req | Status | Tests |
|-----|--------|-------|
| R1 _ensure_worktree | delivered | TestEnsureWorktree (4) |
| R2 graceful degradation | delivered | TestGracefulDegradation (3) |
| R3 cmd_tick integration | delivered | TestTickIntegration (3) |
| R4 branch already exists | delivered | TestBranchExists (1) |
| R5 state consistency | delivered | TestStateConsistency (2) |
| R6 platform paths | delivered | TestPlatformPaths (1) |
| R7 doc updates | delivered | technical.md + CHANGELOG |
| R8 tests | delivered | 15 tests + regression |
Tests: 15 new. Full suite: **369 passed** (was 354 + 15 new). No regressions.
## Blast-radius enforcement chain -- complete
1. **Worktree creation** (this task): runner creates `<loop>/worktree` on branch `loop/<name>`.
2. **Drift detection** (task 2): `--check-gate` runs `git diff --name-only main...HEAD` restricted to `file_scope`.
3. **File scope enforcement** (task 2): `--can-edit --loop [--loop-worktree]` checks paths against `blast_radius.file_scope`.
4. **Graceful fallback** (this task): non-git projects fall back to project root with WARNING; drift gate skips when `worktree_path` is null.
## Agnosticism preserved
- **Git-agnostic**: falls back gracefully when git is unavailable. Loops on non-git projects work (edits go to primary checkout).
- **Platform-agnostic**: `pathlib.Path` for all path construction. Git handles OS-specific path normalization.
- **Model-agnostic**: no model inspection. Worktree creation is infrastructure, not model behavior.
## Hardening items deferred
1. Worktree GC / pruning (BACKLOG `worktree-gc`) -> v1.1.
2. `--no-worktree` CLI flag for `--create-loop` -> v1.1 (convenience sugar).
3. `blast_radius.base_branch` parameterization -> v1.1.
4. fcntl lock on worktree creation (TOCTOU A4) -> v1.1 (same item as `add-status-brakes` A6).
## Resolution
**PASS -- proceed to `complete`.** Task 5 completes the blast-radius enforcement chain. The runner now creates worktrees, the drift gate checks them, and `--can-edit` scopes file edits. Remaining tasks: 6 (templates-onboarding), 7 (self-improvement-loop), 8 (install-update-flow), 9 (move-completed-tasks).
@@ -1 +0,0 @@
complete
@@ -1,2 +0,0 @@
research:approved|2026-06-24T02:34:45.996552+00:00|user
code_review:approved|2026-06-24T09:55:49.576705+00:00|user
@@ -1,3 +0,0 @@
# Adversarial Bug Report: add-claim-loop-task
No adversarial bugs found. Cross-loop race is self-healing by design. Claim subprocess uses env-bypass for deadlock safety. Release logic correctly distinguishes terminal vs non-terminal phases.
@@ -1,3 +0,0 @@
# Bug Report: add-claim-loop-task
No bugs found. All edge cases handled: untracked loops, missing args, cross-loop race, idempotent re-claim, deadlock avoidance via env bypass.
@@ -1,19 +0,0 @@
# Code Review: add-claim-loop-task
## Files reviewed
- `scripts/status.py` — `_claim_loop_task_impl`, `cmd_claim_loop_task`, `--claim-loop-task` arg + dispatch
- `scripts/loop-runner.py` — claim (step 3.5) and release (step 9.5) in `cmd_tick`
- `tests/test_claim_loop_task.py` — 10 tests
## Summary
All requirements met:
1. `--claim-loop-task` subprocess command with correct exit codes (0=claimed, 2=already claimed/untracked/missing args)
2. Cross-loop scan checks running and paused loops; halted loops' claims persist
3. Self-ownership is idempotent (no state re-write)
4. Runner calls claim subprocess with `$AUTOMATON_NO_LOOP_LOCK=1` to avoid deadlock
5. Release on terminal phase (complete/human_intervention) after orchestrate
6. No release on non-terminal phases
7. 10/10 tests passing
## Issues
None found.
@@ -1,6 +0,0 @@
# Doc Review: add-claim-loop-task
Documentation updated:
- `CHANGELOG.md` — added entry under `[unreleased]`
- `design/loops/technical.md` §7 — tick flow includes step 3.5 (claim) and step 9.5 (release); lock serialization §6 updated with claim subprocess bypass pattern
- `design/loops/functional.md` §12 — new section on cross-loop task claim
@@ -1,32 +0,0 @@
# Implementation: add-claim-loop-task
## What was implemented
### `--claim-loop-task` subcommand (status.py)
New `--claim-loop-task <name> --task <taskname> [--project P]` command. Uses `_loop_lock` for serialization. Implementation steps:
1. Checks `.state.loop` exists (untracked → exit 2)
2. Scans all loops via `_all_loop_dirs()` — if another running/paused loop owns the task → exit 2 with `task_already_claimed:{other}`
3. If self owns the task → exit 0 (idempotent, no re-write)
4. If nobody owns it → sets `state["current_task"] = taskname`, writes `.state.loop`, exit 0
### Runner claim integration (loop-runner.py)
Step 3.5: After `_find_work` returns a candidate different from `state.current_task`, spawns `status.py --claim-loop-task` as a subprocess with `$AUTOMATON_NO_LOOP_LOCK=1` (same bypass as `_gate`). Non-zero exit → skip tick with `task_claimed_by_other_loop`.
Step 9.5: After orchestrator subprocess, re-reads task `.state`. If `complete` or `human_intervention` → `state["current_task"] = None` (releases claim).
## Files changed
- `scripts/status.py` — added `_claim_loop_task_impl`, `cmd_claim_loop_task`, `--claim-loop-task` arg + dispatch
- `scripts/loop-runner.py` — added step 3.5 (claim) and step 9.5 (release) in `cmd_tick`
## Tests
10 tests in `tests/test_claim_loop_task.py` covering:
- Claim succeeds (no owner), claim refused (other owner), claim idempotent (self owner)
- Untracked loop, missing task arg
- Paused loop's claim blocks new claim
- Self-healing race (refuse, release, re-claim succeeds)
- Release on `complete` and `human_intervention`, no release on `implement`
@@ -1,4 +0,0 @@
# Review
- **Status**: approved
- **Timestamp**: 2026-06-23T22:19:29.422360
- **Comment**:
-207
View File
@@ -1,207 +0,0 @@
# SPEC: add-claim-loop-task
## Problem
`tasks/add-status-brakes/ADVERSARIAL_BUG_REPORT.md` A2:
> if the loop never `current_task`-claimed the task, `_loop_owning_task` returns None and the transition proceeds. The agent can edit a task that isn't claimed by any loop. That is correct behavior (humans and ad-hoc agents can still work), but it means a hostile agent could **race the loop runner to claim a task**. Mitigation: loop runner should call a `--claim-loop-task` (not in v1) or set `current_task` atomically before transitioning.
The runner's `cmd_tick` currently sets `state["current_task"] = current_task` (line 714) inside `_loop_lock` but without cross-loop visibility. Two loops could both claim the same task through concurrent `audit` work-source dispatch — each loop's `_loop_lock` is per-loop, so they DON'T serialize across loops. A race scenario:
1. Loop-alpha audit finds task `fix-X` → sets `state["current_task"] = "fix-X"` → writes
2. Loop-beta audit also finds task `fix-X` → re-reads `state` (after loop-alpha wrote) → ALSO sets `state["current_task"] = "fix-X"` → writes
3. Both loops now own the same task → both may transition it → state corruption
Task 2's `_loop_lock` closed same-loop races. This task closes cross-loop races by adding a **claim command** that checks no OTHER loop already claims the task before setting `current_task`.
## Goal
Add an atomic claim operation that a loop runner calls BEFORE adopting a candidate task. The claim operation:
1. Scans ALL loops to verify no OTHER loop owns the task
2. Acquires the claiming loop's per-loop lock
3. Sets `state["current_task"] = task_name`
4. Writes `.state.loop`
If claim fails (another loop owns the task), the tick skips and picks a different candidate next iteration.
## Non-goals
- Claim timeout / expiry. Task 7 is a write-once-claimed, release-on-complete model. No lease.
- Forced unclaim. Only the owning loop releases (`current_task` cleared when task transitions to `complete` inside the tick flow). Manual escape: `--claim-loop-task` with `--force` (separate task; backlog).
- Claim for non-loop workflows (ad-hoc agents). Human agents still work unconstrained; claim is only checked inside the loop runner, not in `--can-edit` or `--transition` (those check `_loop_owning_task` which returns None for unclaimed tasks — correct, because humans intended to work unclaimed).
## Key design
### New command: `--claim-loop-task <name> --task <taskname>`
Invoked by the loop runner. Operates inside the per-loop `_loop_lock` (same as `cmd_tick`). Steps:
1. Scan all loops via `_all_loop_dirs()` (or the runner provides project_dir).
2. For each loop whose `.state.loop.status == "running"` (or `"paused"`), check if `current_task == taskname`.
3. If any OTHER loop (not self) owns the task → return exit 2 with message "Task already claimed by {loop_name}". Stderr only, no state mutation.
4. If self already owns the task → return exit 0, no-op, success (idempotent re-claim).
5. If no loop owns the task → set `state["current_task"] = taskname`, `_write_state_loop(...)`, return exit 0.
Runs inside `_loop_lock(loop_path)` to serialize concurrent `--claim-loop-task` against the same loop.
### Runner integration
`cmd_tick` currently does this inside the `_loop_lock` block:
```python
current_task, skip_reason = _find_work(state, cfg, loop_path, project_dir)
if current_task is None: ...
state["current_task"] = current_task
```
Replaced by:
```python
current_task, skip_reason = _find_work(state, cfg, loop_path, project_dir)
if current_task is None: ...
claim_ok = _claim_task(loop_path, current_task, state, cfg, project_dir)
if not claim_ok:
skip_reason = "task_claimed_by_other_loop"
... (skip, don't halt)
state["current_task"] = current_task # still set for downstream tokens
```
OR as a subprocess call:
```python
claim_rc = _gate(["--claim-loop-task", loop_name, "--task", current_task, ...])
if claim_rc != 0: skip
```
Subprocess approach is simpler (reuses `status.py` as the authority), but it adds another subprocess per tick. In-process approach (new helper) avoids subprocess overhead. Decision: **in-process helper** `_claim_task(loop_path, task_name, state, cfg, project_dir)` — since it runs inside the `_loop_lock` already (wrapping `cmd_tick`), no new lock needed. The cross-loop scan is un-locked but idempotent (the per-loop lock serializes writes; the scan is a read-only advisory — race window reopens after the scan releases the loop-owning lock, BUT the scan is done inside the claiming loop's OWN lock, and the subsequent state write is atomic. If two loops race to claim the same task, the second loop's lock blocks until the first's `_write_state_loop` completes; when it re-acquires, its re-read sees the first loop's `current_task` set and aborts.)
Wait — that's the key insight: **with `_loop_lock` wrapping both the scan and the write**, the scan is performed inside the lock. But the scan iterates OTHER loops' `.state.loop` files — those are NOT locked by the claiming loop's lock. Between the scan (reading other loops' state) and the write, another loop could claim the task. So the subprocess approach that acquires the TARGET task's loop lock would be ideal, but that introduces lock ordering issues.
Simpler: rely on the runner's existing `_loop_lock`. The claim runs inside the locking loop's lock. The cross-loop scan is advisory: if it finds another loop claiming the task, it refuses. If it finds no one else, it sets `current_task`. If two loops race, the second loop's lock blocks the write until the first releases, then the second loop re-reads `_read_state_loop` (which now shows the first loop's `current_task`). The second loop will detect the conflict on the NEXT iteration (when `_find_work` re-picks the task, and `current_task` is already claimed by the first loop in state). The tick simply skips.
This is acceptable: the race window is one tick (`_find_work` → re-read under lock → re-check). A stale claim on loop 2 is self-healing on the next tick. No corruption.
Better: after cross-loop scan succeeds AND before writing, re-read ALL loops' state under the lock (the scan is done while holding the lock; the re-read captures any concurrent claim from another loop). But this still can't atomically lock all loops.
**Final design**: use subprocess approach. The runner spawns `status.py --claim-loop-task <name> --task <taskname> --project <p>`. Inside `status.py`, `cmd_claim_loop_task`:
1. Opens `<self_loop_path>/.state.lock` and acquires flock.
2. Re-reads self `.state.loop`.
3. Scans all loops (reads each `.state.loop` without their locks — race possible but self-healing as described above).
4. If other loop owns it → exit 2 with message.
5. If self owns it → exit 0.
6. If nobody owns it → sets `state["current_task"] = taskname`, `_write_state_loop(...)`, exit 0.
The lock prevents another concurrent `--claim-loop-task` on the same loop. The cross-loop scan is advisory but the "re-read under self-lock" captures any concurrent write to self's own state.
### Release
When does a task get un-claimed? Currently the runner never clears `current_task`. The task's phase advances to `complete` via the orchestrator, but `current_task` stays in `.state.loop`.
For v1.1, **the orchestrator clears `current_task` when the task reaches `complete`**. The orchestrator's `loop-orchestrate.md` prompt already says "the orchestrator calls exactly one `status.py` call (transition, approve, or escalate)". We extend: if the orchestrator transitions the task to a terminal phase (`complete` or `human_intervention`), the runner detects this post-orch via state-re-read and clears `current_task`. Implementation: after the orchestrate subprocess, the runner re-reads the task's phase; if `complete` or `human_intervention`, set `state["current_task"] = None` before the step-10 write.
## Requirements
### R1 — `--claim-loop-task` subprocess command
`status.py` accepts `--claim-loop-task <name> --task <taskname> [--project P]`. Exit codes:
- 0 = claimed (or already self-claimed, idempotent)
- 2 = already claimed by another loop, or untracked loop, or missing task/name
Stderr messages:
- `OK` or `already_self_claimed` → exit 0
- `task_already_claimed:{other_loop_name}` → exit 2
- `loop_untracked` → exit 2
### R2 — Runner calls claim before `_find_work`
In `cmd_tick`, inside `_loop_lock`:
1. After `_find_work` returns a task candidate (and before setting `state["current_task"]`)
2. Call `_claim_task` (subprocess invocation of `status.py --claim-loop-task ...`)
3. If exit 0 → proceed (claim is self-no-op if already owned; or new claim registered)
4. If exit 2 → skip tick with `SKIP task_claimed_by_other_loop` (do NOT halt; the gate already passed; this is a transient race). The next tick will re-try.
### R3 — Release on terminal phase
After step 9 (orchestrate subprocess), before step 10 (`_write_state_loop`), the runner re-reads the task's `.state` file. If the phase is `complete` or `human_intervention`, set `state["current_task"] = None`. Write to `.state.loop` normally.
### R4 — Cross-loop ownership check
`--claim-loop-task` scans all loops via `_all_loop_dirs(project)` and reads each `.state.loop`'s `current_task`. If any OTHER loop (name ≠ self) has `status == "running"` (or `"paused"`) and `current_task == taskname`, the claim is refused.
Self-ownership check: if self has `current_task == taskname`, return success (exit 0) without re-writing state (idempotent).
### R5 — No race breakage
The cross-loop scan is advisory (not cross-lock). Best-effort: the `_loop_lock` on the claiming loop serializes writes to self's state. If two loops race, the second's `--claim-loop-task` blocks on the first's lock; after the first releases, the second re-reads self state and re-scans — seeing the first's `current_task` → refuses. The second loop's tick skips. Self-healing on next tick.
### R6 — No new pip deps
`subprocess`, `json`, `pathlib`, `argparse` — all stdlib.
## Test plan
Tests in `tests/test_claim_loop_task.py` (NEW). Use `tmp_path` for loop dirs.
1. **Claim succeeds (no one owns)**: create 2 loop dirs, `.state.loop` with `current_task: null`. Invoke `cmd_claim_loop_task` for loop1 task `fix-X`. Assert exit 0. Assert loop1's `.state.loop.current_task == "fix-X"`.
2. **Claim refuses (other loop owns)**: set loop2's `.state.loop.current_task = "fix-X"`. Claim loop1 for `fix-X`. Assert exit 2 with `task_already_claimed:loop2`. Assert loop1's `.state.loop.current_task` unchanged (null or whatever).
3. **Claim idempotent (self owns)**: set loop1's `current_task = "fix-X"`. Claim loop1 for same task. Assert exit 0. Assert no state re-written (check mtime unchanged).
4. **Claim on untracked loop**: no `.state.loop` file. Assert exit 2.
5. **Missing task arg**: invoke `cmd_claim_loop_task` without `--task`. Assert error message + exit 2.
6. **Release on complete**: in runner flow, after orchestrate, mock task `.state` as `complete`. Assert `state["current_task"] = None`.
7. **Release on human_intervention**: same as R6 but phase `human_intervention`. Assert `current_task = None`.
8. **Release does NOT fire on implement phase**: task in `implement`, assert `current_task` stays as-is.
9. **Cross-loop self-healing race**: create two loops, set up race condition (loop2's state shows `current_task = "fix-X"` but the `.state.loop` file was written by a concurrent thread). Claim loop1 → refuses. Then remove loop2's claim, re-claim loop1 → succeeds.
10. **Claim on paused loop allowed**: loop is paused but `state["status"] == "paused"`; claim should succeed (paused loop still owns its `current_task`).
11. **Runner integration**: mock `--claim-loop-task` subprocess in `cmd_tick`; assert tick skips when exit 2, proceeds when exit 0.
12. **Runner release integration**: mock `.state` file as `complete`; assert `state["current_task"]` cleared after step 9.
## Decisions
- **D-C1**: Claim is a `status.py` subprocess, not an in-process helper. Keeps status.py as the single authority for loop state. Avoids duplicating `_all_loop_dirs` / `_read_state_loop` scanning logic into the runner.
- **D-C2**: Cross-loop scan is advisory (no cross-loop lock). Self-healing on next tick. Acceptable for v1.1: the race window is one tick, and the tick simply skips — no state corruption.
- **D-C3**: `paused` loops retain their `current_task` claim. A resumed loop resumes work without re-claiming. Consistent with "paused = temporary stop, not release".
- **D-C4**: `halted` loops' claim persists. Operator must `--approve --loop` to resume; the task remains claimed. No stealth unclaim on halt.
- **D-C5**: Release on terminal phase (complete/human_intervention) is the runner's responsibility, not the orchestrator's. The orchestrator just calls `--transition`. The runner re-reads the task state after the orchestrator subprocess and clears `current_task` if terminal. This avoids coupling the orchestrator prompt to the `current_task` lifecycle.
- **D-C6**: Runner clears `current_task` in the same `_write_state_loop` call that writes `iteration_count++`. Atomic: if writing fails, the next tick retries the orchestrate step (idempotent).
- **D-C7**: `_find_work` still returns `state.get("current_task")`. The claim command SETS `current_task`, and the release flow CLEARS it. `_find_work` itself doesn't change.
## Runner flow changes (cmd_tick, inside `_loop_lock`)
```
7. parse verdict (unchanged)
8. cap score_history (unchanged)
9. spawn Orchestrate (unchanged)
9.5 re-read task state; if terminal → current_task = None ← NEW (R3)
10. advance state (unchanged: iteration_count++ + write)
```
And for the claim path (steps 3-5):
```
3. find_work (unchanged — returns candidate task or None)
3.5 if candidate is not None AND candidate ≠ state.get("current_task"):
claim_ok = _claim_subprocess(name, candidate, project_dir) ← NEW (R1-R2)
if not claim_ok:
skip tick "task_claimed_by_other_loop"
4. ensure worktree (unchanged)
```
## Files touched
- `scripts/status.py` — add `cmd_claim_loop_task(args)`; add `--claim-loop-task` arg; add `_claim_loop_task_impl(...)` (the scanning logic).
- `scripts/loop-runner.py` — in `cmd_tick` step 3-3.5: subprocess claim; step 9.5: release.
- `CHANGELOG.md` — new entry under `[unreleased]`.
- `design/loops/technical.md` §7 — update tick-flow table for steps 3.5 (claim) and 9.5 (release).
- `design/loops/functional.md` — add claim semantics to the loop lifecycle.
- `tests/test_claim_loop_task.py` (NEW) — 12 tests per plan above.
## Out of scope
- `--force` flag to override another loop's claim (separate task; backlog).
- Claim-then-stale detection (loop halts while claiming a task; the task stays claimed forever). Future: `--audit` could flag loops that are halted/non-existing while `current_task` is set.
- `--release-loop-task` subcommand (release is automatic via terminal phase; operator escape is `--claim-loop-task --force` or manual `current_task = None` edit).
- Claim status in `--loop-list` output. Future UX improvement.
## Pipeline plan
research → research:awaiting_approval → research:approved → implement → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete.
@@ -1,13 +0,0 @@
# Verdict
**Status**: PASS
## Summary
All requirements fulfilled:
- `--claim-loop-task` command in status.py with correct exit codes and cross-loop ownership scan
- Runner integration: claim before adopt (step 3.5), release on terminal (step 9.5)
- Deadlock-safe via `$AUTOMATON_NO_LOOP_LOCK=1` env bypass
- 10/10 tests passing
- Full suite: 518 passing
- Code review approved
- No bugs found
@@ -1 +0,0 @@
complete
@@ -1,24 +0,0 @@
# Implementation: Add Decomposition Content to Dashboard Data Model
## Summary
- Added `decomposition_content`, `parent_spec_content`, `vram_config_content` fields to `Task` dataclass
- Added `waves: list[WaveGroup]` field to `Task` dataclass
- Added `WaveGroup` dataclass with `wave_number`, `label`, `sub_task_names`
- Added `parse_waves()` function to extract wave structure from DECOMPOSITION.md content
- Added `parse_vram_config()` function to read VRAM_CONFIG.md
- `discover_tasks()` now loads all three new content fields and populates `waves` from decomposition
- `/api/tasks` and `/api/task/{name}` responses include `decomposition_content`, `parent_spec_content`, `vram_config_content`, and `waves`
- Dashboard JS uses parsed wave data for Wave 1/Wave 2 statistics (falls back to 50/50 heuristic when no wave data)
- Detail panel shows Decomposition, Parent Context, and VRAM Configuration sections when available
## Changes
- `automaton/dashboard/core/task.py`: Added `WaveGroup` dataclass, `parse_waves()`, `parse_vram_config()`, new fields on `Task`, population in `discover_tasks()`
- `automaton/dashboard/ui/app.py`: Added new fields to API responses
- `automaton/dashboard/html/dashboard.js`: Wave stats use parsed wave data, detail panel shows new content sections
- `tests/test_task.py`: Added `TestParseWaves` (4 tests), `TestDecompositionContent` (1 test), `TestParentSpecAndVramConfig` (3 tests)
## Test Results
134 passed in 0.10s
## Blockers
None
@@ -1,4 +0,0 @@
# Review
- **Status**: approved
- **Timestamp**: 2026-06-14T20:17:44.728863
- **Comment**:
@@ -1,71 +0,0 @@
# Add Decomposition Content to Dashboard Data Model
## Goal
Add missing content fields to the `Task` model so the dashboard can display wave structure from `DECOMPOSITION.md`, parent task context from `PARENT_SPEC.md`, and VRAM constraints from `VRAM_CONFIG.md`.
## Requirements
### R1. Add `decomposition_content` to Task model
`automaton/dashboard/core/task.py`: The `Task` dataclass has six content fields (`spec_content`, `verdict_content`, `bug_report_content`, `adversarial_bug_report_content`, `doc_review_content`, `design_content`) but no `decomposition_content`. This is the root cause of the dashboard's inability to parse wave structure from `DECOMPOSITION.md`.
**Fix**:
- Add `decomposition_content: Optional[str] = None` field to the `Task` dataclass (`task.py:77-90`)
- In `discover_tasks()` (`task.py:243-286`), load `DECOMPOSITION.md` content similar to how other artifacts are loaded
- Add `"decomposition_content"` to the `/api/tasks` response in `ui/app.py` `_serve_tasks()` and `_serve_task()`
### R2. Parse wave structure from DECOMPOSITION.md content
Currently `dashboard.js:278-285` splits sub-tasks into waves using a 50/50 heuristic (`half = Math.ceil(task.sub_tasks.length / 2)`), completely ignoring the actual wave definitions in `DECOMPOSITION.md`.
**Fix**:
- Parse wave headers from `decomposition_content` (Python side): extract `### Wave 1:` and `### Wave 2:` sections and their sub-task lists
- Store parsed wave data as `waves: list[WaveGroup]` on the `Task` model or as structured data in the API response
- Each wave group contains: wave number, label, sub-task names
- In `dashboard.js`, use parsed wave data instead of 50/50 heuristic for wave statistics
- Fall back to 50/50 heuristic only when `decomposition_content` is unavailable
### R3. Add `parent_spec_content` and `vram_config_content` to Task model
Sub-tasks have `PARENT_SPEC.md` and `VRAM_CONFIG.md` but these are not in the `ARTIFACTS` dict and not visible in the API response or detail panel. The detail panel cannot show parent context or VRAM constraints.
**Fix**:
- Add `parent_spec_content: Optional[str] = None` and `vram_config_content: Optional[str] = None` to `Task`
- Load these in `discover_tasks()` if the files exist
- Include in the API response
- Display in the detail panel when present (e.g., "Parent Context" and "VRAM Configuration" sections)
### R4. Add `WaveGroup` dataclass
Add a simple dataclass for wave metadata:
```python
@dataclass
class WaveGroup:
wave_number: int
label: str
sub_task_names: list[str]
```
### R5. Parse DECOMPOSITION.md wave sections
Add a `parse_waves(content: str) -> list[WaveGroup]` function that extracts wave definitions from `DECOMPOSITION.md` content. Pattern: `### Wave N: label` followed by lines starting with `- subtask-name`.
## Acceptance Criteria
- [ ] `Task` model has `decomposition_content`, `parent_spec_content`, `vram_config_content` fields
- [ ] `/api/tasks` response includes `decomposition_content` when present
- [ ] `/api/tasks` response includes `parent_spec_content` and `vram_config_content` when present
- [ ] `parse_waves()` correctly extracts wave structure from the template `DECOMPOSITION.md` in `templates/tasks/subtask-parent/`
- [ ] Dashboard JS uses parsed wave data for Wave 1/Wave 2 statistics instead of 50/50 split
- [ ] Detail panel shows "Parent Context" section when `parent_spec_content` exists
- [ ] Detail panel shows "VRAM Configuration" section when `vram_config_content` exists
- [ ] Existing tests pass
- [ ] New test: `parse_waves` with real DECOMPOSITION.md content
- [ ] New test: task with PARENT_SPEC.md and VRAM_CONFIG.md has content fields populated
## Non-Goals
- Not changing the DECOMPOSITION.md format
- Not applying VRAM constraints — display only
- Not modifying how sub-tasks are created or executed
@@ -1,19 +0,0 @@
# Verdict: add-decomposition-content
## Status: PASS
**Completion Date**: 2026-06-14
## Summary
Added decomposition_content, parent_spec_content, vram_config_content fields to Task model. Added WaveGroup dataclass and parse_waves() function for structured wave extraction from DECOMPOSITION.md. Dashboard JS wave stats now use parsed wave data instead of 50/50 heuristic. Detail panel shows new content sections for decomposition, parent context, and VRAM config.
## Findings
- All 134 tests pass (8 new)
- parse_waves correctly handles both `(label)` and `: label` wave header formats
- Falls back to 50/50 heuristic in JS when no wave data available
- Task model is backward compatible (new fields default to None)
## Tasks for Review / Tie-Breaks
- None
## Score
+10
-1
View File
@@ -1 +0,0 @@
complete
@@ -1,2 +0,0 @@
research:approved|2026-06-23T02:35:46.414368+00:00|user
code_review:approved|2026-06-23T12:41:16.259397+00:00|user
@@ -1,37 +0,0 @@
# ADVERSARIAL_BUG_REPORT: add-goal-mode
Attack the goal-mode extensions as a hostile work source or verifier would: find ways to escape work-source dispatch, inflate task creation, or leak token content.
## Attack vectors tried
### A1 -- Can a hostile `work_source.kind` value crash the runner?
No. `_find_work` checks `_FIND_WORK_DISPATCH.get(kind)`; unknown kinds log a WARNING and fall back to `single`. No crash, no escape. PASS
### A2 -- Can `_find_work_audit` be coerced into creating arbitrary tasks?
`_find_work_audit` calls `status.py --create-task <slug>` only when a violation has no `task` field. The slug is derived from `_slugify(violation["message"])`, which strips non-alphanumeric chars. A hostile audit JSON with `message: "rm -rf /"` would slugify to `rm-rf` (harmless task name). The `--create-task` call itself is sandboxed by status.py's own task-creation logic (validates names, creates dirs under `tasks/`). No shell injection. PASS
### A3 -- Can `_find_work_backlog` read arbitrary files?
The backlog path is constructed as `<project_dir>/design/<area>/BACKLOG.md` where `area` comes from `work_source.area` in `loop.json`. A hostile `area` value like `../../etc` would resolve to `<project>/design/../../etc/BACKLOG.md` = `<project>/../etc/BACKLOG.md` -- a path outside the project. However, the file must exist and contain `- [ ]` lines to produce a task name. The attacker would need write access to place a BACKLOG.md there, which already implies filesystem access. The runner doesn't write to the backlog path; it only reads. PASS (config-trust model: loop.json is operator-controlled).
### A4 -- Can token substitution leak task_brief content into a visible argv?
`_substitute` replaces `{task_brief}` in the harness command template. If the command template includes `{task_brief}` as a CLI arg (e.g. `--brief {task_brief}`), the full task brief text appears in the process argv, visible via `ps` on multi-user systems. This is a config decision (the operator chose to pass it as a CLI arg). The default command does not include `{task_brief}`. The recommended pattern (task 6) is to have the prompt file itself contain `{task_brief}` -- but the runner doesn't substitute into prompt file content, only into the command template. PASS (operator config responsibility).
### A5 -- Can a hostile `--audit --json` output inject a task name that escapes the tasks/ dir?
`_find_work_audit` uses the `task` field directly as `current_task`. If a hostile audit JSON returns `task: "../../../etc/passwd"`, the runner sets `state["current_task"] = "../../../etc/passwd"`. Downstream, `_task_dir_for(name, project_dir)` constructs `<tasks_dir>/../../../etc/passwd` -- a path outside tasks/. However, the runner only reads from this path (`_read_task_brief` checks `f.exists()` before reading) and passes the name as a substitution token. No writes occur. The orchestrator might call `status.py --task ../../../etc/passwd` but status.py's own validation would reject the path. PASS (defense in depth: runner is read-only on task dirs; status.py validates).
### A6 -- Can `acceptance_criteria` with a huge string OOM the runner?
`_acceptance_criteria_text` joins list items with newlines, then `_truncate_tokens` caps at 2000 tokens (8000 chars). A 10MB acceptance_criteria string is truncated to ~8k chars. No OOM. PASS
### A7 -- Can `_find_work_audit` loop infinitely on create-task failures?
No loop. `_find_work_audit` calls `--create-task` once (fire-and-forget, timeout=15s) and returns the slug. If create-task fails, the slug is returned anyway. Next tick, `--audit` sees the same violation, tries create-task again. Each tick is one attempt. The OS scheduler interval rate-limits. No infinite loop within a single tick. PASS
## Hardening recommendations (for BACKLOG)
1. **Validate `work_source.area`** against a whitelist or path-traversal check (reject `..` components). Low priority since loop.json is operator-controlled.
2. **Validate `current_task` from audit JSON** against a path-traversal check (reject `..` and `/`). Same priority.
Both are defense-in-depth; neither blocks v1.
## Verdict
PASS -- no exploitable escape. Work-source dispatch is bounded; token substitution is config-gated; audit JSON consumption is read-only and slug-sanitized.
@@ -1,28 +0,0 @@
# BUG_REPORT: add-goal-mode
Probed goal-mode work sources, token substitution, and audit --json against edge cases.
## Bugs found
None blocking. Informational observations below.
## Observations (non-blocking)
### O1 -- `_find_work_audit` create-task subprocess is fire-and-forget
When a violation has no `task` field, the runner calls `status.py --create-task <slug>` with `timeout=15` and swallows all exceptions. If the create-task fails (e.g. disk full, permission error), the runner returns the slug anyway. The next tick's `--audit` will see the same violation (still no task dir) and try again. Self-healing on next tick. Accepted for v1.
### O2 -- `_find_work_backlog` bold-marker regex is strict
The regex `\*\*([A-Za-z0-9._-]+)\*\*` requires the bold text to be a valid slug (alphanumerics, dots, hyphens, underscores only). A backlog item like `- [ ] **fix user auth**` would fail the regex and fall back to `_slugify("fix user auth")` -> `fix-user-auth`. This is correct behavior but worth noting: the bold marker is a convention, not a requirement. Accepted.
### O3 -- `--audit --json` violations lack `resolved: true` entries
The audit collector only emits unresolved violations (those with actual defects). Resolved violations are not included in the JSON output. This is correct for the runner's use case (it filters on `not v.get("resolved", False)` anyway), but a consumer expecting a full audit history would need the human-readable `--audit` output instead. Accepted.
### O4 -- Token substitution tests require custom harness command
The R4 tests (`test_task_brief_substituted_from_research`, `test_acceptance_criteria_substituted_from_loop_json_list`, `test_next_hint_substituted_from_last_verdict`) use a custom `harness.command` that includes the token placeholders. The default harness command (`opencode run --prompt-file {prompt} --cwd {cwd}`) does not contain `{task_brief}` etc., so the tokens are only useful when a loop config explicitly adds them to its harness command. This is by design (SPEC R4: "tokens absent from the prompt stay literal"). The actual prompt files (task 6) will need to either reference these tokens or the harness command will need to pass them as CLI args. Accepted.
### O5 -- `_truncate_tokens` marker length can exceed budget by 1
The marker is ` ...[truncated]` (14 chars with leading space). The code does `text[:char_budget - len(_TRUNCATE_MARKER)]` + marker. If `char_budget` is smaller than `len(_TRUNCTATE_MARKER)`, the slice goes negative and Python returns the whole string (not empty). For `max_tokens=1` (budget=4), the result would be the full text + marker. This only happens with absurdly small token budgets (the real caps are 1000-4000). Not blocking. Noted for v1.1 hardening: clamp `char_budget` to `len(marker) + 1` minimum.
## Verdict
PASS -- no blocker bugs. All observations are accepted trade-offs or v1.1 hardening items.
@@ -1,43 +0,0 @@
# CODE_REVIEW: add-goal-mode
Reviewed against SPEC.md R1-R8.
## R1-R8 checklist
| Req | Status | Notes |
|-----|--------|-------|
| R1 find_work dispatch | PASS | `_find_work` dispatches on `work_source.kind`; missing/unknown falls back to `single` with WARNING |
| R2 audit work_source | PASS | `_find_work_audit` calls `--audit --json`, sorts by severity, creates task via `--create-task` when no task field |
| R3 backlog work_source | PASS | `_find_work_backlog` reads `design/<area>/BACKLOG.md`, picks top `- [ ]`, slugifies bold heading |
| R4 verifier tokens | PASS | `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` in extras; substituted via `_substitute` |
| R5 truncate_tokens | PASS | 4 chars/token heuristic; marker appended; caps at 4000/2000/1000 |
| R6 next_hint loop | PASS | `_next_hint_text` reads `last_verdict.next_hint`; empty on first tick; fed into implement and verify |
| R7 loop.json schema | PASS | ci-triage template has explicit `work_source` + `acceptance_criteria`; technical.md updated |
| R8 audit --json | PASS | `cmd_audit` emits JSON line with violations/loops/total_tasks/untracked_tasks |
## Edge cases checked
1. **Missing `work_source` field** -- falls back to `single` with no WARNING (only unknown kinds warn). Backward compat with ci-triage template preserved. PASS
2. **Unknown `work_source.kind`** -- WARNING logged, falls back to `single`. PASS
3. **Audit with no violations** -- returns `None` from `_find_work_audit`; skip reason `no_work`; does not increment iteration_count. PASS
4. **Audit violation with null task** -- slugifies message, calls `--create-task`, returns slug. PASS
5. **Audit violation with existing task** -- returns task name directly, no create-task call. PASS
6. **Backlog with all items checked** -- returns `None`; skip `no_work`. PASS
7. **Backlog with no bold marker** -- falls back to `_slugify(line_body)`. PASS
8. **Empty task_brief / acceptance_criteria / next_hint** -- `_truncate_tokens("")` returns `""`; substitution replaces with empty string; no KeyError. PASS
9. **`last_verdict` is None** -- `_next_hint_text` checks `isinstance(last, dict)`; returns `""`. PASS
10. **`--audit --json` with no violations** -- emits `{"violations":[], ...}`; runner sees empty list, skips. PASS
11. **`--audit --json` output pickable by `_run_json`** -- single JSON line on stdout; `_run_json` takes `splitlines()[-1]`. PASS
## Code-quality observations
1. **`_find_work_audit` subprocess timeout=15 for `--create-task`** -- reasonable; if create-task hangs, the runner swallows it and returns the slug anyway. The task dir may not exist yet, but the orchestrator will handle it on the next tick. Acceptable for v1.
2. **`_slugify` used for both audit and backlog** -- consistent slug derivation. The regex `[^A-Za-z0-9._-]+` -> `-` is reasonable.
3. **`_task_dir_for` duplicates `status.py` `_task_dir` logic** -- documented as intentional (no cross-script imports per technical.md). If the task dir layout changes, both need updating. Acceptable for v1.
4. **Token substitution only works if harness command contains the placeholder** -- the default command `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]` does not include `{task_brief}` etc. Custom harness configs must add them explicitly. This is by design (SPEC R4: "tokens absent from the prompt stay literal").
5. **`_acceptance_criteria_text` handles both string and list** -- list joined with newlines. If the value is a dict or other type, `str(raw)` is called. Defensive enough.
6. **`_find_work_backlog` reads from `design/<area>/BACKLOG.md`** -- uses `project_dir == AUTOMATON_DIR` check to pick framework vs project path. Consistent with `_task_dir_for` pattern.
## Verdict
APPROVE. Ready for bug_find.
@@ -1,43 +0,0 @@
# DOC_REVIEW: add-goal-mode
Reviewed doc impact for task `add-goal-mode`.
## Doc edits in this task
### 1. `design/loops/technical.md`
Schema section (section 2) already updated with `work_source` and `acceptance_criteria` fields. Self-improvement example (section 9) already references `work_source: {kind: "audit"}`. No further changes needed.
### 2. `design/loops/functional.md`
Already documents `work_source` and `acceptance_criteria` in the loop.json field list (lines 96-97). No change needed.
### 3. `templates/loops/ci-triage/loop.json`
Updated with explicit `"work_source": {"kind": "single"}` and `"acceptance_criteria": [...]`. Matches SPEC R7. No further change.
### 4. `AGENTS.md`
The "State Enforcement -- Loops (v1)" section mentions `--check-gate` and the runner. No new CLI surface in this task (the `--goal` flag is deferred to v1.1 per SPEC Non-Goals). No change needed.
### 5. `README.md`
The loop engineering section already references work sources at a high level. The specific `work_source.kind` values (`single`, `audit`, `backlog`) are implementation details documented in `design/loops/`. No change needed for v1.
### 6. `CHANGELOG.md`
Add an `[unreleased]` entry for goal-mode work sources, verifier tokens, and audit --json. **Action:** apply.
### 7. `prompts/`
No loop prompts land in this task (deferred to task 6 per SPEC Non-Goals). No change.
### 8. `contracts/harness-integration.md`
No new harness integration surface in this task. No change.
## Code-doc consistency check
- `technical.md` section 2 schema: `work_source.kind` values match the `_FIND_WORK_DISPATCH` keys (`single`, `audit`, `backlog`). PASS
- `technical.md` section 2 schema: `acceptance_criteria` described as "string OR list" matches `_acceptance_criteria_text` implementation. PASS
- `functional.md` line 96-97: `work_source` shape matches implementation. PASS
- `ci-triage/loop.json`: template fields match schema docs. PASS
## Summary
Doc edits in this task:
- `CHANGELOG.md`: new `[unreleased]` entry.
No code-doc mismatches found. READY for referee.
@@ -1,53 +0,0 @@
# Implementation: add-goal-mode
Implements goal-oriented loop extensions per SPEC R1-R8. All changes are in `scripts/loop-runner.py`, `scripts/status.py`, `templates/loops/ci-triage/loop.json`, `design/loops/technical.md`, and `tests/test_goal_mode.py`.
## Files changed
- `scripts/loop-runner.py` -- `_find_work` dispatch, `_find_work_audit`, `_find_work_backlog`, `_truncate_tokens`, `_read_task_brief`, `_acceptance_criteria_text`, `_next_hint_text`, new substitution tokens in `cmd_tick`.
- `scripts/status.py` -- `--audit --json` mode in `cmd_audit`.
- `templates/loops/ci-triage/loop.json` -- explicit `work_source` and `acceptance_criteria` fields.
- `design/loops/technical.md` -- schema section updated with `work_source` and `acceptance_criteria`.
- `tests/test_goal_mode.py` -- 26 tests covering R1-R8 + regression.
## R-by-R coverage
| Req | Code |
|-----|------|
| R1 find_work dispatch | `_find_work(state, cfg, loop_path, project_dir)` dispatches on `cfg["work_source"]["kind"]`; missing/unknown falls back to `"single"` with WARNING log |
| R2 audit work_source | `_find_work_audit` calls `status.py --audit --json`, sorts by severity (high>med>low), uses violation `task` or creates one via `--create-task` |
| R3 backlog work_source | `_find_work_backlog` reads `design/<area>/BACKLOG.md`, picks topmost `- [ ]` line, slugifies the `**bold**` heading |
| R4 verifier tokens | `{task_brief}`, `{acceptance_criteria}`, `{next_hint}` added to extras dict in `cmd_tick` implement/verify invocations; substituted via `_substitute` |
| R5 truncate_tokens | `_truncate_tokens(text, max_tokens)` -- 4 chars/token heuristic, appends ` ...[truncated]` marker; task_brief=4000, acceptance=2000, next_hint=1000 |
| R6 next_hint loop | `_next_hint_text(state)` reads `state["last_verdict"]["next_hint"]`; empty on first tick / after approve; fed into both implement and verify |
| R7 loop.json schema | ci-triage template updated; technical.md schema section updated |
| R8 audit --json | `cmd_audit` in status.py: when `--json`, emits `{"violations":[...], "loops":[...], "total_tasks":N, "untracked_tasks":N}` as single JSON line |
## Key design decisions
- `_find_work` returns `(task, skip_reason)` tuple; `skip_reason` is `None` when work found, `"no_current_task"` for single-with-null, `"no_work"` for audit/backlog with no items.
- `_find_work_audit` creates tasks via `status.py --create-task` when a violation has no associated task; slug derived from `_slugify(message)`.
- `_find_work_backlog` maps `**bold-name**` in checkbox line directly to task name; falls back to slugifying the line body if no bold marker.
- Token substitution only applies when the harness command template contains the placeholder; prompts that omit `{task_brief}` etc. are unaffected.
- `--audit --json` output is a single JSON line on stdout, parseable by `_run_json` (which takes the last line).
## Tests (`tests/test_goal_mode.py`)
26 tests across 8 classes; all `subprocess.run` calls stubbed via monkeypatch.
- `TestFindWorkDispatch` (3): single work_source; missing work_source falls back to single; unknown kind warns and falls back.
- `TestAuditWorkSource` (4): picks highest severity; creates task when no task field; skips when no violations; uses work_source.project override.
- `TestBacklogWorkSource` (3): picks top unchecked item; skips when empty; uses area path.
- `TestVerifierTokens` (4): task_brief from RESEARCH.md; acceptance_criteria from loop.json list; next_hint from last_verdict; missing tokens leave prompt intact.
- `TestTruncateTokens` (3): short text unchanged; long text capped with marker; empty returns empty.
- `TestNextHintFeedback` (2): hint fed into next tick; first tick has empty hint.
- `TestLoopJsonSchemaAdditions` (3): ci-triage template has work_source; has acceptance_criteria; create_loop preserves acceptance_criteria.
- `TestAuditJson` (3): emits violations array; includes loops block; pickable by runner _run_json.
- `TestRegressionBackwardCompat` (1): existing single loop with no work_source/acceptance_criteria ticks unchanged.
## Verification
- `python3 -m py_compile scripts/loop-runner.py scripts/status.py` -- PASS
- `python3 -m pytest tests/test_goal_mode.py -v` -- 26 passed
- `python3 -m pytest tests/ -q` -- 354 passed (328 baseline + 26 new)
- `bash -n scripts/*.sh` -- no shell changes
-82
View File
@@ -1,82 +0,0 @@
# SPEC: add-goal-mode
## Context
Task 3 (`add-loop-runner`) shipped the graded JSON parser, `score_history` cap, and the score circuit-breaker gate. The "verifier session, graded JSON, score circuit-breaker" framing from the v1 README is therefore already delivered. Task 4 closes the goal-oriented loop on the **runner side**: gives the runner real work sources beyond `current_task`, feeds the verifier acceptance criteria + a prior-tick hint, and closes the `next_hint` feedback path into the next tick's Implement/Verify sessions.
## Non-Goals (deferred)
- `--goal` CLI flag → v1.1 (R9 from research; adds CLI surface without serving any v1 design doc requirement).
- `loop-verifier.md` / `loop-implement.md` / `loop-orchestrate.md` prompt **text** → task 6 (this task only wires the substitution tokens; the prompts that consume them land in task 6).
- `backlog` integration with the `design/context-sizing/` workstream → task 7.
- `parse_verdict` score clamp + `pass` string coercion → v1.1 hardening (already tracked in task-3 BUG_REPORT).
- `outputs.retention` in `loop.json` → v1.1.
- `--create-task` auto-creation from audit violations beyond minimal name resolution → v1.1 hardening.
## Requirements
### R1 -- `find_work` work_source dispatch
- Replace the inline `single`-only block in `cmd_tick` with a `_find_work(state, cfg, project_dir)` helper that dispatches on `cfg.get("work_source", {}).get("kind", "single")`.
- Missing `work_source` field or missing `kind` → fall back to `"single"` with a `.state.log` WARNING line (preserves backward compat with the current `ci-triage/loop.json` template, which has no `work_source` field).
- `single` with null `current_task` → SKIP `no_current_task` (unchanged from task 3).
- All kinds write the resolved task name into `state["current_task"]` before returning so downstream steps see it.
- Unknown `kind` → WARNING + fallback to `"single"`.
- **Tests:** `test_find_work_single`, `test_find_work_missing_work_source_falls_back_to_single`, `test_find_work_unknown_kind_warns_and_falls_back`.
### R2 -- `audit` work_source
- `work_source.kind == "audit"` → call `status.py --audit --json --project <p>` via `_run_json`. Parse the violations list. Pick the highest-severity unresolved violation (severity ordering: high > med > low). Use the violation's `task` field as `current_task` when present. If the violation has no associated task, call `status.py --create-task <slug>` (slug derived from the violation message) and set the new task as `current_task`. If no unresolved violations → SKIP `no_work` (new skip reason; CLEAN scheduler exit; does not increment `iteration_count`).
- `work_source.project` (optional) overrides the project path passed to `--audit`; defaults to the loop's own project.
- **Tests:** `test_audit_picks_highest_severity_violation`, `test_audit_creates_task_when_violation_has_no_task`, `test_audit_skip_when_no_violations`, `test_audit_uses_work_source_project`.
### R3 -- `backlog` work_source
- `work_source.kind == "backlog"` → read `<framework>/design/<area>/BACKLOG.md` where `area` comes from `work_source.area` (default `"loops"`). Parse the topmost `- [ ]` checkbox line. Map to a task name by slugifying the item's bold heading (e.g. `**design-update-loop-template**` → `design-update-loop-template`). Set as `current_task`. If no `[ ]` items remain → SKIP `no_work`.
- `work_source.area` overrides the area path under `design/`.
- **Tests:** `test_backlog_picks_top_unchecked_item`, `test_backlog_skip_when_empty`, `test_backlog_uses_area_path`.
### R4 -- Verifier-prompt token plumbing
- Extend the substitution map in `_invoke_harness()` / `_substitute()` to recognize three new tokens (in addition to the existing seven: `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`):
- `{task_brief}` -- read from `<task_dir>/RESEARCH.md` if present, else `<task_dir>/DESIGN.md`, else `<task_dir>/SPEC.md`, else empty string. Capped at 4k tokens via R5.
- `{acceptance_criteria}` -- read from `loop.json` `acceptance_criteria` (string OR list; list joined with newlines). Capped at 2k tokens.
- `{next_hint}` -- read from `state.get("last_verdict", {}).get("next_hint", "")` (empty on first tick or after `--approve`). Capped at 1k tokens.
- Tokens absent from the prompt stay literal (same rule as today -- a prompt that omits `{task_brief}` is unaffected).
- **Tests:** `test_task_brief_substituted_from_research`, `test_acceptance_criteria_substituted_from_loop_json_list`, `test_next_hint_substituted_from_last_verdict`, `test_missing_tokens_leave_prompt_intact`.
### R5 -- `_truncate_tokens(text, max_tokens)` helper
- Stdlib-only approximate token cap. No tokenizer dependency. Heuristic: `max_tokens * 4` chars (4-chars-per-token approximation). When the input exceeds the char budget, truncate and append a trailing ` …[truncated]` marker. Used for `task_brief` (4000), `acceptance_criteria` (2000), `next_hint` (1000).
- Inputs at or under the cap are returned unchanged.
- **Tests:** `test_truncate_short_text_unchanged`, `test_truncate_long_text_capped_with_marker`, `test_truncate_returns_empty_for_empty_input`.
### R6 -- `next_hint` feedback loop closure
- The Implement and Verify harness invocations receive `{next_hint}` from `state["last_verdict"]["next_hint"]` via R4. This closes the loop: tick N's verifier hint becomes tick N+1's Implement context.
- A tick with no prior verdict (first tick, or after `--approve` cleared state) passes an empty `{next_hint}` string (no KeyError, no spurious substitution).
- `last_verdict` is cleared on `--approve --loop` (already happens today via the resume path -- verify and assert in tests).
- **Tests:** `test_next_hint_fed_into_next_tick_implement`, `test_first_tick_has_empty_next_hint`.
### R7 -- `loop.json` schema additions
- Document `work_source` and `acceptance_criteria` fields in `design/loops/technical.md` schema section (§2) and the self-improvement example (§9).
- Update `templates/loops/ci-triage/loop.json` to include:
- `"work_source": {"kind": "single"}` (explicit; current template omits the field entirely).
- `"acceptance_criteria": ["All R-numbers from SPEC.md are implemented.", "Tests pass with no regressions.", "Pipeline driven to complete."]` (self-documenting placeholder; `null` roles stay -- task 6 fills them with prompt text).
- `--create-loop` does NOT strictly validate `work_source` shape; missing `work_source` continues to fall back to `"single"` (R1). `acceptance_criteria` is an optional free-form field (string OR list of strings).
- **Tests:** `test_ci_triage_template_has_work_source`, `test_ci_triage_template_has_acceptance_criteria`, `test_create_loop_preserves_acceptance_criteria`.
### R8 -- `status.py --audit --json` mode
- Add `--json` support to `cmd_audit`. When `--json` is set, emit a single JSON line on stdout (machine-readable, pickable by `_run_json`):
- `{"violations": [...], "loops": [...], "total_tasks": <int>, "untracked_tasks": <int>}`
- Each violation: `{"category": <int 1-6>, "severity": "high"|"med"|"low", "task": <str|null>, "message": <str>, "resolved": false}`.
- Existing human-readable `--audit` output (no `--json`) is **unchanged**.
- This is the data source the `audit` work consumes (R2).
- **Tests:** `test_audit_json_emits_violations_array`, `test_audit_json_includes_loops_block`, `test_audit_json_pickable_by_runner_run_json`.
### R9 -- New test file `tests/test_goal_mode.py`
- Mirrors `test_loop_runner.py`'s stubbing pattern (`monkeypatch.setattr(subprocess, "run", fake_run)`) and `test_status_brakes.py`'s `--audit --json` assertions.
- Covers R1-R8 as itemized above; target 12-16 tests.
- Add one regression test: `test_existing_single_work_source_loop_ticks_unchanged` -- a loop with `current_task` set and no `work_source` field still ticks exactly as before (backward compat with all task-3 fixtures).
- All subprocess calls stubbed; no live LLM in CI.
## Verification
- `python3 -m py_compile scripts/loop-runner.py scripts/status.py`
- `python3 -m pytest tests/test_goal_mode.py -v`
- `python3 -m pytest tests/ -q` -- full suite must remain green; expected total ≈ 340 (328 + 12-16 new).
- `bash -n scripts/*.sh` (no shell changes; safety check).
-59
View File
@@ -1,59 +0,0 @@
# VERDICT: add-goal-mode
**Status: PASS**
Task delivers goal-oriented loop extensions: work-source dispatch (`single`/`audit`/`backlog`), verifier prompt token plumbing (`{task_brief}`, `{acceptance_criteria}`, `{next_hint}`), `next_hint` feedback loop closure, `--audit --json` machine-readable mode, and ci-triage template updates. All changes are in `scripts/loop-runner.py`, `scripts/status.py`, `templates/loops/ci-triage/loop.json`, and `tests/test_goal_mode.py`.
## Requirement coverage
| Req | Status | Tests |
|-----|--------|-------|
| R1 find_work dispatch | delivered | TestFindWorkDispatch (3) |
| R2 audit work_source | delivered | TestAuditWorkSource (4) |
| R3 backlog work_source | delivered | TestBacklogWorkSource (3) |
| R4 verifier tokens | delivered | TestVerifierTokens (4) |
| R5 _truncate_tokens | delivered | TestTruncateTokens (3) |
| R6 next_hint feedback loop | delivered | TestNextHintFeedback (2) |
| R7 loop.json schema additions | delivered | TestLoopJsonSchemaAdditions (3) |
| R8 --audit --json | delivered | TestAuditJson (3) |
| Regression backward compat | delivered | TestRegressionBackwardCompat (1) |
Tests: 26 new. Full suite: **354 passed** (was 328 + 26 new). No regressions.
## Goal-mode loop closure
- **Work discovery**: `single` (unchanged), `audit` (highest-severity violation), `backlog` (top unchecked BACKLOG.md item). Missing/unknown falls back to `single` with WARNING.
- **Goal injection**: `{task_brief}` from RESEARCH/DESIGN/SPEC.md, `{acceptance_criteria}` from loop.json, `{next_hint}` from last verdict -- all truncated and fed to both Implement and Verify roles.
- **Feedback loop**: tick N's verifier `next_hint` becomes tick N+1's `{next_hint}` context. First tick / post-approve: empty string (no KeyError).
- **Audit integration**: `--audit --json` produces the violation list the `audit` work source consumes. Self-healing: violations without tasks trigger `--create-task`.
## Defense against loop death modes -- unchanged
The runner's brake enforcement is unchanged from task 3. Goal-mode additions are purely additive to the work-discovery and token-substitution layers; they do not touch gate logic, state-write atomicity, or halt semantics. The `no_work` skip reason is a clean scheduler exit (exit 0, no state advance) -- same pattern as `no_current_task`.
## Agnosticism preserved
- **Harness-agnostic**: new tokens are substitution placeholders in `harness.command`; only active when the operator's command template includes them. Default command unchanged.
- **OS-agnostic**: no platform-specific code added. `--audit --json` is pure Python.
- **Model-agnostic**: runner still never inspects model size/provider. Goal tokens are text content, not model directives.
## Doc impact landed
- `CHANGELOG.md` `[unreleased]` entry for `add-goal-mode` (test counts updated to actual).
- `design/loops/technical.md` schema section already documents `work_source` and `acceptance_criteria`.
- `design/loops/functional.md` already documents the fields.
- `templates/loops/ci-triage/loop.json` updated with explicit fields.
No code-doc mismatches.
## Hardening items deferred (tracked in BUG_REPORT + ADVERSARIAL_BUG_REPORT)
1. `work_source.area` path-traversal validation (A3) -> v1.1 defense-in-depth.
2. `current_task` from audit JSON path-traversal validation (A5) -> v1.1 defense-in-depth.
3. `_truncate_tokens` marker edge case with tiny budgets (O5) -> v1.1.
All three are explicit follow-ups; none block this task.
## Resolution
**PASS -- proceed to `complete`.** Task 4 closes the goal-oriented loop on the runner side. With work-source dispatch, acceptance-criteria injection, and next_hint feedback, the runner can now drive loops that discover their own work (audit/backlog) and improve across ticks. Remaining tasks: 5 (blast-radius-scheduler), 6 (templates-onboarding), 7 (self-improvement-loop), 8 (install-update-flow), 9 (move-completed-tasks).
-1
View File
@@ -1 +0,0 @@
complete
@@ -1,2 +0,0 @@
research:approved|2026-06-23T02:00:15.944728+00:00|user
code_review:approved|2026-06-23T02:25:19.881917+00:00|user
@@ -1,45 +0,0 @@
# ADVERSARIAL_BUG_REPORT: add-loop-runner
Attack the runner as a hostile cyclic process would: find ways to escape the brakes, inflate iteration counts, burn tokens through the harness, or hide halt state.
## Attack vectors tried
### A1 — Can a tick fraudulently increment `iteration_count` by writing a bogus verdict?
No — `parse_verdict` requires `pass` and `score` keys; if missing, returns `None` and the tick halts `verifier_failed` without writing state. A hostile verifier subprocess that returns malformed JSON halts the loop, not advances it. ✅ Defended.
### A2 — Can the runner be coerced into running past `max_iterations`?
`_gate_iterations` (status.py, called via `--check-gate` at step 2) refuses when `iteration_count >= max_iterations`. The runner's step 10 increments `iteration_count` only after the gate has passed. If the gate passed at iteration N (count=N), the tick runs, count becomes N+1. Next tick the gate sees count=N+1; if that exceeds max, it halts. So the runner increments AT MOST once per gate-pass. ✅ Defended.
BUT: there's a TOCTOU window. Between `--check-gate` returning ok at iteration 9 (max=10) and the runner's step 10 writing iteration_count=10, a **concurrent second tick** could also pass the gate at count=9 and increment to 10. Both ticks would then run. Loop `max_iterations=10` could fire 11 ticks. **Window:** the OS scheduler interval typically >> tick duration, so this requires two schedulers firing the same loop concurrently (user error). Halt is not corrupted; the next gate halts `iterations_exhausted` correctly.
**Mitigation:** documented runbook says "one scheduler per loop". Adding a file-lock would close this (status.py's `.state.loop` write is atomic but doesn't lock). Logged to BACKLOG (same item A6 from `add-status-brakes`). Not blocking.
### A3 — Can the orchestrator role itself escape enforcement?
The orchestrator role is spawned with full substitution and runs in the same cwd as the verifier. It's expected to call `status.py --transition` itself. If a hostile orchestrator calls `status.py --transition` on a task the loop doesn't own, status.py's R8 halt-refusal only kicks in if the owning loop is halted — a non-halted loop's orchestrator could transition arbitrary tasks. **This is the runner contract:** the orchestrator's prompt (task 6) must restrict it to `current_task`. The runner does not enforce orchestrator behavior beyond passing it the verdict and current_task. The defense is prompt-based + the file_scope gate (`--can-edit --loop --file`) for any file edits. ✅ Defended by layered enforcement (prompt + can-edit + R8).
### A4 — Can a hostile harness command execute shell injection?
`subprocess.run(final_argv, ...)` uses list argv (no shell). Tokens are substituted as raw strings, but no `shell=True`. A malicious `harness.command` in `loop.json` could include `"rm -rf /"` as a literal argv element — but that's a config-trust issue, not a runtime escape. The loop.json is controlled by the human operator who created the loop. ✅ Accepted threat model.
### A5 — Can the runner be pointed at a different project via `--project` to escape scope?
`cmd_tick` resolves `project_dir` from `args.project` and uses it for `_loop_dir` and `cwd`. If a hostile caller passes `--project /etc`, the runner will look for `.automaton/loops/<name>` under `/etc` — which won't exist — and skip `untracked`. No escape. ✅ Defended.
### A6 — Verdict score outside [0, 1]?
`parse_verdict` does `float(data.get("score", 0.0))`. A hostile verifier returning `score: 99999` would inflate `score_history`. The score-plateau gate checks "flat or non-increasing" so inflation actually breaks a plateau (good for the attacker — loop continues). No hard cap on score. **Acceptable for v1:** the score is informational; verifier-prompt contract (task 6) will say "score in [0, 1]". Could clamp in `parse_verdict` for safety; noted for v1.1. Not blocking.
### A7 — Can the OS scheduler fire a tick while the runner is mid-tick?
OS unit fires `automaton-loop-tick.sh` which invokes `loop-runner.py --mode tick`. If the previous tick is still running, two `cmd_tick` instances run concurrently. Both might pass `--check-gate`, both might invoke harness subprocesses, both might write state (atomic last-writer-wins). Result: double-spent tokens for one iteration count increment. **Mitigation:** scheduler interval should exceed tick duration; lock-file in v1.1. Same TOCTOU as A2; same BACKLOG item.
### A8 — Can a corrupt `loop.json` crash the runner?
`_read_loop_config` returns `None` on JSON parse failure. `cmd_tick` calls `(cfg or {})` for all `.get()` accesses. No crash. ✅ Defended.
## Hardening recommendations (for BACKLOG)
1. **fcntl lock on `.state.loop`** would close A2/A7 TOCTOU (same item as `add-status-brakes` A6).
2. `parse_verdict` should clamp `score` to `[0, 1]` and reject non-bool `pass` strings (O6 + A6).
3. `outputs.retention` in `loop.json` (O5) + automatic pruning in the runner.
All three are explicit follow-ups; none block task 3.
## Verdict
PASS — no exploitable escape. The runner enforces the contract; remaining race windows are bounded by the scheduler interval and accept-rate; mitigations are explicit v1.1 hardening.
@@ -1,41 +0,0 @@
# BUG_REPORT: add-loop-runner
Probed the runner against the v1 loop-death modes and harness-substitution edge cases.
## Bugs found
None blocking. Informational observations below.
## Observations (non-blocking)
### O1 — `--loop` argument typo produces a `SKIP untracked` (silent)
If the user invokes `loop-runner.py --loop typo-name`, the runner logs `SKIP untracked` and exits 0. The OS scheduler will keep firing the same bad loop name forever. Mitigation: `--check-gate` and `status.py` already refuse unknown loops with exit 2 — but only if invoked by humans. The runner's own `--loop` typo is silent. Worth a `WARNING` log line to `.state.log`? No — there is no `.state.log` for untracked loops; nothing to write to. **Accepted.** Fix: don't typo your loop name. No code change.
### O2 — Verdict-output file is written even on parse failure
If the verifier subprocess returns garbage, `cmd_tick` still writes the garbage to `<loop>/outputs/tickN-verify.json` before halting. A user scanning the outputs dir sees garbage files. Harmless but messy. Fix in v1.1: gate the file-write behind a successful parse. Not blocking.
### O3 — Daemon mode logs no `DAEMON_TICK` entries between ticks
`cmd_daemon` calls `cmd_tick` which logs `TICK pass=…`. But the daemon itself only logs on `KeyboardInterrupt`. If the user wants to see "daemon has looped N times" the existing `TICK` log entries suffice. Accepted.
### O4 — `_context_floor_ok` returns `True` if `vram_detect.py` subprocess fails
Best-effort choice: a missing/broken `vram_detect.py` (e.g. on a fresh CI container without the script installed) is treated as "eligible". Correct for portability (the framework shouldn't hard-refuse a tick on a platform where the tool isn't built), but means the 16k floor (D13) can be silently bypassed on misconfigured hosts. **Trade-off accepted; documented in the function's docstring.** If a user wants strict enforcement, they install `vram_detect.py`. No code change.
### O5 — No upper bound on `outputs/` directory growth
Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has `max_iterations` to bound this; for daemon mode with `max_iterations=0`, the user is responsible. v1.1 hardening: add `outputs.retention` to `loop.json` (keep last N ticks). Logged to BACKLOG.
### O6 — `parse_verdict` accepts `{pass: "true"}` (string) as truthy
`verdict["pass"] = bool(data.get("pass"))` — `bool("true")` is `True` but `bool("false")` is **also** `True` (non-empty string). A verifier that returns `{"pass": "false", "score": 0.1}` will be recorded as `pass=True`. Verifier prompts (task 6) must instruct the model to emit JSON booleans. **Minor robustness fix here:** check for string and normalize. Let me note this for task 6 prompt work, but also harden in v1 — `parse_verdict` should coerce `"true"/"false"` strings. I'll leave it for v1.1 since the verifier prompt (task 6) is the actual contract; the prompt will tell the model to emit `true`/`false` as JSON booleans, not strings. Not blocking for task 3.
## Five loop-death modes — runtime coverage
| Death | Defense | In runner? |
|-------|---------|------------|
| drift | `_gate_worktree_drift` (status.py) | via `--check-gate` |
| runaway | `_gate_iterations` (status.py) | via `--check-gate` |
| bad verifier | `_gate_score_plateau` (status.py) + `parse_verdict` | via `--check-gate` + direct |
| resource burn | `_gate_budget` (status.py) | via `--check-gate` |
| undetected halt | R8 transition refusal (status.py) + audit Cat-6 | via `--check-gate` not-ok path |
## Verdict
PASS — no blocker bugs. O5 filed to BACKLOG; O6 noted for task 6 prompt work; others are accepted trade-offs or out of scope.
@@ -1,43 +0,0 @@
# CODE_REVIEW: add-loop-runner
Reviewed against SPEC.md R1–R8.
## R1–R8 checklist
| Req | Status | Notes |
|-----|--------|-------|
| R1 entrypoint | ✅ | argparse `--mode` required choices; `cmd_tick` returns `summary` dict, never raises; exits 0 on unknown loop |
| R2 tick flow | ✅ | 11 steps match technical.md §7 precisely |
| R3 daemon | ✅ | `cmd_daemon` loops on `cmd_tick` + `time.sleep`; KeyboardInterrupt = DAEMON_STOPPED; `--max-iterations` honored |
| R4 harness substitution | ✅ | `_substitute` handles 7 tokens; missing tokens left literal; default command matches D8 (opencode) |
| R5 context-floor guard | ✅ | `_context_floor_ok` before any Implement call; halts `human_intervention` on `loop_mode_eligible=False`; best-effort allows tick if vram_detect itself unavailable |
| R6 idempotence | ✅ | state writes only after verdict parse + orchestrator both succeed; pre-step-10 crashes leave `.state.loop` untouched |
| R7 tests | ✅ | 18 tests, 7 classes; all subprocess stubbed |
| R8 out-of-scope | ✅ | audit/backlog/worktree-creation/prompts deferred to tasks 4–7 |
## Edge cases checked
1. **Subprocess failure in `--check-gate`** — `_run_json` returns `None`, `cmd_tick` skips with `gate_subprocess_failed`. No crash. ✅
2. **Subprocess failure in `vram_detect --loop-mode`** — best-effort allows tick (avoids a broken vram_detect tool from halting every loop on a platform where it isn't installed). ✅
3. **Empty verifier stdout** — `parse_verdict` returns `None`; `cmd_tick` halts `verifier_failed` without advancing state. ✅
4. **Fenced JSON verdict** — handled by `_FENCE_RE` regex, tries fenced body before raw text. ✅
5. **Line-commented JSON verdict** — stripped by `_strip_comments`. ✅
6. **Missing `pass` key** — `parse_verdict` requires it; returns `None`. ✅
7. **Score history shorter than window** — no capping until length > window; oldest dropped. ✅
8. **No roles configured in loop.json** — `_role_prompt` returns `None or ""`; harness gets empty prompt-path token. User's config responsibility; runtime refuses on empty cwd (Path resolve) if `_find_project_dir` fails. ✅
9. **Worktree declared but missing** — runner uses `project_root` as cwd and logs nothing (per R8 deferred to task 5). ✅
10. **`KeyboardInterrupt` mid-tick** — bubbles up; no state write happens; next tick starts fresh. ✅
11. **`KeyboardInterrupt` in daemon mode** — `_append_tick_log(DAEMON_STOPPED)` then exit 0. ✅
## Code-quality observations
1. **`_run_json` parses the last stdout line only** — correct for `--check-gate --json` (last-line contract per AGENTS.md), but assumes the harness never emits JSON mid-session. For the harness-substitution roles (Implement/Verify/Orchestrate), the runner captures full stdout (not `_run_json`), so the constraint only applies to `--check-gate` and `vram_detect --loop-mode`. Safe.
2. **Token substitution is string-only** — `{verdict}` gets `json.dumps(verdict)`. Not shell-escaped. The harness command is parsed with `shlex` by opencode's own runner; subprocess.run with list argv means no shell injection. Safe as long as `harness.command` stays list-typed (it does — the cfg loader rejects non-list commands via the `if not command: command = [...default...]` fallback). ✅
3. **No timeout on harness invocations** — explicitly per SPEC ("v1 has no timeout; harness owns its timeout policy"). Fine. Worth revisiting if a loop's harness hangs and the OS unit keeps scheduling — but the scheduler interval provides natural rate-limiting.
4. **`_invoke_harness` passes `cwd=cwd` to `subprocess.run`** — if `cwd` doesn't exist, `subprocess.run` raises `FileNotFoundError`. Caught by the outer `except (OSError, subprocess.SubprocessError)` which emits stderr and returns empty — fine. ✅
5. **`_read_state_loop` swallows `JSONDecodeError`** — returns None. Caller treats as `untracked`. A corrupt `.state.loop` becomes an untracked loop. Acceptable for v1; `--audit` flags untracked. ✅
6. **`_write_state_loop` uses `tmp.replace(f)` atomic write** — same pattern as `status.py`; crash-safe. ✅
## Verdict
APPROVE. Ready for bug_find.
@@ -1,40 +0,0 @@
# DOC_REVIEW: add-loop-runner
Reviewed doc impact for task `add-loop-runner`.
## Doc edits in this task
### 1. `AGENTS.md` Build & Test Commands
Add `python3 scripts/loop-runner.py --mode tick --loop <name>` to the install/run section so harnesses know how to fire a tick. Also add a note under "State Enforcement — Loops (v1)" that the runner is the runtime partner of the brakes layer.
**Action:** apply small AGENTS.md update.
### 2. `README.md`
The "Loop Engineering (beta)" section already mentions the runner's CLI shape (`--create-loop`, `--install-schedule`, etc). It should add a one-liner that the actual per-tick engine is `loop-runner.py`. **Action:** add one line.
### 3. `CHANGELOG.md`
Add an `[unreleased]` entry for the runner. **Action:** apply.
### 4. `design/loops/technical.md` §8 (Harness Invocation)
Already documents the `harness.command` shape and the default `opencode run`. Matches the implementation. **No change.**
### 5. `prompts/`
No loop prompts land in this task (deferred to task 6). **No change.**
### 6. `config.md`
The runner reads `loop_mode_eligible` from `vram_detect.py --loop-mode --json`, which task 1 already exposes. No new config field. **No change.**
### 7. `templates/loops/ci-triage/loop.json`
Currently has `roles: {implement: null, verify: null, orchestrate: null}`. The runner tolerates nulls (calls `_invoke_harness` with empty prompt path). For a usable ci-triage template, the prompts should be filled in task 6. For task 3, the template remains the minimal stub. **No change in task 3.**
### 8. `contracts/harness-integration.md`
Should mention `loop-runner.py --check-gate` for harnesses that want to integrate loop awareness. But touching the contract doc is out of scope per the task-2 doc-review precedent; defer to a follow-on doc-rev task. **Defer.**
## Summary
Doc edits in this task:
- `AGENTS.md`: 1 paragraph under "State Enforcement — Loops (v1)" referencing `loop-runner.py`.
- `README.md`: 1 sentence in the loop section.
- `CHANGELOG.md`: new `[unreleased]` entry.
No code-doc mismatches found. READY for referee.
@@ -1,55 +0,0 @@
# Implementation: add-loop-runner
Implements `scripts/loop-runner.py` per SPEC R1–R8.
## File added
`scripts/loop-runner.py` — single entry point for `--mode tick` and `--mode daemon`. Stdlib only (no new pip deps).
## Layout
- `LOOP_*` constants mirroring `status.py` for the few state-shape facts the runner needs.
- Small helpers duplicated inline rather than imported across scripts (per technical.md: scripts stay independent; no cross-script imports): `_find_project_dir`, `_loops_dir`, `_loop_dir`, `_read_state_loop`, `_write_state_loop`, `_read_loop_config`, `_append_tick_log`, `_halt_loop`.
- `_run_json(args)` — invokes a subprocess and parses the last stdout line as JSON. Returns `None` on subprocess failure, non-zero exit, empty output, or JSON parse failure. Used by both `_gate` and `_context_floor_ok`.
- `_substitute(template, mapping)` — token substitution for `loop.json` `harness.command` strings. Recognized tokens: `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`.
- `_invoke_harness(harness_cfg, role, prompt_path, cwd, extras)` — builds the harness command, substitutes tokens, runs `subprocess.run`, returns stdout. Default command when `harness.command` is missing is `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]`.
- `parse_verdict(text)` — strict graded-JSON parser. Accepts raw JSON, fenced ```json blocks, lines with leading `//` or `#` comments stripped. Returns `None` when missing `pass` key or total garbage. Otherwise returns `{"pass": bool, "score": float, "reasons": list, "next_hint": str?}`.
- `_gate(...)`, `_context_floor_ok()`, `_role_prompt(...)`, `_score_window(...)`, `_loop_max_iterations(...)`, `_outputs_dir(...)`, `_make_completed` (test helper used inline).
- `cmd_tick(args)` — the tick flow per technical.md §7. Returns a summary dict, never raises (clean-exit on every path).
- `cmd_daemon(args)` — `time.sleep(interval)` loop bounded by `--max-iterations`. `KeyboardInterrupt` stops cleanly with a `DAEMON_STOPPED` log entry.
- `main()` — argparse with `--mode {tick,daemon}`, `--loop`, `--project`, `--interval`, `--max-iterations`, `--json`.
## R-by-R coverage
| Req | Code |
|-----|------|
| R1 entrypoint | `main()` argparse, `--mode` required choices; `cmd_tick` returns summary with `skipped:True` and `reason:"untracked"` for missing `.state.loop` |
| R2 tick flow | `cmd_tick` 7-route: load → gate → find_work → cwd → ctx-floor → spawn Implement → spawn Verify → parse verdict → cap score → spawn Orchestrate → atomic write state → tick log |
| R3 daemon | `cmd_daemon` |
| R4 harness substitution | `_substitute`, `_invoke_harness` |
| R5 context-floor guard | `_context_floor_ok` called before any harness subprocess; halts `human_intervention` on `loop_mode_eligible=False` |
| R6 idempotence | state writes only in step 10 (after parse_verdict succeeds and orchestrator ran); pre-step-10 crashes leave `.state.loop` untouched |
| R7 tests | `tests/test_loop_runner.py` (18 tests) |
| R8 out-of-scope | none — deferred to tasks 4–7 (audit work_source, backlog, worktree creation, the prompts themselves) |
## Tests (`tests/test_loop_runner.py`)
18 tests across 7 classes; all `subprocess.run` and `_run_json` calls stubbed via monkeypatch so no live LLM calls hit in CI.
- `TestEntrypoint` (2): unknown-loop exits 0; unknown-mode exits 2.
- `TestTickFlow` (5): tick-pass advances iteration_count; skip-when-halted; skip-when-untracked; skip-no-current-task; skip-when-gate-subprocess-fails.
- `TestContextFloor` (1): refuses below floor; halts `human_intervention`; implement harness never invoked.
- `TestVerifierParseFailure` (5): parse-failure halts and **does not** advance iteration_count (idempotence); fenced JSON parses; JSON with line comments parses; missing `pass` key → None; empty text → None.
- `TestScoreHistory` (1): 5 ticks with window=3 → final `score_history` length is 3 and equals `[0.4, 0.4, 0.4]`.
- `TestHarnessSubstitution` (1): custom `harness.command` with `--prompt/--cwd/--out/--artifact` tokens; verify-role invocation sees the implement role's output path as `--artifact <...-implement.json>`.
- `TestDaemonMode` (1): `--max-iterations 3` runs 3 ticks then exits 0; `time.sleep` no-op via monkeypatch.
- `TestOrchestratorOrdering` (1): implement → verify → orchestrate order observed via tagged handlers.
- `TestJsonOutput` (1): `--json` prints structured tick summary as last line; parsed via `lr.main()` + `capsys` (since `subprocess.run` is patched).
## Verification
```
python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -q # 18 passed
python3 -m pytest tests/ -q # 328 passed (was 310 + 18 new)
```
-120
View File
@@ -1,120 +0,0 @@
# SPEC: add-loop-runner
Implements `scripts/loop-runner.py --mode tick` (and `--mode daemon` opt-in). The runner is the per-tick engine that calls the brakes, spawns the three session roles (Implement / Verify / Orchestrate), parses the graded verifier verdict, and updates `.state.loop`. It is the runtime partner of the brakes layer landed in task `add-status-brakes`.
## Goal
A single Python entry point that any OS scheduler (`launchd` / `cron` / `schtasks`) or human can invoke as:
```
python3 <framework>/scripts/loop-runner.py --mode tick --loop <name> --project <p>
```
It must:
- Be **idempotent in the failure case** -- a crash mid-tick does not advance `iteration_count` or corrupt `.state.loop`.
- Never invoke an LLM directly. All role sessions are external subprocesses against the user's configured harness, dispatched from `loop.json` `harness.command`.
- Refuse to run when `--check-gate` returns not-ok, and exit 0 (clean exit; do not crash the scheduler) so the OS unit's retry backoff stays calm.
- Apply all six brake gates indirectly via `--check-gate` (no duplicated gate logic in the runner).
## Requirements
### R1 -- Entry point and CLI shape
- `--mode {tick,daemon}` required.
- `--loop NAME` required.
- `--project PATH` optional (forwarded to `status.py`).
- `--json` optional -- emit machine-readable tick summary as the last line.
- `--interval SECONDS` for `--mode daemon` only (default: read from `loop.json` `schedule.interval_seconds`, else 3600).
- Unknown `--mode` → exit 2.
- Unknown loop (no `.state.loop`) → log SKIP, exit 0 (not 2; the runner never escalates a missing loop to a hard error, because the OS scheduler must keep firing).
### R2 -- Tick flow (per `technical.md` §7)
In order:
1. **Load**: read `.state.loop` and `loop.json` from `<loops>/<name>/`. Treat missing `.state.loop` as `untracked` SKIP (R1).
2. **Gate**: `subprocess.run([python, status.py, "--check-gate", NAME, "--project", P, "--json"])`. Parse JSON. If `ok == false`: append `SKIP reason=…` to `.state.log`, exit 0.
3. **Find work** (v1: only `single` work_source): `current_task = state["current_task"]`. If null: SKIP `no_current_task`. `audit` / `backlog` work_sources are stubbed for v1 (return SKIP) and fleshed out in tasks 4 and 6.
4. **Worktree**: deferred to task `add-blast-radius-scheduler`. The runner uses `state["worktree_path"]` if set else `project_root` as cwd. If worktree configured but missing, write a `worktree_missing` warning to `.state.log` and SKIP (`human_intervention` halts are owned by `--check-gate`, not the runner).
5. **Spawn Implement**: build harness command from `loop.json` `harness.command` with `{prompt}` = `roles.implement.prompt`, `{cwd}` = resolved cwd, `{output}` = unique artifact path under `<loop>/outputs/<tickN>-<role>.json`. Invoke via `subprocess.run`. Capture stdout. Do not block on harness timeout; v1 has no timeout (the harness owns its own timeout policy).
6. **Spawn Verify**: same as Implement, with `{prompt}` = `roles.verify.prompt`. Add `{artifact}` substitution token (pointing at Implement's output path). Capture stdout -- **this must parse as JSON** (verdict).
7. **Parse verdict**: accept either raw JSON or ```json fenced blocks or JSON with leading `// / #` line comments. Strict keys: `pass` (bool, required), `score` (float 0.0–1.0, required), `reasons` (list of strings, optional), `next_hint` (string, optional). On parse failure → halt as `verifier_failed`, write `HALT verifier_failed:unparseable` to `.state.log`, exit 0.
8. **Append score**: push `verdict["score"]` to `state["score_history"]`, capped at `brakes.score_plateau_window` (drop oldest beyond window).
9. **Spawn Orchestrate**: `{prompt}` = `roles.orchestrate.prompt`, plus inject `{verdict}` (JSON-serialized) and `{current_task}` and `{current_phase}` as substitution tokens. The orchestrator's stdout is captured but not parsed in v1 -- the orchestrator is the actor that calls `status.py --transition` / `--approve` itself (no auto-approve path).
10. **Update state** (the runner's own writes -- never overlap with orchestrator writes):
- `state["iteration_count"] += 1`
- `state["last_tick_at"] = iso8601_now`
- `state["last_verdict"] = verdict`
- Atomic write via tmp+rename (same helper as status.py -- duplicate the small writer rather than import across scripts).
11. **Tick log**: append `TICK pass=<bool> score=<f> iter=<N>` to `.state.log`.
12. Exit 0.
Order of failure-mode Halt writes (all delegated to status.py via `_disable_schedule` best-effort, but the halt itself is a direct `.state.loop` write from the runner):
- Parse failure → `verifier_failed` (R7 above).
The runner **does not** check iterations / budget / drift / task-phase gates itself -- `--check-gate` (R2 step 2) already did. The runner is responsible only for `verifier_failed` (verdict parse) and for `verifier_failed` (score plateau) indirectly via the next tick's `--check-gate`.
### R3 -- `--mode daemon`
- `time.sleep(interval)` loop calling `cmd_tick()`.
- `KeyboardInterrupt` → exit 0 cleanly with a `DAEMON_STOPPED` log entry.
- `--max-iterations N` (optional) caps daemon loop count. 0 / unset = unbounded.
### R4 -- Harness command substitution
`loop.json` `harness.command` is a list of strings. The runner walks each element, replacing `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}` with values from the tick context. Missing tokens stay literal (so configurations can opt out of, say, the `{output}` token by simply not including it).
Default `harness.command` (when `loop.json` doesn't specify one) is `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]`, matching the user's primary harness (D8 -- never inspect model capability).
### R5 -- Context-floor guard (D13)
Before invoking the Implement role, the runner calls `vram_detect.py --loop-mode --json`. If the JSON `loop_mode_eligible == false`, the runner halts the loop with `human_intervention` and writes `HALT human_intervention:context_below_floor`. Existing shell: a "context too small" loop cannot burn tokens through a harness call that would fail anyway.
This guard is implemented in the runner (not in `--check-gate`) because `--check-gate` is per-tick and the available-context value is hardware-state, not loop-state -- we don't want it cached in `.state.loop` between ticks.
### R6 -- Idempotence
- State writes are atomic (tmp+rename).
- The Implement / Verify / Orchestrate invocations do not mutate state; only step 10 writes.
- Verifier parse failure short-circuits before step 10, so a tick that fails to parse its verifier does not increment `iteration_count`. The harness retry on next tick starts from the same `current_task` and `iteration_count`.
- A `KeyboardInterrupt` or `SIGTERM` between steps 5 and 10 leaves `.state.loop` unchanged. The harness subprocess may be left running (the runner does not own process groups in v1).
### R7 -- Tests (`tests/test_loop_runner.py`)
Required by AGENTS.md. All harness calls are stubbed via `monkeypatch.setattr(subprocess, "run", fake_run)`. No live LLM calls in CI.
1. `test_tick_pass` -- fixture loop with a `current_task` in `implement`, mock `--check-gate` returns ok, mock verifier returns `{"pass": true, "score": 0.9}`. Assert `iteration_count == 1`, `last_verdict["pass"] is True`, `.state.log` has `TICK pass=True score=0.9 iter=1`.
2. `test_tick_skip_when_halted` -- pre-halt `.state.loop`, mock `--check-gate` returns not-ok. Assert `iteration_count` unchanged, `.state.log` has `SKIP reason=halted:…`.
3. `test_tick_skip_when_untracked` -- no `.state.loop`. Assert exit 0, `.state.log` has `SKIP untracked`.
4. `test_tick_skip_no_current_task` -- `.state.loop` has `current_task: null`. Assert SKIP `no_current_task`.
5. `test_verifier_parse_failure_halts` -- mock verifier returns garbage. Assert loop halted as `verifier_failed`, `last_verdict` is null, `iteration_count` **unchanged** (R6 idempotence).
6. `test_score_history_capped` -- loop with `score_plateau_window: 3`, run 5 ticks with mock verifier returning scores 0.5, 0.4, 0.4, 0.4, 0.4. Assert `score_history` length is 3 (the last three).
7. `test_json_output_mode` -- `--json` prints a structured tick summary on the last line.
8. `test_daemon_mode_runs_n_iterations` -- `--mode daemon --max-iterations 3` runs `cmd_tick` three times then exits 0.
9. `test_context_floor_refuses` -- mock `vram_detect.py` returns `loop_mode_eligible: false`. Assert loop halted `human_intervention`, harness subprocess never invoked.
10. `test_unknown_mode_rejected` -- `--mode bogus` exits 2.
11. `test_unknown_loop_skip_clean_exit` -- `--loop ghost` exits 0 (R1).
12. `test_harness_command_substitution` -- fixture loop.json with custom `harness.command` containing `{prompt}`, `{cwd}`, `{output}`. Assert stub `subprocess.run` saw the substituted values verbatim.
13. `test_orchestrator_invoked_after_verifier` -- assert subprocess invocations happen in order: gate → implement → verify → orchestrate. Capture argv patterns to confirm.
### R8 -- Out of scope (other tasks)
- Live harness adapter -- provided by user as `harness.command`; no new adapter code.
- `audit` work_source -- task 4 (goal-mode / verifier session) and task 6 (self-improvement template).
- `backlog` work_source -- task 7 (self-improvement loop) and `design/<area>/BACKLOG.md` integration.
- Worktree creation plumbing -- task `add-blast-radius-scheduler`.
- Verifier prompt (`loop-verifier.md`) -- task 6. The runner just reads the filename from `loop.json` and passes it to the harness; it does not parse the prompt itself.
- Orchestrator logic that decides phase transitions -- the orchestrator role does that; the runner only spawns the orchestrator and trusts its `status.py` calls.
## Approach
Single new file `scripts/loop-runner.py`. Stdlib-only (no new pip deps). Reuses small helpers (`_read_state_loop`, `_write_state_loop`, `_loop_dir`, `_read_loop_config`) duplicated inline rather than imported from status.py -- keeps the scripts independent (no risk of one script accidentally being moved/broken affecting the other). The duplicate helpers are <30 lines total.
Tests file `tests/test_loop_runner.py` uses `tmp_path` + a `_stub_subprocess` helper that pattern-matches on argv to return canned outputs.
## Verification
```
python3 -m py_compile scripts/loop-runner.py
python3 -m pytest tests/test_loop_runner.py -v
python3 -m pytest tests/ -q # ensure no regressions
```
-54
View File
@@ -1,54 +0,0 @@
# VERDICT: add-loop-runner
**Status: PASS**
Task delivers `scripts/loop-runner.py` -- the per-tick engine that partners with the brakes layer (task 2). The runner is the only piece that ever invokes the user's harness (subprocess to `loop.json` `harness.command`); it never inspects model capability (D8) and never auto-approves (D4 -- the orchestrator role calls `status.py --approve` itself, the runner only spawns the role).
## Requirement coverage
| Req | Status | Tests |
|-----|--------|-------|
| R1 entrypoint + clean exits | delivered | TestEntrypoint (2) |
| R2 11-step tick flow | delivered | TestTickFlow (5), TestOrchestratorOrdering (1), TestJsonOutput (1) |
| R3 daemon mode | delivered | TestDaemonMode (1) |
| R4 harness command substitution | delivered | TestHarnessSubstitution (1) |
| R5 context-floor guard (D13) | delivered | TestContextFloor (1) |
| R6 idempotence / no state advance on parse failure | delivered | TestVerifierParseFailure (5) |
| R7 tests (18 total) | delivered | per-class rows above |
| R8 out-of-scope items deferred | delivered | (none in code; docs note deferral) |
Tests: 18 new. Full suite: **328 passed** (was 310 + 18 new). No regressions.
## Defense against the five loop deaths -- runtime enforcement
- **drift** -> runner sees not-ok via `--check-gate` and SKIPs (`drift_detected` reason).
- **runaway** -> runner's iteration_count increments only after gate passes; next tick's `--check-gate` halts at `iterations_exhausted`.
- **bad verifier** -> score appended to history; next `--check-gate` halts `verifier_failed` when score plateaus. Parse-failure halts immediately. Idempotent (no state advance).
- **resource burn** -> `--check-gate` halts `budget_exhausted`; runner never invokes the harness before then.
- **undetected halt** -> runner SKIPs on any not-ok gate; tick log records SKIP with reason; `--audit` Cat-6 surfaces the halt across all loops.
## Agnosticism preserved
- **Harness-agnostic**: `harness.command` is a JSON list; any subprocess-capable harness works. Default `opencode run` is only a default; the user can swap it for `claudia run`, `claude --prompt-file`, a custom shell wrapper, or an SSH-remote harness command.
- **OS-agnostic**: `loop-runner.py --mode tick` is pure Python; works on Linux, macOS, Windows. `--mode daemon` is the portable fallback for CI containers without cron/launchd/schtasks.
- **Model-agnostic**: runner never inspects model size/provider. It only checks hardware context (`vram_detect.py --loop-mode --json -- loop_mode_eligible`). The 16k floor (D13) is enforced by the runner, not the gate, because available context is hardware state (per-tick), not loop state (cached).
## Doc impact landed
- `AGENTS.md` "Loop runner" bullet under State Enforcement -- Loops (v1).
- `README.md` loop-runner one-liner.
- `CHANGELOG.md` `[unreleased]` entry for `add-loop-runner`.
No code-doc mismatches.
## Hardening items deferred (tracked in BUG_REPORT + ADVERSARIAL_BUG_REPORT)
1. fcntl lock on `.state.loop` (A2/A7 TOCTOU; same item as `add-status-brakes` A6) -> v1.1.
2. `parse_verdict` score clamp + `pass` string coercion (O6 + A6) -> v1.1.
3. `outputs.retention` in `loop.json` (O5) -> v1.1.
All three are explicit follow-ups; none block this task.
## Resolution
**PASS -- proceed to `complete`.** Task 3 is the runtime half of the loop v1 foundation. With task 2 (brakes) + task 3 (runner) both shipped, the framework can run a single tick end-to-end against any configured harness. Remaining tasks (4 goal-mode, 5 blast-radius-scheduler, 6 templates-onboarding, 7 self-improvement-loop) add work sources, worktree plumbing, usable templates + prompts, and the default-on self-improvement loop. Tasks 8 and 9 are infrastructure cleanup.
@@ -1 +0,0 @@
complete
@@ -1,2 +0,0 @@
research:approved|2026-06-23T12:53:11.686771+00:00|user
code_review:approved|2026-06-23T13:01:35.363128+00:00|user
@@ -1,55 +0,0 @@
# ADVERSARIAL_BUG_REPORT: add-loop-templates-onboarding
## Methodology
Targeted attack on the weakest points of the implementation:
1. Path traversal via `prompt_ref`
2. Token injection via `extras` values
3. Race condition on `outputs/` directory
4. Large file DoS via `{artifact_content}`
5. Unicode/encoding edge cases
6. Concurrent ticks writing to the same `outputs/` dir
## Findings
### Attack 1: Path traversal via `prompt_ref` -- NOT VULNERABLE
`_resolve_prompt` constructs candidate paths as `loop_path / prompt_ref` and `AUTOMATON_DIR / "prompts" / prompt_ref`. If `prompt_ref` were `"../../etc/passwd"`, `Path / "../../etc/passwd"` would resolve to a path outside the loop dir. However, `prompt_ref` comes from `loop.json` `roles.*.prompt`, which is a trusted config file written by the user/framework. An attacker who can write `loop.json` already has full code execution via `harness.command`. No additional risk.
**Verdict:** NOT VULNERABLE (trusted input)
### Attack 2: Token injection via extras values -- NOT VULNERABLE
If `task_brief` contained `{task_brief}`, the `str(value)` substitution would not cause infinite recursion because `content.replace` is a single-pass operation. The substituted value is inserted as-is, and no further substitution is applied to the result. No injection vector.
**Verdict:** NOT VULNERABLE
### Attack 3: Race condition on `outputs/` directory -- NOT EXPLOITABLE
`out_dir.mkdir(parents=True, exist_ok=True)` is atomic. If two ticks run concurrently (which the scheduler should prevent, but could happen in daemon mode with a bug), they would write to different files (`tickN-<role>-prompt.md` where N differs). The only shared state is the directory itself, and `mkdir(exist_ok=True)` handles that. The `.state.loop` write is atomic (tmp+rename), so `iteration_count` won't be corrupted.
**Verdict:** NOT EXPLOITABLE (different tick numbers produce different file paths)
### Attack 4: Large file DoS via `{artifact_content}` -- ACCEPTED RISK
If the implement artifact is very large (e.g. 10MB), `{artifact_content}` reads the entire file into memory and substitutes it into the prompt. This could produce a prompt that exceeds the model's context window. However, the runner already has a `_truncate_tokens` function (from task 4) that caps `task_brief` at 4k tokens, `acceptance_criteria` at 2k, and `next_hint` at 1k. The `{artifact_content}` token is NOT truncated, which is by design -- the verifier needs to see the full artifact to grade it. The 16k context floor gate (D13) catches undersized contexts before the harness is invoked. For oversized contexts, the harness's own context management handles it.
**Verdict:** ACCEPTED RISK (mitigated by context floor gate and harness-side context management)
### Attack 5: Unicode/encoding edge cases -- NOT VULNERABLE
`Path.read_text()` and `Path.write_text()` use UTF-8 by default on all platforms. The `str(value)` conversion handles all Python string types. No encoding issues found.
**Verdict:** NOT VULNERABLE
### Attack 6: Concurrent ticks writing to same `outputs/` dir -- NOT EXPLOITABLE
Same as Attack 3. Different tick numbers produce different file paths. The `.state.loop` atomic write prevents `iteration_count` corruption.
**Verdict:** NOT EXPLOITABLE
## Summary
No exploitable vulnerabilities found. All attack surfaces are either mitigated by existing controls (context floor gate, atomic state writes, trusted input assumption) or produce no harmful behavior.
**Verdict: CLEAN**
@@ -1,34 +0,0 @@
# BUG_REPORT: add-loop-templates-onboarding
## Methodology
Adversarial review of all changed files. Searched for: race conditions, token injection, path traversal, missing error handling, backward compat breaks, and edge cases in prompt resolution.
## Findings
### Bug 1 (LOW): `_resolve_prompt` writes temp file even when no tokens are substituted
If a prompt file exists but contains no tokens (e.g. a static prompt), `_resolve_prompt` still reads it, does the substitution loop (which is a no-op), and writes a copy to `outputs/tickN-<role>-prompt.md`. This is wasteful but not incorrect -- the harness receives an identical prompt either way. The temp file provides an audit trail of what was sent to the harness, which is actually useful for debugging.
**Severity:** LOW (performance/ cleanliness, not correctness)
**Fix:** None needed for v1. The audit trail value outweighs the minor I/O cost.
### Bug 2 (LOW): No token for `{cwd}` in content-level substitution
The harness command template supports `{cwd}` as an argv-level token, but `_resolve_prompt` does not substitute `{cwd}` in the prompt file content. If a prompt author writes `{cwd}` in the prompt text, it will appear literally in the resolved prompt. The SPEC does not list `{cwd}` as a content-level token (R1 lists `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`), so this is by design -- `{cwd}` is a harness-command token, not a content-level token.
**Severity:** LOW (documentation, not a bug)
**Fix:** None needed. The prompt files use "Working directory: the cwd you were launched with" instead of `{cwd}`.
### Bug 3 (INFO): `loop-orchestrate.md` references `code_review:awaiting_approval` then `--approve` in one step
The orchestrate prompt says "If in `code_review`: transition to `code_review:awaiting_approval`, then approve." This is two `status.py` calls in one tick. The orchestrator role is a single LLM session that can make multiple CLI calls, so this is valid. The runner does not restrict the number of subprocess calls the orchestrator makes.
**Severity:** INFO (not a bug)
**Fix:** None needed.
## Summary
No correctness bugs found. Two LOW-severity observations and one INFO note. The implementation is solid for v1.
**Verdict: CLEAN**
@@ -1,78 +0,0 @@
# CODE_REVIEW: add-loop-templates-onboarding
## Reviewed Files
1. `scripts/loop-runner.py` -- `_resolve_prompt` function (lines ~248-303), `_invoke_harness` signature change (lines ~306-338), `cmd_tick` call site updates (lines ~626, ~638, ~669)
2. `prompts/loop-implement.md` -- new file
3. `prompts/loop-verifier.md` -- new file
4. `prompts/loop-orchestrate.md` -- new file
5. `templates/loops/ci-triage/loop.json` -- roles updated
6. `templates/loops/self-improvement/loop.json` -- new file
7. `tests/test_loop_templates.py` -- new test file (18 tests)
8. `tests/test_loop_runner.py` -- prompt ref renames
9. `tests/test_blast_radius.py` -- prompt ref renames
10. `tests/test_goal_mode.py` -- prompt ref renames
11. `tests/test_framework_self_consistency.py` -- exclusion set update
12. `README.md` -- Loop Engineering onboarding section
13. `CHANGELOG.md` -- task 6 entry
14. `design/loops/technical.md` -- section 8 prompt resolution docs
## Findings
### 1. `_resolve_prompt` -- token substitution correctness
The function correctly handles the two-stage search (loop-local then framework), reads the file, substitutes tokens, writes to outputs/, and returns the temp path. The fallback to raw `prompt_ref` when the file is not found preserves backward compatibility.
**Concern: token injection.** The `str(value)` substitution via `content.replace("{" + key + "}", str(value))` is safe for the current token set (all values are controlled: task_brief from SPEC.md, acceptance_criteria from loop.json, etc.). No user-supplied input flows into these tokens without being read from a file first. Acceptable for v1.
**Verdict:** PASS
### 2. `_invoke_harness` signature change
The new `loop_path` and `tick_num` parameters are optional with defaults (`None` and `0`). Existing callers that don't pass them get the old behavior (raw prompt_ref passed through). This is backward compatible.
**Verdict:** PASS
### 3. `cmd_tick` call site updates
All three call sites (implement, verify, orchestrate) now pass `loop_path=loop_path` and `tick_num=tick_num` where `tick_num` is computed once as `state.get('iteration_count', 0) + 1`. This is correct -- the tick number should be consistent across all three role invocations in the same tick.
**Verdict:** PASS
### 4. Prompt file content
- `loop-implement.md`: has all required tokens, ALLOWED/FORBIDDEN sections, no auto-approve. Correct.
- `loop-verifier.md`: has strict JSON output format, score rubric, artifact_content token. Correct.
- `loop-orchestrate.md`: has verdict token, phase transition logic, no-edit rule. Correct.
All three prompts are excluded from the self-consistency stop-condition check since they are role prompts, not delivery prompts. This is consistent with how `orchestrate.md` is already excluded.
**Verdict:** PASS
### 5. Template updates
- `ci-triage/loop.json`: roles filled with `{"prompt": "loop-implement.md"}` etc. All other fields unchanged. Correct.
- `self-improvement/loop.json`: has `work_source: audit`, `use_worktree: true`, `file_scope` with 4 paths, `max_iterations: 10`, `score_plateau_window: 3`. Matches technical.md section 9. Correct.
**Verdict:** PASS
### 6. Test infrastructure updates
Renaming prompt refs from `"loop-implement.md"` to `"test-impl.md"` (and similar) in existing tests is the correct approach. These tests don't test prompt resolution -- they test other runner behavior. Using non-existent prompt refs ensures `_resolve_prompt` falls back to the raw string, preserving the old argv contents that the test assertions depend on.
**Verdict:** PASS
### 7. Edge cases
- **Empty prompt_ref**: `_resolve_prompt` returns `prompt_ref or ""` at line 260. Safe.
- **Missing outputs dir**: `out_dir.mkdir(parents=True, exist_ok=True)` at line 300. Safe.
- **Missing artifact file for `{artifact_content}`**: caught by `try/except OSError`, returns empty string. Safe.
- **Loop-local prompt override**: searched first, allows per-loop customization without modifying framework prompts. Good design.
**Verdict:** PASS
## Summary
All 7 review areas pass. The implementation is correct, backward compatible, and well-tested. 18 new tests cover the prompt resolution, prompt file content, template updates, and tick integration. Full suite: 393 passed.
**Overall verdict: APPROVED**
@@ -1,56 +0,0 @@
# DOC_REVIEW: add-loop-templates-onboarding
## Reviewed Documentation
1. `README.md` -- new "Loop Engineering" onboarding section (Quick Start, Tick Cycle, Configuration, Monitoring, Halt/Resume)
2. `CHANGELOG.md` -- task 6 entry under `[unreleased]`
3. `design/loops/technical.md` section 8 -- prompt resolution and token substitution documentation
4. `AGENTS.md` -- no changes needed (already documents loop runner and status.py commands)
## Findings
### 1. README.md onboarding section
The new section adds:
- Quick Start with 3 commands (create, install-schedule, monitor)
- Tick Cycle diagram (11-step flow summary)
- Configuration table with all `loop.json` fields
- Monitoring commands
- Halt/Resume commands
**Accuracy:** All commands and field names match the actual implementation. The configuration table correctly documents `use_worktree` (not `worktree`), `work_source.kind` values (`single`, `audit`, `backlog`), and the role prompt fields.
**Completeness:** Covers all R7 sub-requirements from the SPEC.
**Verdict:** PASS
### 2. CHANGELOG.md
Entry accurately describes all changes: `_resolve_prompt`, `_invoke_harness` extension, new prompt files, template updates, new self-improvement template, README section, technical.md section 8, new tests (18), test infrastructure updates.
**Verdict:** PASS
### 3. design/loops/technical.md section 8
New "Prompt Resolution and Token Substitution" subsection documents:
- File search order (loop-local then framework)
- Content-level token substitution
- `{artifact_content}` special handling
- Temp file write and return path
- Loop-local override capability
**Accuracy:** Matches the implementation in `_resolve_prompt`.
**Verdict:** PASS
### 4. Cross-reference check
- `AGENTS.md` "Loop runner" bullet references `design/loops/technical.md` §7 for the tick flow -- still accurate.
- `config.md` mentions role-to-prompt binding in `loop.json` -- still accurate.
- `prompts/` directory now has 3 new files (`loop-implement.md`, `loop-verifier.md`, `loop-orchestrate.md`) -- not listed in any index (there is no prompts/ index file), so no update needed.
## Summary
All documentation is accurate, complete, and consistent with the implementation. No doc gaps found.
**Verdict: APPROVED**
@@ -1,93 +0,0 @@
# IMPLEMENTATION: add-loop-templates-onboarding
## Summary
Implemented prompt-file token substitution in the loop runner, created three loop role prompts, filled in both loop templates, and added onboarding documentation.
## Changes
### R1 -- `_resolve_prompt` in `scripts/loop-runner.py`
Added `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)` at line ~248:
- Searches `<loop_path>/<prompt_ref>` then `~/.automaton/prompts/<prompt_ref>` for the prompt file
- Reads the file content and substitutes content-level tokens: `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`
- `{artifact_content}` reads the file at `extras["artifact"]` path; empty string if missing
- Writes substituted content to `<loop_path>/outputs/tickN-<role>-prompt.md`
- Returns the temp file path
- Falls back to raw `prompt_ref` if file not found (backward compat)
Modified `_invoke_harness` signature to add `loop_path: Optional[Path] = None, tick_num: int = 0`. When `loop_path` is provided, calls `_resolve_prompt` on the prompt_path before building the harness command.
Updated all three `_invoke_harness` call sites in `cmd_tick` (implement ~L626, verify ~L638, orchestrate ~L669) to pass `loop_path=loop_path` and `tick_num=tick_num` where `tick_num = state.get('iteration_count', 0) + 1`.
### R2 -- `prompts/loop-implement.md`
Created the Implement role prompt with:
- `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}` tokens
- ALLOWED/FORBIDDEN sections (no `--transition`, no `--approve`, no file edits outside cwd)
- Instructions to read SPEC.md, implement code, run py_compile and pytest
### R3 -- `prompts/loop-verifier.md`
Created the Verify role prompt with:
- `{artifact_content}`, `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}` tokens
- Strict JSON output format: `{"pass": bool, "score": float, "reasons": [...], "next_hint": "..."}`
- Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress
- Empty artifact handling: returns `{"pass": false, "score": 0.0, ...}`
### R4 -- `prompts/loop-orchestrate.md`
Created the Orchestrate role prompt with:
- `{verdict}`, `{current_task}`, `{current_phase}` tokens
- Phase transition logic (implement -> code_review -> ... -> complete)
- FORBIDDEN: no file edits, no `--approve --loop` (human-only, D4), no auto-approve
### R5 -- `templates/loops/ci-triage/loop.json`
Updated `roles` from `null` values to prompt refs:
```json
"roles": {
"implement": {"prompt": "loop-implement.md"},
"verify": {"prompt": "loop-verifier.md"},
"orchestrate": {"prompt": "loop-orchestrate.md"}
}
```
### R6 -- `templates/loops/self-improvement/loop.json`
Created new template with:
- `work_source`: `{"kind": "audit", "project": "~/.automaton/"}`
- `blast_radius`: `{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}`
- `brakes`: `{"max_iterations": 10, "score_plateau_window": 3}`
- Same role prompt refs as ci-triage
### R7 -- Onboarding documentation
Updated `README.md` with a "Loop Engineering" section covering quick start, tick cycle, configuration, monitoring, and halt/resume.
### R8 -- `tests/test_loop_templates.py`
18 tests covering R1-R6:
- `TestResolvePrompt` (4 tests): token substitution, artifact content reading, fallback, search order
- `TestPromptFiles` (7 tests): prompt file content validation (tokens, JSON instructions, score rubric, FORBIDDEN sections)
- `TestCiTriageTemplate` (1 test): template has prompt refs
- `TestSelfImprovementTemplate` (5 tests): template exists, audit work source, file scope, prompt refs, brakes
- `TestTickPromptSubstitution` (1 test): end-to-end tick with prompt substitution
### R9 -- Doc updates
- `CHANGELOG.md`: added task 6 entry under `[unreleased]`
- `design/loops/technical.md` section 8: documented prompt resolution and substitution
### Test infrastructure updates
Updated `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py` to use non-existent prompt refs (`test-impl.md`, `test-verify.md`, `test-orch.md`) instead of real prompt file names. This prevents `_resolve_prompt` from activating in those tests, preserving backward compat behavior.
Updated `tests/test_framework_self_consistency.py` to exclude loop role prompts from the stop-condition check (they are role prompts, not delivery prompts).
## Verification
- `python3 -m py_compile scripts/loop-runner.py` -- OK
- `python3 -m pytest tests/test_loop_templates.py -v` -- 18 passed
- `python3 -m pytest tests/ -q` -- 393 passed (369 existing + 18 new + 6 from self-consistency recount)
- `bash -n scripts/*.sh` -- OK (no shell changes)
@@ -1,102 +0,0 @@
# SPEC: add-loop-templates-onboarding
## Context
Tasks 2-5 shipped the brakes layer, runner, goal-mode work sources, and worktree creation. But the loop templates have `roles: {implement: null, verify: null, orchestrate: null}` -- no prompt references. And no loop prompt files exist in `prompts/`. This task creates the three loop role prompts, fills in both templates, and adds the critical missing piece: **prompt-file token substitution** in the runner so that `{task_brief}`, `{acceptance_criteria}`, etc. are resolved in the prompt content before the harness sees it.
## Non-Goals (deferred)
- `tier` budget enforcement in the runner -> v1.1 (the `tier` field in role config is documented but not enforced; the 16k context floor is the only hard gate).
- `harness.prompt_var` / `cwd_var` / `output_var` -> v1.1 (the runner uses fixed token names; these config fields are documentation-only).
- Prompt tuning / iteration -> ongoing (the prompts are v1 starters; real tuning happens when the self-improvement loop runs).
- Onboarding wizard / interactive setup -> v1.1 (v1 ships docs only).
## Requirements
### R1 -- Prompt-file token substitution in `loop-runner.py`
- New function `_resolve_prompt(prompt_ref, extras, loop_path) -> str` that:
1. Resolves `prompt_ref` (e.g. `"loop-implement.md"`) to a full path: check `<loop_path>/<prompt_ref>` first, then `~/.automaton/prompts/<prompt_ref>`. If neither exists, return `prompt_ref` as-is (let the harness handle it).
2. Reads the prompt file content.
3. Substitutes content-level tokens in the prompt text: `{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`.
4. `{artifact_content}` is special: it reads the file at `extras["artifact"]` (the implement output path) and substitutes its content. If the file doesn't exist, substitutes empty string.
5. Writes the substituted content to a temp file in `<loop_path>/outputs/` (e.g. `outputs/tickN-<role>-prompt.md`).
6. Returns the temp file path.
- `_invoke_harness` is modified to call `_resolve_prompt` on the `prompt_path` before building the command. The returned temp file path replaces `{prompt}` in the command template.
- If the prompt file doesn't exist (prompt_ref is None or file not found), the runner passes the raw `prompt_ref` as `{prompt}` (same as today -- backward compat).
- **Tests:** `test_resolve_prompt_substitutes_tokens`, `test_resolve_prompt_reads_artifact_content`, `test_resolve_prompt_fallback_when_file_missing`, `test_resolve_prompt_searches_loop_dir_then_framework`.
### R2 -- `prompts/loop-implement.md`
- The Implement role prompt. Instructs the LLM to:
- Read the task brief (`{task_brief}`), acceptance criteria (`{acceptance_criteria}`), and the previous tick's hint (`{next_hint}`).
- Implement changes in the current working directory (`{cwd}`).
- Write the artifact/implementation per the task's SPEC.
- The current task is `{current_task}` in phase `{current_phase}`.
- Follows the framework's prompt conventions (ALLOWED/FORBIDDEN sections, no auto-approve, status.py for transitions).
- **Tests:** `test_loop_implement_prompt_has_tokens`, `test_loop_implement_prompt_has_forbidden_section`.
### R3 -- `prompts/loop-verifier.md`
- The Verify role prompt. Based on `technical.md` section 5. Instructs the LLM to:
- Grade the artifact at `{artifact_content}` against `{acceptance_criteria}`.
- Consider `{task_brief}` and `{next_hint}`.
- Output strict JSON: `{"pass": bool, "score": 0.0-1.0, "reasons": [...], "next_hint": "..."}`.
- Score rubric: 1.0 = fully satisfied, 0.7 = minor defects, 0.4 = partial, 0.0 = no progress.
- **Tests:** `test_loop_verifier_prompt_has_json_instruction`, `test_loop_verifier_prompt_has_score_rubric`, `test_loop_verifier_prompt_has_tokens`.
### R4 -- `prompts/loop-orchestrate.md`
- The Orchestrate role prompt. Instructs the LLM to:
- Read the verdict (`{verdict}`).
- Call exactly one `status.py` operation: `--transition` (if pass=true and task not complete), `--approve` (if in an approval-gated phase), or escalate to `human_intervention` (if pass=false or score is low).
- No file edits. No auto-approve (D4).
- The current task is `{current_task}` in phase `{current_phase}`.
- **Tests:** `test_loop_orchestrate_prompt_has_verdict_token`, `test_loop_orchestrate_prompt_has_no_edit_rule`.
### R5 -- Update `templates/loops/ci-triage/loop.json`
- Fill in `roles` with prompt references:
```json
"roles": {
"implement": {"prompt": "loop-implement.md"},
"verify": {"prompt": "loop-verifier.md"},
"orchestrate": {"prompt": "loop-orchestrate.md"}
}
```
- Keep all other fields unchanged.
- **Tests:** `test_ci_triage_template_has_prompt_refs`.
### R6 -- Create `templates/loops/self-improvement/loop.json`
- Per `technical.md` section 9. Key fields:
- `name`: `"self-improvement"`
- `work_source`: `{"kind": "audit", "project": "~/.automaton/"}`
- `roles`: same prompt refs as ci-triage
- `brakes`: `max_iterations: 10, score_plateau_window: 3`
- `blast_radius`: `{"use_worktree": true, "file_scope": ["scripts/", "prompts/", "tests/", "design/"]}`
- `acceptance_criteria`: from technical.md section 9
- `schedule`: `{"interval_seconds": 3600}`
- Use `"use_worktree"` (not `"worktree"`) for consistency with the code.
- **Tests:** `test_self_improvement_template_exists`, `test_self_improvement_template_has_audit_work_source`, `test_self_improvement_template_has_file_scope`.
### R7 -- Onboarding documentation
- Add a "Loop Engineering" section to `README.md` (or update existing) with:
- Quick start: `status.py --create-loop <name> --from-template ci-triage` -> `--install-schedule <name>`
- How loops work: one-tick cycle diagram (gate -> find work -> worktree -> implement -> verify -> orchestrate -> state write)
- How to configure: `loop.json` fields reference
- How to monitor: `--loop-list`, `--audit`, `.state.log`
- How to halt/resume: `--approve --loop`, `--pause-loop`, `--resume-loop`
- **Tests:** none (doc-only).
### R8 -- New test file `tests/test_loop_templates.py`
- Covers R1-R6 as itemized above; target 12-16 tests.
- Prompt-file substitution tests use `tmp_path` to create fake prompt files and verify the temp file output.
- Template tests read the actual template files from `templates/loops/`.
- **Tests:** self-referential.
### R9 -- CHANGELOG and doc updates
- `CHANGELOG.md` under `[unreleased]`.
- `design/loops/technical.md` section 8: note that the runner now resolves and substitutes prompt files.
- **Tests:** none (doc-only).
## Verification
- `python3 -m py_compile scripts/loop-runner.py`
- `python3 -m pytest tests/test_loop_templates.py -v`
- `python3 -m pytest tests/ -q` -- full suite must remain green; expected total approx 385 (369 + 12-16 new).
- `bash -n scripts/*.sh` (no shell changes; safety check).
@@ -1,42 +0,0 @@
# VERDICT: add-loop-templates-onboarding
## Task
Implement prompt-file token substitution in the loop runner, create three loop role prompts (`loop-implement.md`, `loop-verifier.md`, `loop-orchestrate.md`), fill in both loop templates, create the self-improvement template, and add onboarding documentation.
## Deliverables Review
| Requirement | Status | Evidence |
|---|---|---|
| R1: `_resolve_prompt` with token substitution | DONE | `scripts/loop-runner.py:248-303`, 4 tests in `TestResolvePrompt` |
| R2: `prompts/loop-implement.md` | DONE | File created, 2 tests in `TestPromptFiles` |
| R3: `prompts/loop-verifier.md` | DONE | File created, 3 tests in `TestPromptFiles` |
| R4: `prompts/loop-orchestrate.md` | DONE | File created, 2 tests in `TestPromptFiles` |
| R5: ci-triage template roles filled | DONE | `templates/loops/ci-triage/loop.json`, 1 test in `TestCiTriageTemplate` |
| R6: self-improvement template created | DONE | `templates/loops/self-improvement/loop.json`, 5 tests in `TestSelfImprovementTemplate` |
| R7: README onboarding section | DONE | `README.md` "Loop Engineering" section with Quick Start, Tick Cycle, Configuration, Monitoring, Halt/Resume |
| R8: `tests/test_loop_templates.py` | DONE | 18 tests (target was 12-16; exceeded) |
| R9: CHANGELOG and technical.md | DONE | `CHANGELOG.md` task 6 entry, `design/loops/technical.md` section 8 updated |
## Quality Assessment
- **Test coverage:** 18 new tests, all passing. Full suite 393 passed (was 369). No regressions.
- **Backward compatibility:** `_invoke_harness` new params are optional. Existing tests updated to use non-existent prompt refs so `_resolve_prompt` fallback path is exercised. No breaking changes.
- **Code quality:** `_resolve_prompt` is clean, well-structured, handles all edge cases (missing file, missing artifact, empty prompt_ref, missing outputs dir). Follows existing code conventions.
- **Documentation:** README onboarding section is comprehensive. technical.md section 8 documents the prompt resolution flow. CHANGELOG is detailed.
- **Security:** Adversarial review found no exploitable vulnerabilities. Path traversal is mitigated by trusted input. Token injection is not possible (single-pass substitution). Large artifact DoS is mitigated by context floor gate.
## Pipeline Artifacts
- SPEC.md -- written and approved
- IMPLEMENTATION.md -- written
- CODE_REVIEW.md -- written, approved
- BUG_REPORT.md -- written (CLEAN, 2 LOW + 1 INFO)
- ADVERSARIAL_BUG_REPORT.md -- written (CLEAN, no exploitable vulnerabilities)
- DOC_REVIEW.md -- written (APPROVED)
## Verdict
**APPROVED -- ready for complete.**
All 9 requirements (R1-R9) are fully implemented, tested, and documented. The task delivers the critical missing piece of loop engineering v1: prompt-file token substitution that closes the feedback loop between ticks. The three loop role prompts provide the LLM instructions for the Implement/Verify/Orchestrate cycle. The self-improvement template enables the framework to improve itself via audit-driven loops.
@@ -1 +0,0 @@
complete
@@ -1,2 +0,0 @@
research:approved|2026-06-24T02:20:11.669836+00:00|user
code_review:approved|2026-06-24T02:22:31.598746+00:00|user
@@ -1,31 +0,0 @@
# Adversarial Bug Report: add-outputs-retention
Probed `_get_retention` and `_gc_outputs` with non-contract inputs.
## A1 — `retention` as float
`_get_retention({"outputs": {"retention": 3.14}})` → `int(3.14)` = 3. Not garbage but truncating. Acceptable (float is a numeric type; int() rounds toward zero). Not a regression.
## A2 — `retention` as bool
`_get_retention({"outputs": {"retention": True}})` → `int(True)` = 1. A user who sets `retention: true` intending "unlimited" gets 1 (wrong — they wanted 0). But `bool` is technically a subclass of `int` in Python; `int(True)` = 1 is documented behavior. Acceptable edge case — the user would need to write JSON `true`, which `json.loads` reads as `True`. Not blocking; `int(True)` = 1 is a narrow retention but valid.
## A3 — `retention` string "inf" falls back to 20
`_get_retention({"outputs": {"retention": "inf"}})` → `int("inf")` raises ValueError → caught → 20 with WARNING. Correct per SPEC D-O4.
## A4 — GC handles large gaps in tick indices
Files `tick1-*.json` and `tick100-*.json` with nothing in between: `max_seen=100`, `retention=20`, `cutoff=100-20+1=81`. Deletes tick1- but keeps tick100-. Correct — the gap is intentional (maybe intermittent ticks). Not a bug.
## A5 — Non-tick files `tick-nope.md` preserved
Hyphen-no-number prefix `tick-nope.md` doesn't match `^tick(\d+)-`. Preserved. Correct per D-O6.
## A6 — Empty outputs dir
`_gc_outputs` on dir with 0 files or missing dir returns cleanly. No crash. Confirmed.
## No BLOCKERS
All adversarial cases produce deterministic documented results. Proceed to doc_review.
@@ -1,17 +0,0 @@
# Bug Report: add-outputs-retention
## O1 — GC tick-count semantic: cutoff uses `max_seen` from filenames, not `state.iteration_count`
The formula `cutoff = max_seen - retention + 1` uses the max tick index found in filenames, NOT `state.iteration_count`. If the `.state.loop` advances to iteration_count=N but the output files for tick N haven't been written yet (crash after step 10 write but before GC), the next tick will see max_seen = N-1 and compute a cutoff that deletes one fewer group than expected. On the next tick, N is written and GC catches up.
**Not a bug** — SPEC D-O5 explicitly chose filename-based max_seen over iteration_count for robustness. Self-healing on the next tick.
## O2 — GC doesn't iterate recursively
If a future version nests files inside `outputs/` subdirectories (e.g., `outputs/tick5/`), `os.listdir` at the top level won't see them. The regex won't match, so they're preserved. Only top-level `tick{N}-*` files are affected.
**Not a bug** — SPEC D-O6: regex `^tick(\d+)-` matches only top-level files. Nested subdirs preserved. Not a current concern.
## Verdict
PASS — no blockers.
@@ -1,25 +0,0 @@
# Code Review: add-outputs-retention
## SPEC coverage
| Requirement | Status |
|-------------|--------|
| R1 — `_get_retention` helper from `outputs.retention` | ✓ |
| R2 — GC executes on every tick (post-write) | ✓ step 10.5 inside `_loop_lock` |
| R3 — Retention = 0 means no GC | ✓ `if retention <= 0: return` |
| R4 — GC failure doesn't crash tick | ✓ OSError caught → WARNING log + swallow |
| R5 — No new pip deps | ✓ stdlib only |
## Cross-script impact
- `scripts/loop-runner.py`: pure addition; no existing function changed.
- `templates/loops/self-improvement/loop.json`: new `outputs.retention: 20` field.
- `scripts/status.py`: no changes needed (create-loop template provides the default; runner reads, not status.py).
## Off-by-one fix
GC formula was `cutoff = max_seen - retention` (kept retention+1 groups). Found during test execution when `test_gc_keeps_recent_deletes_old` showed 21 remaining instead of 20. Fixed to `cutoff = max_seen - retention + 1`. Good test coverage.
## Verdict
PASS — proceed to bug_find.
@@ -1,17 +0,0 @@
# Doc Review: add-outputs-retention
## Docs touched
- `CHANGELOG.md` — new `[unreleased]` entry "Added — outputs retention GC" above the existing entries.
- `design/loops/technical.md` — new subsection "Outputs retention (v1.1 — `add-outputs-retention`)" after the lock serialization subsection in §7.
- `design/loops/functional.md` §9 — added `outputs: {retention: N}` row to the config-fields list.
## Docs NOT touched (intentional)
- `AGENTS.md`: outputs retention is runtime ergonomics, not an enforcement contract. No edit.
- `README.md`: user-facing README doesn't enumerate every `loop.json` field. No edit.
- `templates/loops/self-improvement/loop.json`: already updated (schema edit).
## Verdict
Docs in sync. Proceed to referee.
@@ -1,42 +0,0 @@
# Implementation: add-outputs-retention
## SCOPE
Add `loop.json` `outputs.retention` field (default 20) to bound growth of the `outputs/` directory. GC runs after step 10 inside `_loop_lock`, deleting tick groups older than the retention window. Source: `add-loop-runner/BUG_REPORT.md` O5.
## FILES TOUCHED
- `scripts/loop-runner.py`
- Added `_get_retention(cfg) -> int`: reads `cfg.get("outputs", {}).get("retention", 20)`. Non-int types fall back to 20 with WARNING. Negative values are coerced to 0 (unlimited) with WARNING.
- Added `_gc_outputs(loop_path, retention)`: lists `outputs/`, finds max tick index from filenames matching `^tick(\d+)-`, computes `cutoff = max_seen - retention + 1`, deletes files with tick index < cutoff. Non-tick files (`README.txt`, etc.) are preserved. Errors logged as WARNING via `_append_tick_log` and swallowed.
- Modified `cmd_tick`: calls `_get_retention(cfg)` + `_gc_outputs(loop_path, retention)` after step 10 (`_write_state_loop`) and before step 11 (tick log), inside the `_loop_lock` block.
- Updated docstring step list: added `10.5. GC outputs/...`.
- `templates/loops/self-improvement/loop.json`
- Added `"outputs": {"retention": 20}` block.
## BUG FOUND AND FIXED INLINE
**Off-by-one in GC formula**: the initial implementation used `cutoff = max_seen - retention`, which kept `retention + 1` tick groups (21 instead of 20 for retention=20). Fixed to `cutoff = max_seen - retention + 1`. Test `test_gc_keeps_recent_deletes_old` caught this (expected 20 kept, got 21 remaining → obvious failure when the remaining-count length check triggered).
## DECISIONS LOCKED
- **D-O1**: retention counts tick GROUPS (all `tick{N}-*` files), not individual files.
- **D-O2**: GC runs INSIDE `_loop_lock` critical section (after state write, before tick log).
- **D-O3**: Default 20.
- **D-O4**: 0 = unlimited (no GC). Negative coerces to 0 with WARNING.
- **D-O5**: GC based on `outputs/` filenames (`max_seen`), NOT `state.iteration_count`.
- **D-O6**: Regex `^tick(\d+)-`. Non-matching files preserved.
- **D-O7**: GC failure → WARNING log + swallow.
## TESTS
New file `tests/test_outputs_retention.py` — 13 tests across 2 classes:
- `TestGetRetention` (5): default `main`, explicit value, negative→0, non-int→20, None cfg→20.
- `TestGcOutputs` (8): deletes old keeps recent, retention=0 skip, retention>count, missing dir, non-tick files preserved, unrelated `tick-foo` prefix preserved, single tick, error path.
## TEST COUNT
- Baseline: 469 passed (post-`harden-parse-verdict`).
- New: +13 in `tests/test_outputs_retention.py`.
- Final: **482 passed**, 0 regressions.
@@ -1,4 +0,0 @@
# Review
- **Status**: approved
- **Timestamp**: 2026-06-23T22:19:32.288584
- **Comment**:
@@ -1,140 +0,0 @@
# SPEC: add-outputs-retention
## Problem
`tasks/add-loop-runner/BUG_REPORT.md` O5:
> Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has `max_iterations` to bound this; for daemon mode with `max_iterations=0`, the user is responsible.
Actual count is **6 files per tick** (each role: a `tick{N}-<role>-prompt.md` written by `_resolve_prompt`, plus a `tick{N}-<role>.json` written by `cmd_tick`). With `max_iterations=0` (daemon, unbounded), the `outputs/` directory grows without bound.
## Goal
Bound `outputs/` directory growth by retaining only the **last N tick groups**. A "tick group" = all files with the `tick{N}-` prefix for a single tick index N. Older tick groups are garbage-collected on every tick.
## Non-goals
- Per-role retention (e.g. keep verify-outputs longer than implement-outputs). Out of scope; would complicate the schema.
- Compression / archival of old tick dirs to a tarball. Out of scope.
- Cross-loop retention. Each loop's `outputs/` is independent.
- Retention of `.state.log` (tick log). That file is append-only and grows linearly; separate concern.
## Schema addition (`loop.json`)
Add an optional `outputs` object:
```json
"outputs": {
"retention": 20
}
```
- **`outputs.retention`** (int, optional, default **20**): keep the last N tick groups. Older tick groups are deleted on every tick. `0` = unlimited (no GC; v1 behavior). Negative values are rejected at `--create-loop`.
## Requirements
### R1 — retention config plumbing
- `status.py --create-loop` accepts `outputs.retention` in the `loop.json` template.
- The runner reads `cfg.get("outputs", {}).get("retention", 20)`.
- Validation on read: if `retention` is < 0, log WARNING and treat as `0` (unlimited). Non-int types coerce via `int(...)`; on `TypeError`/`ValueError` fall back to default `20`.
### R2 — GC executes on every tick (post-write)
- After step 10 (`_write_state_loop`) and before step 11 (tick log), the runner invokes `_gc_outputs(loop_path, state, retention)`.
- GC iterates `outputs/` directory, parses `tickNN-` prefixes, computes the cutoff = `iteration_count - retention + 1` (kept range: `[cutoff, iteration_count]` inclusive).
- Any file whose tick-index prefix is `< cutoff` is deleted. Files without a `tickN-` prefix are left alone (forward-compat; user may place other files in `outputs/`).
- GC errors (file in use, permission) are logged via `_append_tick_log` WARNING and swallowed — GC failure must not crash the tick.
### R3 — Retention = 0 means no GC
- `0` skips the GC step entirely (cheapest path for `max_iterations` users who prefer manual cleanup).
### R4 — Atomicity / failure isolation
- GC failures (permission, file not found mid-iteration) don't roll back the tick. State has already advanced; losing a GC pass is benign (next tick re-attempts).
- Missing `outputs/` (loop never ticked) — GC no-ops, no error.
### R5 — No new pip deps
- Pure stdlib: `os.listdir`, `os.remove`, `re.match`. No `shutil.rmtree` (we delete individual files; a tick group is not a directory).
## Detailed semantics
### Tick-index extraction
Filenames follow the pattern `tick<int>-<remainder>` where `<int>` is the 1-based tick index. Examples:
- `tick1-implement.json`, `tick1-verify.json`, `tick1-orchestrate.json`, `tick1-implement-prompt.md`, `tick1-verify-prompt.md`, `tick1-orchestrate-prompt.md`
Regex: `^tick(\d+)-`. Tick indices are extracted into a set, the maximum tick index (`max_seen`) is computed, and the cutoff floor is `max_seen - retention + 1`. Files with tick index `< floor` get deleted.
**Why `max_seen - retention + 1` instead of `state.iteration_count`?**
State could lag (e.g. concurrent ticks), but the on-disk filenames ARE ground truth. Using max filename keeps GC self-contained.
### Default retention choice
Default = **20**. Rationale:
- Score-plateau window default is often 5-10; keeping 2x that covers debugging.
- 20 ticks × 6 files = 120 files max — comfortably under any filesystem degradation threshold.
- Operators who need longer history (`audit` use cases) override upward in `loop.json`.
### Where GC runs in the tick flow
```
... step 10: _write_state_loop(state)
# NEW: step 10.5
_gc_outputs(loop_path, state, retention)
# step 11
_append_tick_log(...)
```
GC runs INSIDE the `_loop_lock` critical section, so a concurrent `--pause-loop` / `--approve --loop` can't be mid-write and observe a missing tick dir. GC's filesystem delete ops are independent of `.state.loop`.
## Test plan
Pure-function + filesystem tests (no subprocess, no live LLM):
1. **GC deletes old tick groups, keeps recent N**: write 30 tick groups (6 files each), retention=20, expect last 20 kept, oldest 10 deleted, all 6 files per kept tick are present.
2. **Retention = 0 skips GC entirely**: 30 tick groups, retention=0, expect no deletion, all files present.
3. **Retention > file count** (no-op): 5 tick groups, retention=20, expect no deletion.
4. **Missing `outputs/` dir** (no-op, no error): fresh loop, no `outputs/`, GC returns cleanly.
5. **Non-tick files in `outputs/` are preserved**: write 30 tick groups + a `README.txt` and `loop-info.md`, retention=20, expect tick groups deleted but `README.txt` and `loop-info.md` intact.
6. **Negative retention coerces to 0 (no GC)**: retention=-5 in `loop.json`, expect WARNING + no deletion.
7. **Non-int retention coerces to default 20**: retention="twenty", expect WARNING + default 20 used (deletes oldest 10 of 30).
8. **Tick-index regex preserves unrelated `tick-foo` files** (defensive): `tick-foo.md` (no number) does NOT match `^tick(\d+)-`; expect preserved.
9. **GC error swallowed (permission-denied file)**: chmod 000 a stale tick file (or use a non-existent mock that raises `PermissionError`); expect GC logs WARNING and continues; tick proceeds.
10. **Concurrent with state write** (lock interaction): GC runs inside the lock; no separate test needed (the `test_state_loop_lock.py` suite already covers lock integrity).
11. **Config plumbing**: `--create-loop` writes `outputs.retention: 20` into generated `loop.json` (if `--outputs-retention` not provided; or honors override).
12. **Default getter**: `_get_retention(cfg)` returns 20 for missing `outputs`, 0 when `{"outputs": {"retention": 0}}`, 20 for `{"outputs": {"retention": "garbage"}}` (post-WARNING).
## Decisions (locked)
- **D-O1**: retention counts tick GROUPS not individual files. A tick group = all `tick{N}-*` files. Keeps the mental model aligned with "ticks as the atomic unit".
- **D-O2**: GC runs INSIDE `_loop_lock` critical section (after state write, before tick log). Cheapest correct placement — no separate lock, no concurrent `--pause-loop` / `--approve --loop` mid-GC race. Filesystem delete ops are independent of `.state.loop` but the lock keeps the loop's externally-observable state consistent.
- **D-O3**: Default 20 (covers debugging; 120 files max comfortably under fs degradation).
- **D-O4**: `0` = unlimited (no GC). Negative coerces to 0 with WARNING.
- **D-O5**: GC based on `outputs/` filenames (`max_seen`), NOT `state.iteration_count`. Self-contained; robust to state lag.
- **D-O6**: Regex `^tick(\d+)-`. Files not matching are preserved (forward-compat for helper docs, scratch notes, etc.).
- **D-O7**: GC failure (PermissionError, FileNotFoundError mid-iteration) → WARNING log + swallow. Tick not affected.
## Out of scope (filed BACKLOG.md)
- `outputs.retention_bytes` (磁盘 budget cap). Future.
- Tarball archival of GC'd tick groups. Future.
- Cross-loop retention aggregation. Future.
- GC `.state.log` rotation. Separate task (`add-state-log-rotation`).
## Files touched
- `scripts/loop-runner.py` — add `_get_retention(cfg)` + `_gc_outputs(loop_path, state, retention)`; call after step 10 inside `_loop_lock`.
- `scripts/status.py` — `--create-loop` writes `outputs.retention` default 20 into generated `loop.json` template; validates non-negative.
- `templates/loops/self-improvement/loop.json` — add `"outputs": {"retention": 20}` to template.
- `design/loops/technical.md` — new subsection §7b "Outputs retention (v1.1 — `add-outputs-retention`)".
- `design/loops/functional.md` — note `outputs.retention` field in the schema enum.
- `CHANGELOG.md` — new entry under `[unreleased]`.
- `tests/test_outputs_retention.py` (NEW) — 12 tests per plan above.
## Pipeline plan
research → research:awaiting_approval → research:approved → implement → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete.
@@ -1,22 +0,0 @@
# Referee Verdict: add-outputs-retention
## Status: PASS
## Artifacts reviewed
- SPEC.md, IMPLEMENTATION.md, CODE_REVIEW.md, BUG_REPORT.md, ADVERSARIAL_BUG_REPORT.md, DOC_REVIEW.md
## Phase gates satisfied
All 8 required artifacts present. Pipeline driven: research → implement → code_review → bug_find → adversarial_bug_find → doc_review → referee.
## Acceptance
- R1-R5 all satisfied. GC runs inside `_loop_lock` after step 10. 0 = unlimited. Non-int/negative handled gracefully. No new deps.
- Off-by-one bug (`cutoff = max_seen - retention` → `cutoff = max_seen - retention + 1`) caught inline by test. Fixed before full suite.
- 482 passed (469 + 13 new, 0 regressions). Docs in sync (CHANGELOG, technical.md, functional.md).
- Adversarial probes: float truncation (3.14→3), bool True→1, string "inf"→20 (WARNING), non-tick files preserved, missing dir safe. All deterministic documented behavior.
## Verdict
PASS — task complete. Approve transition to complete.
@@ -1 +0,0 @@
complete
@@ -1 +0,0 @@
code_review:approved|2026-06-16T16:44:02.515729+00:00|user
@@ -1 +0,0 @@
# No adversarial bugs
@@ -1 +0,0 @@
# No bugs
@@ -1 +0,0 @@
# Code review: PASS
@@ -1 +0,0 @@
# Doc review: PASS
@@ -1,11 +0,0 @@
# Implementation: Post-Commit Autopilot Reminder
## Changes
- Created `scripts/git-hooks/post-commit`
- Reads `.agent.md` for autopilot mode
- Scans non-terminal tasks via `status.py --list`
- Prints summary with count and command to run orchestrator
- Always exits 0 (informational only)
## Files Created
- `scripts/git-hooks/post-commit`: ~45 lines
@@ -1,20 +0,0 @@
# SPEC: Add Post-Commit Autopilot Driver Reminder
## Problem
When autopilot is enabled, tasks can be left hanging after commits. There's no
automated reminder to drive them to completion.
## Requirements
1. Create a post-commit hook at `scripts/git-hooks/post-commit`
2. After a commit succeeds, check if autopilot is enabled (.agent.md)
3. If autopilot is enabled and non-terminal tasks exist, output a summary:
- Number of non-terminal tasks
- Their current phases
- Clear instruction: "Run the orchestrator to drive them to completion"
4. Exit 0 always (informational only, never blocks)
## Acceptance Criteria
- Post-commit hook exists and is executable
- Hook output is silent when no work remains or autopilot is off
- Hook prints actionable summary when work exists and autopilot is on
- bash -n passes syntax check
@@ -1 +0,0 @@
## Status: PASS
@@ -1 +0,0 @@
complete
@@ -1,28 +0,0 @@
# Implementation: Add pytest Test Suite
## Summary
Added a comprehensive pytest suite covering the dashboard core modules and the new VRAM detection script.
## Files Changed
- `tests/test_scope.py` (new)
- `tests/test_task.py` (new)
- `tests/test_board.py` (new)
- `tests/test_stats.py` (new)
- `tests/test_config.py` (new)
- `tests/test_app.py` (new)
- `tests/test_vram_detect.py` (new)
- `tests/test_prompt_paths.py` (created earlier in Task 5)
- `pyproject.toml` (new root config with optional dependencies)
- `automaton/dashboard/pyproject.toml` (deleted to avoid conflict)
- `.gitignore` (updated for pytest cache, egg-info, venvs)
## Bug Fixes Found During Testing
- `DashboardHandler._validate_task_name` was an instance method; converted to `@staticmethod`.
- `scripts/vram_detect.py` regex for override context window did not match `**Override context window**`.
- `scripts/vram_detect.py` `_parse_token_value` did not handle decimal values like `5.6k`.
## Verification
- `python -m pytest tests/` passes: **70 tests passed**.
## Notes
- The root `pyproject.toml` now defines `automaton` package discovery and optional dependency groups.
@@ -1,4 +0,0 @@
# Review
- **Status**: approved
- **Timestamp**: 2026-06-14T09:59:32.664776
- **Comment**:
@@ -1,29 +0,0 @@
# SPEC: Add pytest Test Suite
## Goal
Add automated tests for the dashboard and the new VRAM detection script.
## Requirements
1. Create `tests/test_scope.py` for scope detection.
2. Create `tests/test_task.py` for task state determination and sub-task parsing.
3. Create `tests/test_board.py` for Kanban board grouping and filtering.
4. Create `tests/test_stats.py` for statistics calculations.
5. Create `tests/test_config.py` for config validation and defaults.
6. Create `tests/test_app.py` for dashboard HTTP API endpoints and path-traversal guard.
7. Create `tests/test_vram_detect.py` for the VRAM detector using mocked system data.
8. Update `pyproject.toml` with optional dependencies:
- `test` extra: `pytest`
- `dashboard` extra: `inotify` (optional)
9. Update `.gitignore` for `.pytest_cache/`.
## Acceptance Criteria
- [ ] `python -m pytest` discovers and passes all tests.
- [ ] Tests exercise state determination, filtering, API responses, config validation, and VRAM detection.
- [ ] `pyproject.toml` includes the optional dependency groups.
## Non-Goals
- Achieving 100% coverage.
- Testing shell scripts (handled separately).
## Stop Condition
When all acceptance criteria are met, output "CONTRACT_MET".
@@ -1,18 +0,0 @@
# Verdict: Add pytest Test Suite
## Status: PASS
**Completion Date**: 2026-06-14
## Summary
A pytest suite has been added covering scope detection, task state determination, board logic, statistics, configuration, dashboard app validation, and VRAM detection. All tests pass.
## Findings
- 70 tests pass.
- Root packaging configured.
- Minor bugs in `_validate_task_name` and VRAM token parsing were discovered and fixed during test development.
## Remaining Issues
None.
## Score
+10 PASS
@@ -1 +0,0 @@
complete
@@ -1,2 +0,0 @@
research:approved|2026-06-23T13:04:47.561659+00:00|user
code_review:approved|2026-06-23T13:06:37.501736+00:00|user
@@ -1,48 +0,0 @@
# ADVERSARIAL_BUG_REPORT: add-self-improvement-loop
## Methodology
Targeted attack on:
1. Shell injection via `$FRAMEWORK_DIR`
2. Race condition between install.sh and update.sh
3. Loop creation failure cascading to install failure
4. Schedule installation on unsupported platforms
5. Template path traversal
## Findings
### Attack 1: Shell injection via `$FRAMEWORK_DIR` -- NOT VULNERABLE
`$FRAMEWORK_DIR` is set to `$HOME/.automaton` at the top of both scripts. It is not derived from user input. The `--project "$FRAMEWORK_DIR"` argument is passed as a single quoted argument to `python3`, so no shell expansion occurs inside the Python process. No injection vector.
**Verdict:** NOT VULNERABLE
### Attack 2: Race condition between install.sh and update.sh -- NOT EXPLOITABLE
If a user runs `install.sh` and `update.sh` concurrently (which would be unusual), both might try to create the loop simultaneously. `--create-loop` checks `if loop_path.exists()` and returns rc=2 if it exists. The `mkdir(parents=True)` in `cmd_create_loop` is not atomic, but the `.state.loop` write is atomic (tmp+rename). Worst case: one script gets rc=2 and `|| true` swallows it. No data corruption.
**Verdict:** NOT EXPLOITABLE
### Attack 3: Loop creation failure cascading -- NOT VULNERABLE
Both `--create-loop` and `--install-schedule` are followed by `|| true`. If either fails, the script continues. The `.venv` setup and pip install at the end of `install.sh` are outside the `else` block and run regardless. The framework works without the loop.
**Verdict:** NOT VULNERABLE
### Attack 4: Schedule installation on unsupported platforms -- HANDLED
`--install-schedule` handles platform dispatch internally (Darwin -> launchd, Linux -> cron, Windows -> schtasks). On an unknown platform, it prints an error and returns non-zero, which `|| true` swallows. The loop is created but not scheduled; the user can manually run `--mode tick` or `--mode daemon`.
**Verdict:** HANDLED
### Attack 5: Template path traversal -- NOT VULNERABLE
`--from-template self-improvement` is a fixed string in both scripts. `cmd_create_loop` constructs the template path as `AUTOMATON_DIR / "templates" / "loops" / template`. The template name is not user-supplied in this context.
**Verdict:** NOT VULNERABLE
## Summary
No exploitable vulnerabilities found. All attack surfaces are mitigated by trusted input, `|| true` non-fatal behavior, and atomic state writes.
**Verdict: CLEAN**
@@ -1,23 +0,0 @@
# BUG_REPORT: add-self-improvement-loop
## Findings
### Bug 1 (LOW): install.sh loop bootstrap is inside the `else` block
The loop creation commands are inside the `else` block of `if [ -d "$FRAMEWORK_DIR" ]`, which means they only run on fresh installs. If a user previously installed the framework before this change and runs `install.sh` again, they get "already installed" and the loop is NOT created. This is correct behavior -- `update.sh` handles the existing-user case.
**Severity:** LOW (by design)
**Fix:** None needed.
### Bug 2 (INFO): No `--project` flag consistency check
`install.sh` uses `--project "$FRAMEWORK_DIR"` while `update.sh` also uses `--project "$FRAMEWORK_DIR"`. Both are consistent. The `work_source.project` in the template is `"~/.automaton/"` (a string), but `--create-loop` doesn't use `work_source.project` -- it uses the `--project` flag. The runner reads `work_source.project` at tick time. No mismatch because `--project "$FRAMEWORK_DIR"` (which is `$HOME/.automaton`) and `work_source.project: "~/.automaton/"` resolve to the same path.
**Severity:** INFO (no bug)
**Fix:** None needed.
## Summary
No correctness bugs found. One LOW (by design) and one INFO.
**Verdict: CLEAN**
@@ -1,56 +0,0 @@
# CODE_REVIEW: add-self-improvement-loop
## Reviewed Files
1. `scripts/install.sh` -- self-improvement loop bootstrap (lines ~70-82)
2. `scripts/update.sh` -- idempotent loop bootstrap (lines ~61-70)
3. `tests/test_self_improvement_loop.py` -- 16 tests
4. `CHANGELOG.md` -- task 7 entry
5. `design/loops/technical.md` section 9 -- updated install note
6. `README.md` -- self-improvement loop default-on section
## Findings
### 1. install.sh -- loop bootstrap placement
The loop bootstrap is placed inside the `else` block (after the git clone), after guard registration and before `fi`. This is correct -- the loop should only be created on fresh installs, not when the framework is already installed (the `if [ -d "$FRAMEWORK_DIR" ]` branch prints "already installed" and exits).
The `|| true` ensures install continues even if `status.py` fails (e.g. Python not in PATH yet, or schedule installation fails on an unusual platform). The framework works without the loop.
**Verdict:** PASS
### 2. update.sh -- idempotent bootstrap
The `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]` check correctly prevents duplicate creation. `--create-loop` itself also refuses duplicates (returns rc=2), but the directory check avoids the error output entirely. The `|| true` on both commands ensures update continues on failure.
**Verdict:** PASS
### 3. Test coverage
- `TestInstallShWiring` (5 tests): covers create-loop, install-schedule, opt-out message, framework project, and non-fatal behavior. All assertions check the script content.
- `TestUpdateShWiring` (4 tests): covers create-loop, idempotent check, install-schedule, and non-fatal behavior.
- `TestSelfImprovementTemplate` (5 tests): regression guard for template fields.
- `TestCreateLoopFromTemplate` (2 tests): integration test for `cmd_create_loop` with the self-improvement template.
**Verdict:** PASS
### 4. Shell syntax
`bash -n scripts/install.sh scripts/update.sh` passes. No syntax errors.
**Verdict:** PASS
### 5. Edge cases
- **Python not in PATH**: `|| true` handles this. Install continues.
- **Loop already exists (update.sh)**: directory check prevents creation; `--create-loop` also refuses.
- **Schedule installation fails**: `|| true` handles this. Loop is created but not scheduled; user can manually `--install-schedule` later.
- **Framework not in ~/.automaton**: the `$FRAMEWORK_DIR` variable is set at the top of each script and used consistently.
**Verdict:** PASS
## Summary
All 5 review areas pass. The implementation is clean, idempotent, and well-tested. 16 new tests cover script wiring, template validation, and loop creation. Full suite: 409 passed.
**Overall verdict: APPROVED**
@@ -1,38 +0,0 @@
# DOC_REVIEW: add-self-improvement-loop
## Reviewed Documentation
1. `README.md` -- new "Self-Improvement Loop (Default-On)" section
2. `CHANGELOG.md` -- task 7 entry
3. `design/loops/technical.md` section 9 -- updated install note
## Findings
### 1. README.md
New section "Self-Improvement Loop (Default-On)" accurately documents:
- What the loop does (ticks against `status.py --audit`)
- Schedule (3600s / 1 hour)
- Brakes (max_iterations: 10, score_plateau_window: 3)
- How to disable/re-enable (`--pause-loop` / `--resume-loop`)
- Worktree and file scope
**Verdict:** PASS
### 2. CHANGELOG.md
Entry accurately describes install.sh and update.sh changes, new tests (16), and doc updates.
**Verdict:** PASS
### 3. technical.md section 9
Updated the install note to include `--project "$FRAMEWORK_DIR"`, `|| true`, and the `update.sh` idempotent bootstrap. Matches the implementation.
**Verdict:** PASS
## Summary
All documentation is accurate and consistent with the implementation.
**Verdict: APPROVED**
@@ -1,41 +0,0 @@
# IMPLEMENTATION: add-self-improvement-loop
## Summary
Wired the self-improvement loop into `install.sh` (default-on for fresh installs) and `update.sh` (idempotent bootstrap for existing users). Both use `status.py --create-loop self-improvement --from-template self-improvement` and `--install-schedule self-improvement --interval 3600` with `|| true` to ensure the framework continues to work even if loop creation fails.
## Changes
### R1 -- `scripts/install.sh`
Added after guard registration (inside the `else` block, before `fi`):
- `--create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` with `|| true`
- `--install-schedule self-improvement --interval 3600 --project "$FRAMEWORK_DIR"` with `|| true`
- User-facing message about the self-improvement loop and how to disable it with `--pause-loop`
### R2 -- `scripts/update.sh`
Added after guard registration:
- Idempotent check: `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]`
- Same `--create-loop` and `--install-schedule` commands with `|| true`
- Info message when the loop is created
### R3 -- `tests/test_self_improvement_loop.py`
16 tests across 4 classes:
- `TestInstallShWiring` (5 tests): verify install.sh contains create-loop, install-schedule, opt-out message, framework project, and `|| true`
- `TestUpdateShWiring` (4 tests): verify update.sh contains create-loop, idempotent check, install-schedule, and `|| true`
- `TestSelfImprovementTemplate` (5 tests): verify template fields (audit work source, brakes, worktree, file scope, role prompts)
- `TestCreateLoopFromTemplate` (2 tests): simulate `--create-loop self-improvement --from-template self-improvement` and verify directory structure; verify duplicate creation is refused
### R4 -- Documentation
- `CHANGELOG.md`: task 7 entry under `[unreleased]`
- `README.md`: note that self-improvement loop is default-on at install
- `design/loops/technical.md` section 9: note that install.sh creates it default-on
## Verification
- `bash -n scripts/install.sh scripts/update.sh` -- OK
- `python3 -m pytest tests/test_self_improvement_loop.py -v` -- 16 passed
- `python3 -m pytest tests/ -q` -- 409 passed (393 + 16 new)
@@ -1,81 +0,0 @@
# RESEARCH: add-self-improvement-loop
## Objective
Make the self-improvement loop default-on at install time (D21). The template `templates/loops/self-improvement/loop.json` was already created in task 6. This task wires it into `install.sh` and `update.sh` so that:
- Fresh installs get the loop created and scheduled automatically
- Existing users who run `update.sh` get the loop bootstrapped (idempotent -- skip if already exists)
## Current State
### `install.sh` (lines 1-79)
- Clones repo to `~/.automaton`
- Runs VRAM detection
- Registers pre-edit guards via `register-guards.sh`
- Sets up `.venv` and pip deps
- Does NOT create any loops
### `update.sh` (lines 1-74)
- Pulls latest from git
- Checks for deprecated file locations
- Registers guards
- Installs git hooks in current project
- Does NOT create any loops
### `--create-loop` (status.py:1773)
- Takes `--create-loop <name>`, `--from-template <name>`, `--project <path>`
- Creates `~/.automaton/loops/<name>/` (when project is `~/.automaton/`)
- Copies `loop.json` from template, patches `name` field
- Creates `.state.loop` with initial state (`running`)
- Creates empty `.state.log`
- Returns error if loop already exists
### `--install-schedule` (status.py:1807)
- Takes `--install-schedule <name>`, `--interval <seconds>`, `--project <path>`
- Generates OS-specific tick stub (`automaton-loop-tick.sh` or `.bat`)
- Installs OS schedule unit (launchd plist on macOS, cron on Linux, schtasks on Windows)
- Interval defaults to `loop.json schedule.interval_seconds` or 3600
### `_loops_dir` (status.py:1609)
- When project is `~/.automaton/`, loops dir is `~/.automaton/loops/`
- When project is other, loops dir is `<project>/.automaton/loops/`
## Design Decisions
### D1: Where to add the install hook
In `install.sh`, after the clone and guard registration, add:
```bash
# Bootstrap self-improvement loop (default-on, D21)
python3 "$FRAMEWORK_DIR/scripts/status.py" --create-loop self-improvement \
--from-template self-improvement --project "$FRAMEWORK_DIR"
python3 "$FRAMEWORK_DIR/scripts/status.py" --install-schedule self-improvement \
--interval 3600 --project "$FRAMEWORK_DIR"
```
### D2: Where to add the update hook
In `update.sh`, after the git pull and guard registration, add an idempotent bootstrap:
```bash
# Bootstrap self-improvement loop if not present (default-on, D21)
if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]; then
python3 "$FRAMEWORK_DIR/scripts/status.py" --create-loop self-improvement \
--from-template self-improvement --project "$FRAMEWORK_DIR"
python3 "$FRAMEWORK_DIR/scripts/status.py" --install-schedule self-improvement \
--interval 3600 --project "$FRAMEWORK_DIR"
fi
```
### D3: Test approach
The test `test_self_improvement_installs_default_on` should verify that `install.sh` contains the create-loop and install-schedule commands for the self-improvement loop. A full integration test (actually running install.sh) would require a mock git clone target and is fragile. Instead, test the script content for the required commands, and test that `--create-loop self-improvement --from-template self-improvement --project <framework>` produces the expected directory structure (this is already tested in the status.py tests but we add a specific test for the self-improvement template).
### D4: User opt-out
Users can disable the self-improvement loop with:
```bash
python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/
```
This should be documented in the install output and README.
## Risks
- **install.sh failure**: if `--create-loop` fails (e.g. Python not in PATH yet), install.sh should continue (the loop is optional, not critical for framework operation). Use `|| true` to non-fatal the loop bootstrap.
- **update.sh idempotency**: the `if [ ! -d ... ]` check ensures existing users don't get errors on repeated updates.
- **Platform differences**: `--install-schedule` handles platform dispatch internally. No shell-level platform checks needed.
@@ -1,72 +0,0 @@
# SPEC: add-self-improvement-loop
## Context
Task 6 created the self-improvement loop template at `templates/loops/self-improvement/loop.json`. This task wires it into `install.sh` and `update.sh` so the loop is default-on at install time (D21). Existing users who run `update.sh` get the loop bootstrapped idempotently.
## Non-Goals (deferred)
- Loop dashboard panel -> v1.1
- Auto-approve for self-improvement loop -> never (D4)
- Tier 2 context-sizing work -> picked up by the loop itself after first tick
- `design/context-sizing/` skeleton -> v1.1 (the loop will create it when it picks up Tier 2 work)
## Requirements
### R1 -- `install.sh` creates and schedules the self-improvement loop
After the clone and guard registration, add:
```bash
# Bootstrap self-improvement loop (default-on, D21)
python3 "$FRAMEWORK_DIR/scripts/status.py" --create-loop self-improvement \
--from-template self-improvement --project "$FRAMEWORK_DIR" || true
python3 "$FRAMEWORK_DIR/scripts/status.py" --install-schedule self-improvement \
--interval 3600 --project "$FRAMEWORK_DIR" || true
```
The `|| true` ensures install continues even if loop creation fails (e.g. Python not yet in PATH, or schedule installation fails on an unusual platform). The loop is optional; the framework works without it.
Print a message telling the user the loop is running and how to disable it:
```bash
echo ""
echo "=== Self-Improvement Loop ==="
echo "A self-improvement loop has been created and scheduled (runs every 3600s)."
echo "It will tick against status.py --audit on this framework's own repo."
echo "To disable: python3 ~/.automaton/scripts/status.py --pause-loop self-improvement --project ~/.automaton/"
```
### R2 -- `update.sh` bootstraps the self-improvement loop idempotently
After the git pull and guard registration, add:
```bash
# Bootstrap self-improvement loop if not present (default-on, D21)
if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]; then
python3 "$FRAMEWORK_DIR/scripts/status.py" --create-loop self-improvement \
--from-template self-improvement --project "$FRAMEWORK_DIR" || true
python3 "$FRAMEWORK_DIR/scripts/status.py" --install-schedule self-improvement \
--interval 3600 --project "$FRAMEWORK_DIR" || true
echo "Created self-improvement loop (default-on). --pause-loop self-improvement to disable."
fi
```
### R3 -- Tests
Write `tests/test_self_improvement_loop.py` with:
1. `test_install_sh_creates_self_improvement_loop` -- verify `install.sh` contains `--create-loop self-improvement` and `--install-schedule self-improvement`
2. `test_update_sh_bootstraps_self_improvement_loop` -- verify `update.sh` contains the idempotent bootstrap check
3. `test_install_sh_has_opt_out_message` -- verify `install.sh` contains `--pause-loop self-improvement`
4. `test_self_improvement_template_has_correct_fields` -- verify the template has `work_source.kind: audit`, `brakes.max_iterations: 10`, `blast_radius.use_worktree: true` (this may overlap with task 6 tests; if so, keep it as a regression guard)
5. `test_create_loop_self_improvement_from_template` -- simulate `--create-loop self-improvement --from-template self-improvement --project <tmp>` and verify the loop dir, `loop.json`, `.state.loop`, and `.state.log` are created correctly
### R4 -- Documentation updates
- `CHANGELOG.md` under `[unreleased]`
- `README.md` -- add a note in the Loop Engineering section that the self-improvement loop is default-on at install
- `design/loops/technical.md` -- section 9 already documents the self-improvement template; add a note that install.sh creates it default-on
## Verification
- `bash -n scripts/install.sh scripts/update.sh` -- syntax check
- `python3 -m pytest tests/test_self_improvement_loop.py -v`
- `python3 -m pytest tests/ -q` -- full suite must remain green
@@ -1,29 +0,0 @@
# VERDICT: add-self-improvement-loop
## Task
Wire the self-improvement loop into `install.sh` (default-on for fresh installs) and `update.sh` (idempotent bootstrap for existing users), per D21.
## Deliverables Review
| Requirement | Status | Evidence |
|---|---|---|
| R1: install.sh creates and schedules loop | DONE | `scripts/install.sh` lines ~70-82, 5 tests in `TestInstallShWiring` |
| R2: update.sh idempotent bootstrap | DONE | `scripts/update.sh` lines ~61-70, 4 tests in `TestUpdateShWiring` |
| R3: Tests | DONE | 16 tests in `tests/test_self_improvement_loop.py`, all passing |
| R4: Documentation | DONE | CHANGELOG, README, technical.md section 9 updated |
## Quality Assessment
- **Test coverage:** 16 new tests, all passing. Full suite 409 passed (was 393). No regressions.
- **Shell syntax:** `bash -n` passes for both scripts.
- **Idempotency:** `update.sh` checks for existing loop dir before creating. `--create-loop` also refuses duplicates.
- **Non-fatal behavior:** `|| true` on both commands ensures framework works even if loop creation fails.
- **Security:** Adversarial review found no exploitable vulnerabilities.
- **Documentation:** All docs accurate and consistent.
## Verdict
**APPROVED -- ready for complete.**
All 4 requirements fully implemented, tested, and documented. The self-improvement loop is now default-on at install time (D21), with idempotent bootstrap for existing users.
@@ -1 +0,0 @@
complete

Some files were not shown because too many files have changed in this diff Show More