Complete tasks 3-7: harden verdict parsing, outputs retention, base branch, linux schedule parity, claim loop task
CI / build (push) Has been cancelled
CI / build (push) Has been cancelled
This commit is contained in:
+168
@@ -2,6 +2,174 @@
|
||||
|
||||
## [unreleased]
|
||||
|
||||
### Added — cross-loop task claim (task `add-claim-loop-task`)
|
||||
|
||||
- **`scripts/status.py`**: New `--claim-loop-task <name> --task <taskname> [--project P]` command. Exit 0 = claimed (or already self-claimed, idempotent). Exit 2 = already claimed by another running/paused loop (`task_already_claimed:{other}` on stderr) or untracked loop (`loop_untracked`). Uses `_loop_lock` to serialize writes; cross-loop scan is advisory (self-healing on next tick).
|
||||
- **`scripts/loop-runner.py`**: `cmd_tick` step 3.5: after `_find_work` returns a candidate different from `current_task`, spawn `status.py --claim-loop-task` as subprocess with `$AUTOMATON_NO_LOOP_LOCK=1` (same bypass pattern as `_gate`). Step 9.5: after orchestrator, re-read `.state`; if phase is `complete` or `human_intervention`, clear `current_task` to None. Closes `add-status-brakes/ADVERSARIAL_BUG_REPORT.md` A2 (cross-loop task race).
|
||||
- **New tests**: `tests/test_claim_loop_task.py` — 10 tests covering claim success, refusal, idempotency, untracked loop, missing task, paused-loop ownership, self-healing race, and release on terminal phases.
|
||||
- **Full suite**: **518 passed** (was 508; +10 new; 0 regressions).
|
||||
|
||||
### Added — Linux schedule parity (task `linux-schedule-parity`)
|
||||
|
||||
- **`scripts/status.py`**: New `_install_cron_block(name, project, interval_seconds)` for Linux cron support. Writes `# automaton-loop:<name>` / `# end automaton-loop:<name>` blocks into crontab via `crontab -`. Interval rounded to full minutes, minimum 1. Strips prior block before insert (idempotent).
|
||||
- New `_enable_schedule(name, project)`: platform dispatch — Linux (cron insert via `_install_cron_block`), Darwin (rename `.plist.disabled` back), Windows (no-op). Reads `loop.json > schedule.interval_seconds`; non-int falls back to 3600.
|
||||
- `_disable_schedule` already existed; verified correct.
|
||||
- **New tests**: `tests/test_linux_schedule_parity.py` — 13 tests covering cron block writes, error handling, platform dispatch, and interval edge cases.
|
||||
- **Full suite**: **508 passed** (was 495; +13 new; 0 regressions).
|
||||
|
||||
### Fixed — parametrize-base-branch (task `parametrize-base-branch`)
|
||||
|
||||
- **`scripts/status.py`**: Replaced hardcoded `"main"` in `_gate_worktree_drift`'s `git diff` call with `_base_branch(cfg) -> str`. New helper returns `blast_radius.base_branch` if configured (default `"main"`). Empty string and non-string types produce WARNING and fall back to `"main"`. Closes `add-status-brakes/BUG_REPORT.md` O3: projects on `master`/`trunk`/`develop` no longer have a silently-disabled drift gate — operator sets `"base_branch": "master"` in `loop.json`.
|
||||
- **`templates/loops/self-improvement/loop.json`**: Added `"base_branch": "main"` to `blast_radius`.
|
||||
- **`design/loops/functional.md` §9**: Updated `blast_radius` field list to include `base_branch`.
|
||||
- **New tests**: `tests/test_base_branch.py` — 13 tests covering `_base_branch` helper (5) and drift-gate branch usage (8, with mocked `subprocess.run`).
|
||||
- **Full suite**: **495 passed** (was 482; +13 new; 0 regressions).
|
||||
|
||||
### Added — outputs retention GC (task `add-outputs-retention`)
|
||||
|
||||
- **`scripts/loop-runner.py`**: Added `_get_retention(cfg)` and `_gc_outputs(loop_path, retention)` to bound `outputs/` directory growth. Every tick (`cmd_tick` step 10.5, inside `_loop_lock`), older tick groups are deleted, keeping only the last N (default 20). Closes `add-loop-runner/BUG_REPORT.md` O5 (tick dirs accumulate without bound).
|
||||
- `_get_retention` reads `cfg.get("outputs", {}).get("retention", 20)`. Non-int types fall back to 20 with WARNING. Negative values coerce to 0 (unlimited) with WARNING. 0 = no GC (v1 behavior).
|
||||
- `_gc_outputs` lists `outputs/`, parses `tickNN` indices via `^tick(\d+)-` regex, computes `cutoff = max_seen - retention + 1`, deletes files with tick index < cutoff. Non-tick files (e.g. `README.txt`, `tick-foo.md`) are preserved. Errors (permission, missing file) are logged as WARNING and swallowed — GC failure never crashes the tick.
|
||||
- Off-by-one bug found and fixed inline: initial formula `cutoff = max_seen - retention` kept `retention+1` groups; caught by `test_gc_keeps_recent_deletes_old` length assertion.
|
||||
- **`templates/loops/self-improvement/loop.json`**: Added `"outputs": {"retention": 20}` to the template schema.
|
||||
- **New tests**: `tests/test_outputs_retention.py` — 13 tests across `TestGetRetention` (5) and `TestGcOutputs` (8). Pure-file-system; no subprocess, no live LLM.
|
||||
- **Full suite**: **482 passed** (was 469; +13 new; 0 regressions).
|
||||
|
||||
### Fixed — `parse_verdict` defensive coercion (task `harden-parse-verdict`)
|
||||
|
||||
- **`scripts/loop-runner.py`**: Hardened `parse_verdict` against malformed verifier output. Closes `add-loop-runner/BUG_REPORT.md` O6 (string-typed `pass`) and the un-noted sibling issue (unclamped `score`).
|
||||
- `pass` field: now accepts bool OR the strings `"true"`/`"false"` (case-insensitive, whitespace-stripped). The pre-fix `bool(data.get("pass"))` returned `True` for `"false"` (non-empty string is truthy) — a verifier emitting `{"pass": "false", "score": 0.1}` was recorded as `pass=True`, advancing the loop on a failed verdict. Non-`"true"`/`"false"` strings fall through to `bool(raw_pass.strip())` for backwards compat (`"yes"` stays truthy; `""` stays falsy; `None`/`0`/`[]`/`{}` keep v1's `bool(...)` semantics).
|
||||
- `score` field: now clamped to `[0, 1]` via `max(0.0, min(1.0, score))`. `NaN`, `Infinity`, and `-Infinity` (which `json.loads` accepts as bare tokens because `float(...)` happily returns `math.nan`/`inf`) default to `0.5` (neutral midpoint) via a `math.isfinite` guard. Non-numeric types (`None`, `[]`, `"high"`, etc.) and non-numeric strings default to `0.5` via a `try/except (TypeError, ValueError)` around `float(...)`.
|
||||
- Pure-function change; no CLI surface; no schema change; no new deps (`math` is stdlib). Strict emitters (those already emitting `bool pass` and `0.0 ≤ score ≤ 1.0`) are unaffected.
|
||||
- **`design/loops/technical.md` §7**: Documented the new coercion contract inline in the tick-flow step 7 (parse verdict).
|
||||
- **`design/loops/functional.md` §10**: Updated the Verifier Contract section to note the runner's defensive coercion for `pass` strings and `score` clamping.
|
||||
- **New tests**: `tests/test_parse_verdict.py` — 22 tests across 5 classes (`TestStrictBaseline`, `TestPassStringCoercion`, `TestScoreClamping`, `TestFenceBlockStillWorks`, `TestOptionalKeysPreserved`). Pure-functional; no subprocess, no live LLM.
|
||||
- **Adversarial probe**: 12-row sweep against `pass` as non-string non-bool types (`None`/`[]`/`{}`/`0`/`1`/`-1`/`1.5`/`[False]`/`[True]`), `score` as `[]`/`{}`/`[1, 2]`/`"high"`/numeric strings/JSON literal `NaN`/`Infinity`/`-Infinity`/`true`/`false`. All inputs yield deterministic, documented results; no crashes; no silent truthy/coercion regressions vs v1.
|
||||
- **Inline bug found and fixed during implementation**: the first iteration of the three-way branch set `verdict_pass = raw_pass.strip().lower() == "true"` — which mapped EVERY non-`"true"` string to False, breaking the SPEC R1 `bool(...)` fallback clause (a `"yes"` string would have become False, silently regressing v1). Fixed to the explicit `if "true" / elif "false" / else: bool(...)` form; test `test_other_truthy_string_pass` was added immediately to lock the contract.
|
||||
- **Full suite**: **469 passed** (was 447; +22 new; 0 regressions).
|
||||
|
||||
### Added — `.state.loop` file lock (task `add-state-loop-lock`)
|
||||
|
||||
- **`scripts/status.py`**: New `_loop_lock(loop_path, exclusive=True)` context manager wrapping the read-modify-write cycle of `.state.loop` with a cross-process file lock (POSIX `fcntl.flock(LOCK_EX)`, Windows `msvcrt.locking(LK_LOCK, 1)`) on `<loop_path>/.state.lock`. Closes the TOCTOU races flagged in `add-status-brakes/ADVERSARIAL_BUG_REPORT.md` (A6) and `add-loop-runner/ADVERSARIAL_BUG_REPORT.md` (A2, A7) where concurrent ticks (two scheduler firings on the same loop) or a concurrent `--pause-loop` / `--approve --loop` write could overwrite a tick's iteration increment or lose an approve's `resumed_count` increment. Per-loop granularity; blocking acquire; no timeout in v1.1 (operators notice a wedged tick via `--loop-list` stale `last_tick_at`).
|
||||
- **`scripts/status.py` callsites wrapped**: `cmd_pause_loop`, `cmd_resume_loop`, `cmd_approve_loop`, `cmd_check_gate` now acquire `_loop_lock` around their read-modify-write blocks. `--create-loop` intentionally unwrapped (no prior state to race against; create is name-unique-refused). `_disable_schedule` / `_enable_schedule` side-effect toggles moved OUTSIDE the lock to keep the critical section tight; the per-tick `--check-gate` self-skip on non-`running` state makes the transient ~100ms window harmless.
|
||||
- **`scripts/loop-runner.py`**: Mirrored `_loop_lock` helper (no env bypass — the runner is the lock holder). Wraps the entire `cmd_tick` critical section from the `_gate` subprocess call through step-10's state write. `_read_state_loop` is invoked twice: once unlocked (fast-fail untracked) and once inside the lock (re-read fresh to capture any concurrent mutate).
|
||||
- **Subprocess-deadlock avoidance (D-L6)**: When the runner spawns `status.py --check-gate` as a subprocess inside its held lock, status.py's own `cmd_check_gate` would otherwise deadlock waiting on the parent's held flock. Resolved by introducing the `$AUTOMATON_NO_LOOP_LOCK=1` env var, scoped ONLY to the `--check-gate` subprocess's env (set via `subprocess.run`'s `env` kwarg in `_gate`). status.py's `_loop_lock` checks this env var; if set, it yields without flocking (trusting the caller's outer lock). Harness subprocesses (Implement/Verify/Orchestrate) do NOT inherit the env var, so any `status.py --transition` calls the harness transitively invokes lock normally and serialize correctly.
|
||||
- **Scope of the env var**: `$AUTOMATON_NO_LOOP_LOCK` is read-only-internal; the runner never sets it in `os.environ` globally, only in the `_gate` subprocess's explicit env dict. Manual CLI users who set it in their shell bypass locking (documented escape hatch; same trust boundary as "shell user can kill the runner").
|
||||
- **`design/loops/technical.md` §7**: Added a "Lock serialization" subsection documenting the lock shape, the env-bypass mechanism, the re-entry forbidding contract, and the harness-contract implication (loop-control commands deadlethal inside a tick; task commands are safe).
|
||||
- **`AGENTS.md`**: Updated "State Enforcement — Loops (v1)" section to mention `_loop_lock` and the env var.
|
||||
- **`README.md`**: Added a row in the loop engineering monitoring table mentioning `.state.lock` per-loop serialization.
|
||||
- **Stdlib-only**: `fcntl` (POSIX) and `msvcrt` (Windows) are stdlib. No `filelock` package. No new pip deps.
|
||||
- **Backwards-compatible**: existing scripts that don't invoke `--pause-loop` / `--resume-loop` / `--approve --loop` / `--check-gate` concurrently are unaffected. Primary lock surface is per-tick blocking when schedulers fire on the same loop concurrently.
|
||||
- **New tests**: `tests/test_state_loop_lock.py` -- 7 tests covering (1) concurrent-acquire serialization, (2) clean-release + re-acquire, (3) release-on-exception, (4) per-loop granularity, (5) `--create-loop` doesn't create `.state.lock` (D-L3) but `--check-gate` does, (6) `--pause-loop` blocks under concurrent lock holder, (7) `cmd_tick` holds the lock across the harness subprocess while `--approve --loop` blocks. All use `tmp_path`, stdlib only, no live LLM.
|
||||
- **Adversarial findings**: A1 (concurrent approve: race closed, only 1 of 5 won, `resumed_count` incremented once), A2 (concurrent pause: idempotent but consistent), A3 (lock releases on mid-tick exception: verified), A4 (harness calls loop-control inside tick: theoretical deadlock, LOW severity, deferred to harness-integration contract docs follow-up), A5 (manual env var escape hatch: documented), A6 (orphaned `.state.lock` after crash: benign; POSIX auto-releases flock on process exit), A7 (NFS: documented assumption), A8 (cyclomatic complexity: acceptable). No blockers.
|
||||
- **Full suite**: **447 passed** (was 440; +7 new; 0 regressions).
|
||||
|
||||
### Fixed — harness command template (task `fix-harness-command-template`)
|
||||
|
||||
- **`scripts/loop-runner.py`**: Fixed the broken default `harness.command`. The old default `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]` used flags that do not exist in `opencode run` (`--prompt-file`, `--cwd`) — the v1 runner has only been exercised via mocked subprocess tests, so the bug was never caught. New default is `["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]`. Introduces a new `{prompt_content}` substitution token that carries the resolved prompt's text as a single argv element under `subprocess.run` list mode (no shell expansion, safe for prompts containing quotes/special chars).
|
||||
- **`{prompt}` and `{cwd}` tokens retained**: backwards-compatible. Users with custom `harness.command` in `loop.json` using these tokens are unaffected.
|
||||
- **`UnicodeDecodeError` now caught** alongside `OSError` when reading the resolved prompt file (defensive; falls back to empty prompt content rather than crashing mid-tick).
|
||||
- **Harness-agnostic contract preserved and extended**: runner core has zero harness awareness; the new `{prompt_content}` token covers harnesses that prefer a message argument (Pi Dev, aider, any CLI taking a prompt as a positional). Non-opencode users override `harness.command` in `loop.json` (e.g. `["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]`). D8 (no model/provider inspection) intact.
|
||||
- **`design/loops/technical.md` §8 and §9**: Updated default command shape; documented the new `{prompt_content}` token alongside existing `{prompt}`/`{cwd}`/`{output}`/`{artifact}` tokens; added Pi Dev, aider, and generic shell-wrapper examples in `loop.json` form.
|
||||
- **`templates/loops/self-improvement/loop.json`**: Updated `harness.command` to the new default.
|
||||
- **Test scaffolding updates**: `_make_loop` helpers in `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py`, `tests/test_loop_templates.py` now write loop-local prompt stubs (containing role-marker content `prompt: <ref>`) when the framework prompt at `~/.automaton/prompts/<ref>` does not already exist. This preserves the substring-matcher strategy used by tick-flow tests while not overriding real framework prompts (which contain `{current_task}` etc. substitution tokens). Custom-`{prompt}`-command tests in `test_goal_mode.py` updated their matcher substrings from `test-impl`/`test-verify`/`test-orch` to `implement-prompt`/`verify-prompt`/`orchestrate-prompt` (the resolved temp-file path contains those markers).
|
||||
- **New tests**: `tests/test_harness_command.py` -- 7 tests covering: default uses `--dir` not `--cwd`/`--prompt-file`; default passes prompt content as last argv element; prompt content preserves single quotes/double quotes/dollar signs as a single argv element; `{prompt}` token still available for custom commands; `{cwd}` token still works in custom commands; empty `command` falls back to the new default; Pi Dev-shaped command substitution (proves the substitution mechanism is harness-agnostic — any binary works).
|
||||
- **Backwards-compat correction**: Updated the v1 `add-loop-runner` CHANGELOG bullet that documented the old default — no retroactive edit to the released entry shape, but the new `### Fixed` entry supersedes the default-command wording. The default command works end-to-end now.
|
||||
- Full suite: **440 passed** (was 433 baseline; +7 new).
|
||||
|
||||
### Changed — completed tasks moved to `tasks/complete/` (task `move-completed-tasks-to-complete-folder`)
|
||||
|
||||
- **`scripts/status.py`**: `--transition complete` now moves the task directory from `tasks/<name>/` to `tasks/complete/<name>/`. `_task_dir` has a fallback to find completed tasks. `--list` and `--audit` exclude completed tasks (visible only via `--task <name>` fallback).
|
||||
- **New tests**: `tests/test_move_completed.py` -- 9 tests covering directory move, fallback, listing exclusion, create-task refusal, and transition refusal from complete. Full suite: 433 passed.
|
||||
|
||||
### Fixed — install/update flow (task `fix-install-update-flow`)
|
||||
|
||||
- **`scripts/install.sh`**: Replaced hardcoded private git URL with user-supplied `GIT_URL="${1:-}"` argument (D11). Refuses with usage and irreversibility warning if absent. Fixed `.venv` cwd bug: venv now created in `$FRAMEWORK_DIR/.venv` instead of CWD. Added Windows venv path support (`.venv/Scripts/python.exe`). Uses `"$VENV_PY" -m pip` for cross-platform pip invocation. Added `status.py --version` smoke test after install.
|
||||
- **`scripts/update.sh`**: Changed hook installation from `ln -sf` (symlink) to `cp` + `chmod +x` (copy), matching `install-hooks.sh`.
|
||||
- **`scripts/upgrade.sh`**: Replaced all `ln -sf` and symlink-checking logic (`readlink`, `-L`) with `cp` + `chmod +x`. Simplified hook-exists warning to point to `install-hooks.sh`.
|
||||
- **New tests**: `tests/test_install_update_flow.py` -- 15 tests covering git URL, venv paths, version check, and hook consistency. Full suite: 424 passed.
|
||||
|
||||
### Added — loop engineering v1, self-improvement loop default-on (task `add-self-improvement-loop`)
|
||||
|
||||
- **`scripts/install.sh`**: after clone and guard registration, creates and schedules the self-improvement loop (`--create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` + `--install-schedule self-improvement --interval 3600`). Both commands use `|| true` so the framework continues to work even if loop creation fails. Prints a user-facing message with opt-out instructions (`--pause-loop self-improvement`).
|
||||
- **`scripts/update.sh`**: idempotent bootstrap for existing users. Checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]` before creating. Same `|| true` non-fatal behavior.
|
||||
- **New tests**: `tests/test_self_improvement_loop.py` -- 16 tests covering install.sh wiring (5), update.sh wiring (4), template fields (5), and loop creation from template (2). Full suite: 409 passed.
|
||||
- **Doc updates**: `README.md` Loop Engineering section notes default-on; `design/loops/technical.md` section 9 notes install.sh creates it.
|
||||
|
||||
### Added — loop engineering v1, templates and onboarding (task `add-loop-templates-onboarding`)
|
||||
|
||||
- **Prompt-file token substitution**: `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)` in `scripts/loop-runner.py` resolves prompt refs to full paths (searches `<loop_path>/<ref>` then `~/.automaton/prompts/<ref>`), reads the file, substitutes content-level tokens (`{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`), writes to `<loop_path>/outputs/tickN-<role>-prompt.md`, and returns the temp path. Falls back to raw `prompt_ref` when file not found (backward compat). `{artifact_content}` reads the file at `extras["artifact"]` and substitutes its content (empty string if missing).
|
||||
- **`_invoke_harness` extended**: new optional `loop_path` and `tick_num` params. When `loop_path` is provided, calls `_resolve_prompt` before building the harness command. All three call sites in `cmd_tick` (implement, verify, orchestrate) now pass `loop_path` and `tick_num`.
|
||||
- **New prompt files**: `prompts/loop-implement.md` (Implement role with task_brief/criteria/hint tokens, ALLOWED/FORBIDDEN sections), `prompts/loop-verifier.md` (Verify role with strict JSON output, score rubric 0.0-1.0, artifact_content token), `prompts/loop-orchestrate.md` (Orchestrate role with verdict token, phase transition logic, no-edit/no-auto-approve rules).
|
||||
- **`templates/loops/ci-triage/loop.json`**: `roles` filled from `null` to `{"prompt": "loop-implement.md"}` etc.
|
||||
- **`templates/loops/self-improvement/loop.json`**: new template with `work_source: audit`, `blast_radius.use_worktree: true`, `file_scope: ["scripts/", "prompts/", "tests/", "design/"]`, `brakes.max_iterations: 10`, `score_plateau_window: 3`.
|
||||
- **README.md**: new "Loop Engineering" onboarding section (quick start, tick cycle, configuration, monitoring, halt/resume).
|
||||
- **`design/loops/technical.md` section 8**: documented prompt resolution and token substitution flow.
|
||||
- **New tests**: `tests/test_loop_templates.py` -- 18 tests covering R1-R6 (resolve_prompt, prompt file content, ci-triage template, self-improvement template, tick integration). Full suite: 393 passed.
|
||||
- **Test infrastructure**: updated `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py` to use non-existent prompt refs (`test-impl.md` etc.) so `_resolve_prompt` fallback path is exercised in those tests. Added loop prompts to self-consistency test exclusion set.
|
||||
|
||||
### Added — loop engineering v1, blast-radius scheduler (task `add-blast-radius-scheduler`)
|
||||
|
||||
- **`scripts/loop-runner.py` extended**: `_ensure_worktree(state, cfg, loop_path, project_dir)` creates a per-loop git worktree at `<loop>/worktree` on branch `loop/<name>` when `blast_radius.use_worktree` is true (default) and no worktree exists yet. Records `worktree_path` and `worktree_branch` in `.state.loop` atomically. Reuses existing worktree on subsequent ticks. Falls back to project root with a WARNING log when: not a git repo, git binary missing, or `git worktree add` fails. Handles branch-already-exists by retrying without `-b`. Stale `worktree_path` (directory deleted) is cleared and worktree recreated.
|
||||
- **New tests**: `tests/test_blast_radius.py` -- 15 tests covering worktree creation, reuse, fallback, branch-exists retry, state consistency, tick integration, platform paths, and backward compat. Full suite: 369 passed (was 354 + 15 new).
|
||||
|
||||
### Added — loop engineering v1, goal-mode (task `add-goal-mode`)
|
||||
|
||||
- **`scripts/loop-runner.py` extended**: `_find_work(state, cfg, loop_path, project_dir)` replaces the inline `single`-only block, dispatching on `loop.json` `work_source.kind`:
|
||||
- `"single"`: unchanged behavior (uses `.state.loop` `current_task`).
|
||||
- `"audit"`: runs `status.py --audit --json --project <p>`, picks the highest-severity unresolved violation. If the violation has a task, uses it; otherwise slugifies the message and calls `--create-task`, then sets the new task as `current_task`. When no unresolved violations → SKIP `no_work` (clean scheduler exit; does not increment `iteration_count`). Optional `work_source.project` overrides the audit target.
|
||||
- `"backlog"`: reads `<root>/design/<area>/BACKLOG.md` (`work_source.area` defaults to `"loops"`); picks the topmost `- [ ]` item. Empty backlog → SKIP `no_work`.
|
||||
- **Goal-oriented substitution tokens**: three new tokens available in `harness.command`:
|
||||
- `{task_brief}` -- from `<task>/RESEARCH.md` or `DESIGN.md` or `SPEC.md` (first hit), capped at 4k tokens.
|
||||
- `{acceptance_criteria}` -- from `loop.json` `acceptance_criteria` (string OR list joined by newlines), capped at 2k tokens.
|
||||
- `{next_hint}` -- from `state.last_verdict.next_hint` (empty on first tick / after `--approve`), capped at 1k tokens.
|
||||
- Closes the `next_hint` feedback loop: tick N's verifier hint becomes tick N+1's Implement/Verify context (D7 graded verifier feedback).
|
||||
- **`_truncate_tokens(text, max_tokens)`** stdlib-only approximate cap (4-chars-per-token heuristic, `…[truncated]` marker). No tokenizer dependency.
|
||||
- **`scripts/status.py --audit --json`**: machine-readable audit mode. Emits a single JSON line on stdout: `{"violations": [...], "loops": [...], "total_tasks": N, "untracked_tasks": M}`. Each violation carries `category`, `severity` (`high`/`med`/`low`), `task`, `message`, `resolved: false`. Existing human-readable `--audit` output is unchanged when `--json` is absent. Backed by new `_audit_collect` and `_audit_category3_paths` helpers.
|
||||
- **`templates/loops/ci-triage/loop.json`**: added `"work_source": {"kind": "single"}` and an `"acceptance_criteria"` example so the template is self-documenting.
|
||||
- **Backward compat**: missing `work_source` or unknown `kind` falls back to `"single"` with a `.state.log` WARNING entry. Existing task-3 fixtures tick identically.
|
||||
- **New tests**: `tests/test_goal_mode.py` -- 26 tests covering R1-R8 and one regression. All subprocess calls stubbed via `monkeypatch`; no live LLM in CI. Full suite: 354 passed (was 328 + 26 new).
|
||||
- **Doc updates**: `design/loops/technical.md` §9 self-improvement template now shows `acceptance_criteria`; `design/loops/functional.md` §9 loop.json fields list documents `work_source` and `acceptance_criteria`.
|
||||
|
||||
### Added — loop engineering v1, loop runner (task `add-loop-runner`)
|
||||
|
||||
- **`scripts/loop-runner.py`**: per-tick engine. `--mode tick` runs one tick (gate -> find work -> spawn Implement -> spawn Verify -> parse graded JSON verdict -> spawn Orchestrate -> atomic state write -> log); `--mode daemon` runs a sleep loop bounded by `--max-iterations`. Calls `status.py --check-gate` first; any non-ok gate exits 0 (clean scheduler exit). Calls `vram_detect.py --loop-mode --json` before any harness subprocess; refuses below the 16k floor (D13) with `human_intervention`. Idempotent in failure -- pre-step-10 crashes do not corrupt `.state.loop` or advance `iteration_count`.
|
||||
- **Graded verifier protocol**: `parse_verdict` accepts raw JSON, ```json fenced blocks, or JSON with `//`/`#` line comments. Required keys: `pass` (bool), `score` (float). Optional: `reasons` (list[str]), `next_hint` (str). Parse failure halts as `verifier_failed` without advancing state.
|
||||
- **Harness command substitution**: `loop.json` `harness.command` (list of strings) with tokens `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`. Default command (v1.1, see `fix-harness-command-template` above): `["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]` (old default with `--prompt-file`/`--cwd` was non-functional; fixed in v1.1). Stdlib only, no new pip deps.
|
||||
- **Score history capping**: per-tick score appended to `score_history`, capped at `brakes.score_plateau_window` (oldest dropped). Score plateau halts via the next tick's `--check-gate`.
|
||||
- **Tick log**: every state-changing op appends an ISO-timestamped `TICK pass=<bool> score=<f> iter=<N>` line. Daemon mode appends `DAEMON_STOPPED` on KeyboardInterrupt.
|
||||
- **New tests**: `tests/test_loop_runner.py` -- 18 tests across 7 classes cover R1--R8. All subprocess calls stubbed via monkeypatch (no live LLM in CI). Full suite: 328 passed (was 310 + 18 new).
|
||||
- **Doc updates**: `AGENTS.md` "Loop runner" bullet under State Enforcement -- Loops (v1); `README.md` loop-runner one-liner in the Loop Engineering (beta) section.
|
||||
|
||||
### Added — loop engineering v1, brakes layer (task `add-status-brakes`)
|
||||
|
||||
- **`.state.loop` runtime state**: single source of truth per loop at `{project}/.automaton/loops/<name>/.state.loop` (framework-internal at `~/.automaton/loops/<name>/`). Schema v1 with 13 fields (`status`, `halt_reason`, `iteration_count`, `resumed_count`, `last_tick_at`, `last_verdict`, `score_history`, `current_task`, `worktree_branch`, `worktree_path`). Atomic tmp+rename writes.
|
||||
- **`--create-loop NAME [--from-template T]`**: only way to bootstrap a loop dir + `.state.loop` + `loop.json`. Refuses non-kebab names, duplicates, unknown templates; patches `name` into the copied `loop.json`.
|
||||
- **`--check-gate NAME [--json]`**: runs 6 brake gates in order (status, iterations, budget, task phase, worktree drift, score plateau). First failure halts the loop, best-effort disables the OS schedule unit, emits structured JSON verdict.
|
||||
- **`--can-continue NAME [--json]`**: cheap pre-tick probe — `ok := status == "running"`.
|
||||
- **`--approve --loop NAME`**: the only way to clear a halt. Increments `resumed_count`. No auto-approve in v1 (D4).
|
||||
- **`--pause-loop` / `--resume-loop`**: user-controlled soft stop; cannot clear halts.
|
||||
- **`--install-schedule NAME [--interval S]`**: triple-dispatched (Darwin launchd plist / Linux crontab block / Windows schtasks) native schedule unit generator per `platform.system()`. Generates `automaton-loop-tick.sh` / `automaton-loop-tick.bat` stub (self-documenting: when the scheduler fires it, the filename alone says what it does).
|
||||
- **`--can-edit --loop NAME [--loop-worktree] --file P`**: worktree scope check — is the file inside this loop's declared `blast_radius.file_scope`?
|
||||
- **`--transition` halt refusal**: refuses to transition any task owned by a HALTED loop until `--approve --loop` clears the halt.
|
||||
- **`--audit` Cat-6 Loops block + `--loop-list`**: flags halted / untracked / stale-running loops and loops whose `current_task` no longer exists. `_audit_loops_block` runs even when no tasks are present.
|
||||
- **`--version`**: prints framework version parsed from `config.md`'s `## Framework Version` section.
|
||||
- **`.state.log` tick trail**: every state-changing loop op appends an ISO-timestamped line (PAUSED / RESUMED / APPROVED / HALT).
|
||||
- **New loop template**: `templates/loops/ci-triage/loop.json` (minimal; full template expansion lands in task `add-loop-templates-onboarding`).
|
||||
- **New tests**: `tests/test_status_brakes.py` — 46 tests across 10 classes cover R1–R10. Full suite: 310 passed (was 264 + 46 new).
|
||||
- **Doc updates**: `AGENTS.md` Harness Integration modes block extended with the `--loop` worktree-scope mode; new "State Enforcement — Loops (v1)" section in `AGENTS.md`.
|
||||
|
||||
### Added — loop engineering v1 (design only; implementation pending bootstrap tasks)
|
||||
|
||||
- **Loop system design**: `design/loops/{functional,technical,README,BACKLOG}.md` — locked v1 design for an unattended, state-enforced loop runner built on top of the existing `status.py` phase machine. No second enforcement surface.
|
||||
- **Five deaths halt model**: `iterations_exhausted`, `budget_exhausted`, `verifier_failed`, `drift_detected`, `human_intervention`. All halts require human `--approve --loop` to resume; no auto-approve path (D4).
|
||||
- **Tiered role budgets**: three session roles (`Implement:`, `Verify:`, `Orchestrate:`) with explicit context tiers. 16k floor hard refuse (D13). Session divergence mandatory; model divergence only when `Verify:` != `Implement:` (D12).
|
||||
- **Per-loop git worktree** as blast radius default (D2); `--no-worktree` opt-out.
|
||||
- **Native scheduler generator**: `platform.system()`-dispatched `launchd` / `crontab` / `schtasks` unit generation via `status.py --install-schedule` (D1); `--daemon` opt-in fallback.
|
||||
- **Self-improvement loop template** at `templates/loops/self-improvement/`, **default-on at install** (D21). Ticks against `status.py --audit` on the framework's own repo — the literal seed of self-management.
|
||||
- **Bootstrap task plan** (8 tasks) committed to `design/loops/README.md` — `fix-context-sizing`, `add-status-brakes`, `add-loop-runner`, `add-goal-mode`, `add-blast-radius-scheduler`, `add-loop-templates-onboarding`, `add-self-improvement-loop`, `fix-install-update-flow`. These are the **last tasks a human creates by hand**; after task 7 lands, loops create subsequent tasks from `BACKLOG.md` and `--audit` output (D24, D25).
|
||||
- **Scoped deferrals** recorded in `BACKLOG.md`: Scope 2 (design-update loop) → v1.1; Scope 3 (self-designing loops) + parallel-mode-default + auto-approve-relax → deferred indefinitely (D20).
|
||||
|
||||
### Added
|
||||
- **Added:** `requirements.txt` pinning `pytest==7.4.4` for reproducible test runs.
|
||||
- **Added:** `scripts/install.sh` now creates `.venv/` and installs pytest into it.
|
||||
|
||||
Reference in New Issue
Block a user