Files
automaton/CHANGELOG.md
T

47 KiB
Raw Blame History

Changelog

[unreleased]

Added — cross-loop task claim (task add-claim-loop-task)

  • scripts/status.py: New --claim-loop-task <name> --task <taskname> [--project P] command. Exit 0 = claimed (or already self-claimed, idempotent). Exit 2 = already claimed by another running/paused loop (task_already_claimed:{other} on stderr) or untracked loop (loop_untracked). Uses _loop_lock to serialize writes; cross-loop scan is advisory (self-healing on next tick).
  • scripts/loop-runner.py: cmd_tick step 3.5: after _find_work returns a candidate different from current_task, spawn status.py --claim-loop-task as subprocess with $AUTOMATON_NO_LOOP_LOCK=1 (same bypass pattern as _gate). Step 9.5: after orchestrator, re-read .state; if phase is complete or human_intervention, clear current_task to None. Closes add-status-brakes/ADVERSARIAL_BUG_REPORT.md A2 (cross-loop task race).
  • New tests: tests/test_claim_loop_task.py — 10 tests covering claim success, refusal, idempotency, untracked loop, missing task, paused-loop ownership, self-healing race, and release on terminal phases.
  • Full suite: 518 passed (was 508; +10 new; 0 regressions).

Added — Linux schedule parity (task linux-schedule-parity)

  • scripts/status.py: New _install_cron_block(name, project, interval_seconds) for Linux cron support. Writes # automaton-loop:<name> / # end automaton-loop:<name> blocks into crontab via crontab -. Interval rounded to full minutes, minimum 1. Strips prior block before insert (idempotent).
  • New _enable_schedule(name, project): platform dispatch — Linux (cron insert via _install_cron_block), Darwin (rename .plist.disabled back), Windows (no-op). Reads loop.json > schedule.interval_seconds; non-int falls back to 3600.
  • _disable_schedule already existed; verified correct.
  • New tests: tests/test_linux_schedule_parity.py — 13 tests covering cron block writes, error handling, platform dispatch, and interval edge cases.
  • Full suite: 508 passed (was 495; +13 new; 0 regressions).

Fixed — parametrize-base-branch (task parametrize-base-branch)

  • scripts/status.py: Replaced hardcoded "main" in _gate_worktree_drift's git diff call with _base_branch(cfg) -> str. New helper returns blast_radius.base_branch if configured (default "main"). Empty string and non-string types produce WARNING and fall back to "main". Closes add-status-brakes/BUG_REPORT.md O3: projects on master/trunk/develop no longer have a silently-disabled drift gate — operator sets "base_branch": "master" in loop.json.
  • templates/loops/self-improvement/loop.json: Added "base_branch": "main" to blast_radius.
  • design/loops/functional.md §9: Updated blast_radius field list to include base_branch.
  • New tests: tests/test_base_branch.py — 13 tests covering _base_branch helper (5) and drift-gate branch usage (8, with mocked subprocess.run).
  • Full suite: 495 passed (was 482; +13 new; 0 regressions).

Added — outputs retention GC (task add-outputs-retention)

  • scripts/loop-runner.py: Added _get_retention(cfg) and _gc_outputs(loop_path, retention) to bound outputs/ directory growth. Every tick (cmd_tick step 10.5, inside _loop_lock), older tick groups are deleted, keeping only the last N (default 20). Closes add-loop-runner/BUG_REPORT.md O5 (tick dirs accumulate without bound).
    • _get_retention reads cfg.get("outputs", {}).get("retention", 20). Non-int types fall back to 20 with WARNING. Negative values coerce to 0 (unlimited) with WARNING. 0 = no GC (v1 behavior).
    • _gc_outputs lists outputs/, parses tickNN indices via ^tick(\d+)- regex, computes cutoff = max_seen - retention + 1, deletes files with tick index < cutoff. Non-tick files (e.g. README.txt, tick-foo.md) are preserved. Errors (permission, missing file) are logged as WARNING and swallowed — GC failure never crashes the tick.
    • Off-by-one bug found and fixed inline: initial formula cutoff = max_seen - retention kept retention+1 groups; caught by test_gc_keeps_recent_deletes_old length assertion.
  • templates/loops/self-improvement/loop.json: Added "outputs": {"retention": 20} to the template schema.
  • New tests: tests/test_outputs_retention.py — 13 tests across TestGetRetention (5) and TestGcOutputs (8). Pure-file-system; no subprocess, no live LLM.
  • Full suite: 482 passed (was 469; +13 new; 0 regressions).

Fixed — parse_verdict defensive coercion (task harden-parse-verdict)

  • scripts/loop-runner.py: Hardened parse_verdict against malformed verifier output. Closes add-loop-runner/BUG_REPORT.md O6 (string-typed pass) and the un-noted sibling issue (unclamped score).
    • pass field: now accepts bool OR the strings "true"/"false" (case-insensitive, whitespace-stripped). The pre-fix bool(data.get("pass")) returned True for "false" (non-empty string is truthy) — a verifier emitting {"pass": "false", "score": 0.1} was recorded as pass=True, advancing the loop on a failed verdict. Non-"true"/"false" strings fall through to bool(raw_pass.strip()) for backwards compat ("yes" stays truthy; "" stays falsy; None/0/[]/{} keep v1's bool(...) semantics).
    • score field: now clamped to [0, 1] via max(0.0, min(1.0, score)). NaN, Infinity, and -Infinity (which json.loads accepts as bare tokens because float(...) happily returns math.nan/inf) default to 0.5 (neutral midpoint) via a math.isfinite guard. Non-numeric types (None, [], "high", etc.) and non-numeric strings default to 0.5 via a try/except (TypeError, ValueError) around float(...).
    • Pure-function change; no CLI surface; no schema change; no new deps (math is stdlib). Strict emitters (those already emitting bool pass and 0.0 ≤ score ≤ 1.0) are unaffected.
  • design/loops/technical.md §7: Documented the new coercion contract inline in the tick-flow step 7 (parse verdict).
  • design/loops/functional.md §10: Updated the Verifier Contract section to note the runner's defensive coercion for pass strings and score clamping.
  • New tests: tests/test_parse_verdict.py — 22 tests across 5 classes (TestStrictBaseline, TestPassStringCoercion, TestScoreClamping, TestFenceBlockStillWorks, TestOptionalKeysPreserved). Pure-functional; no subprocess, no live LLM.
  • Adversarial probe: 12-row sweep against pass as non-string non-bool types (None/[]/{}/0/1/-1/1.5/[False]/[True]), score as []/{}/[1, 2]/"high"/numeric strings/JSON literal NaN/Infinity/-Infinity/true/false. All inputs yield deterministic, documented results; no crashes; no silent truthy/coercion regressions vs v1.
  • Inline bug found and fixed during implementation: the first iteration of the three-way branch set verdict_pass = raw_pass.strip().lower() == "true" — which mapped EVERY non-"true" string to False, breaking the SPEC R1 bool(...) fallback clause (a "yes" string would have become False, silently regressing v1). Fixed to the explicit if "true" / elif "false" / else: bool(...) form; test test_other_truthy_string_pass was added immediately to lock the contract.
  • Full suite: 469 passed (was 447; +22 new; 0 regressions).

Added — .state.loop file lock (task add-state-loop-lock)

  • scripts/status.py: New _loop_lock(loop_path, exclusive=True) context manager wrapping the read-modify-write cycle of .state.loop with a cross-process file lock (POSIX fcntl.flock(LOCK_EX), Windows msvcrt.locking(LK_LOCK, 1)) on <loop_path>/.state.lock. Closes the TOCTOU races flagged in add-status-brakes/ADVERSARIAL_BUG_REPORT.md (A6) and add-loop-runner/ADVERSARIAL_BUG_REPORT.md (A2, A7) where concurrent ticks (two scheduler firings on the same loop) or a concurrent --pause-loop / --approve --loop write could overwrite a tick's iteration increment or lose an approve's resumed_count increment. Per-loop granularity; blocking acquire; no timeout in v1.1 (operators notice a wedged tick via --loop-list stale last_tick_at).
  • scripts/status.py callsites wrapped: cmd_pause_loop, cmd_resume_loop, cmd_approve_loop, cmd_check_gate now acquire _loop_lock around their read-modify-write blocks. --create-loop intentionally unwrapped (no prior state to race against; create is name-unique-refused). _disable_schedule / _enable_schedule side-effect toggles moved OUTSIDE the lock to keep the critical section tight; the per-tick --check-gate self-skip on non-running state makes the transient ~100ms window harmless.
  • scripts/loop-runner.py: Mirrored _loop_lock helper (no env bypass — the runner is the lock holder). Wraps the entire cmd_tick critical section from the _gate subprocess call through step-10's state write. _read_state_loop is invoked twice: once unlocked (fast-fail untracked) and once inside the lock (re-read fresh to capture any concurrent mutate).
  • Subprocess-deadlock avoidance (D-L6): When the runner spawns status.py --check-gate as a subprocess inside its held lock, status.py's own cmd_check_gate would otherwise deadlock waiting on the parent's held flock. Resolved by introducing the $AUTOMATON_NO_LOOP_LOCK=1 env var, scoped ONLY to the --check-gate subprocess's env (set via subprocess.run's env kwarg in _gate). status.py's _loop_lock checks this env var; if set, it yields without flocking (trusting the caller's outer lock). Harness subprocesses (Implement/Verify/Orchestrate) do NOT inherit the env var, so any status.py --transition calls the harness transitively invokes lock normally and serialize correctly.
  • Scope of the env var: $AUTOMATON_NO_LOOP_LOCK is read-only-internal; the runner never sets it in os.environ globally, only in the _gate subprocess's explicit env dict. Manual CLI users who set it in their shell bypass locking (documented escape hatch; same trust boundary as "shell user can kill the runner").
  • design/loops/technical.md §7: Added a "Lock serialization" subsection documenting the lock shape, the env-bypass mechanism, the re-entry forbidding contract, and the harness-contract implication (loop-control commands deadlethal inside a tick; task commands are safe).
  • AGENTS.md: Updated "State Enforcement — Loops (v1)" section to mention _loop_lock and the env var.
  • README.md: Added a row in the loop engineering monitoring table mentioning .state.lock per-loop serialization.
  • Stdlib-only: fcntl (POSIX) and msvcrt (Windows) are stdlib. No filelock package. No new pip deps.
  • Backwards-compatible: existing scripts that don't invoke --pause-loop / --resume-loop / --approve --loop / --check-gate concurrently are unaffected. Primary lock surface is per-tick blocking when schedulers fire on the same loop concurrently.
  • New tests: tests/test_state_loop_lock.py -- 7 tests covering (1) concurrent-acquire serialization, (2) clean-release + re-acquire, (3) release-on-exception, (4) per-loop granularity, (5) --create-loop doesn't create .state.lock (D-L3) but --check-gate does, (6) --pause-loop blocks under concurrent lock holder, (7) cmd_tick holds the lock across the harness subprocess while --approve --loop blocks. All use tmp_path, stdlib only, no live LLM.
  • Adversarial findings: A1 (concurrent approve: race closed, only 1 of 5 won, resumed_count incremented once), A2 (concurrent pause: idempotent but consistent), A3 (lock releases on mid-tick exception: verified), A4 (harness calls loop-control inside tick: theoretical deadlock, LOW severity, deferred to harness-integration contract docs follow-up), A5 (manual env var escape hatch: documented), A6 (orphaned .state.lock after crash: benign; POSIX auto-releases flock on process exit), A7 (NFS: documented assumption), A8 (cyclomatic complexity: acceptable). No blockers.
  • Full suite: 447 passed (was 440; +7 new; 0 regressions).

Fixed — harness command template (task fix-harness-command-template)

  • scripts/loop-runner.py: Fixed the broken default harness.command. The old default ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"] used flags that do not exist in opencode run (--prompt-file, --cwd) — the v1 runner has only been exercised via mocked subprocess tests, so the bug was never caught. New default is ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]. Introduces a new {prompt_content} substitution token that carries the resolved prompt's text as a single argv element under subprocess.run list mode (no shell expansion, safe for prompts containing quotes/special chars).
  • {prompt} and {cwd} tokens retained: backwards-compatible. Users with custom harness.command in loop.json using these tokens are unaffected.
  • UnicodeDecodeError now caught alongside OSError when reading the resolved prompt file (defensive; falls back to empty prompt content rather than crashing mid-tick).
  • Harness-agnostic contract preserved and extended: runner core has zero harness awareness; the new {prompt_content} token covers harnesses that prefer a message argument (Pi Dev, aider, any CLI taking a prompt as a positional). Non-opencode users override harness.command in loop.json (e.g. ["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]). D8 (no model/provider inspection) intact.
  • design/loops/technical.md §8 and §9: Updated default command shape; documented the new {prompt_content} token alongside existing {prompt}/{cwd}/{output}/{artifact} tokens; added Pi Dev, aider, and generic shell-wrapper examples in loop.json form.
  • templates/loops/self-improvement/loop.json: Updated harness.command to the new default.
  • Test scaffolding updates: _make_loop helpers in tests/test_loop_runner.py, tests/test_blast_radius.py, tests/test_goal_mode.py, tests/test_loop_templates.py now write loop-local prompt stubs (containing role-marker content prompt: <ref>) when the framework prompt at ~/.automaton/prompts/<ref> does not already exist. This preserves the substring-matcher strategy used by tick-flow tests while not overriding real framework prompts (which contain {current_task} etc. substitution tokens). Custom-{prompt}-command tests in test_goal_mode.py updated their matcher substrings from test-impl/test-verify/test-orch to implement-prompt/verify-prompt/orchestrate-prompt (the resolved temp-file path contains those markers).
  • New tests: tests/test_harness_command.py -- 7 tests covering: default uses --dir not --cwd/--prompt-file; default passes prompt content as last argv element; prompt content preserves single quotes/double quotes/dollar signs as a single argv element; {prompt} token still available for custom commands; {cwd} token still works in custom commands; empty command falls back to the new default; Pi Dev-shaped command substitution (proves the substitution mechanism is harness-agnostic — any binary works).
  • Backwards-compat correction: Updated the v1 add-loop-runner CHANGELOG bullet that documented the old default — no retroactive edit to the released entry shape, but the new ### Fixed entry supersedes the default-command wording. The default command works end-to-end now.
  • Full suite: 440 passed (was 433 baseline; +7 new).

Changed — completed tasks moved to tasks/complete/ (task move-completed-tasks-to-complete-folder)

  • scripts/status.py: --transition complete now moves the task directory from tasks/<name>/ to tasks/complete/<name>/. _task_dir has a fallback to find completed tasks. --list and --audit exclude completed tasks (visible only via --task <name> fallback).
  • New tests: tests/test_move_completed.py -- 9 tests covering directory move, fallback, listing exclusion, create-task refusal, and transition refusal from complete. Full suite: 433 passed.

Fixed — install/update flow (task fix-install-update-flow)

  • scripts/install.sh: Replaced hardcoded private git URL with user-supplied GIT_URL="${1:-}" argument (D11). Refuses with usage and irreversibility warning if absent. Fixed .venv cwd bug: venv now created in $FRAMEWORK_DIR/.venv instead of CWD. Added Windows venv path support (.venv/Scripts/python.exe). Uses "$VENV_PY" -m pip for cross-platform pip invocation. Added status.py --version smoke test after install.
  • scripts/update.sh: Changed hook installation from ln -sf (symlink) to cp + chmod +x (copy), matching install-hooks.sh.
  • scripts/upgrade.sh: Replaced all ln -sf and symlink-checking logic (readlink, -L) with cp + chmod +x. Simplified hook-exists warning to point to install-hooks.sh.
  • New tests: tests/test_install_update_flow.py -- 15 tests covering git URL, venv paths, version check, and hook consistency. Full suite: 424 passed.

Added — loop engineering v1, self-improvement loop default-on (task add-self-improvement-loop)

  • scripts/install.sh: after clone and guard registration, creates and schedules the self-improvement loop (--create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR" + --install-schedule self-improvement --interval 3600). Both commands use || true so the framework continues to work even if loop creation fails. Prints a user-facing message with opt-out instructions (--pause-loop self-improvement).
  • scripts/update.sh: idempotent bootstrap for existing users. Checks if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ] before creating. Same || true non-fatal behavior.
  • New tests: tests/test_self_improvement_loop.py -- 16 tests covering install.sh wiring (5), update.sh wiring (4), template fields (5), and loop creation from template (2). Full suite: 409 passed.
  • Doc updates: README.md Loop Engineering section notes default-on; design/loops/technical.md section 9 notes install.sh creates it.

Added — loop engineering v1, templates and onboarding (task add-loop-templates-onboarding)

  • Prompt-file token substitution: _resolve_prompt(prompt_ref, extras, loop_path, tick_num, role) in scripts/loop-runner.py resolves prompt refs to full paths (searches <loop_path>/<ref> then ~/.automaton/prompts/<ref>), reads the file, substitutes content-level tokens ({task_brief}, {acceptance_criteria}, {next_hint}, {current_task}, {current_phase}, {verdict}, {artifact_content}), writes to <loop_path>/outputs/tickN-<role>-prompt.md, and returns the temp path. Falls back to raw prompt_ref when file not found (backward compat). {artifact_content} reads the file at extras["artifact"] and substitutes its content (empty string if missing).
  • _invoke_harness extended: new optional loop_path and tick_num params. When loop_path is provided, calls _resolve_prompt before building the harness command. All three call sites in cmd_tick (implement, verify, orchestrate) now pass loop_path and tick_num.
  • New prompt files: prompts/loop-implement.md (Implement role with task_brief/criteria/hint tokens, ALLOWED/FORBIDDEN sections), prompts/loop-verifier.md (Verify role with strict JSON output, score rubric 0.0-1.0, artifact_content token), prompts/loop-orchestrate.md (Orchestrate role with verdict token, phase transition logic, no-edit/no-auto-approve rules).
  • templates/loops/ci-triage/loop.json: roles filled from null to {"prompt": "loop-implement.md"} etc.
  • templates/loops/self-improvement/loop.json: new template with work_source: audit, blast_radius.use_worktree: true, file_scope: ["scripts/", "prompts/", "tests/", "design/"], brakes.max_iterations: 10, score_plateau_window: 3.
  • README.md: new "Loop Engineering" onboarding section (quick start, tick cycle, configuration, monitoring, halt/resume).
  • design/loops/technical.md section 8: documented prompt resolution and token substitution flow.
  • New tests: tests/test_loop_templates.py -- 18 tests covering R1-R6 (resolve_prompt, prompt file content, ci-triage template, self-improvement template, tick integration). Full suite: 393 passed.
  • Test infrastructure: updated tests/test_loop_runner.py, tests/test_blast_radius.py, tests/test_goal_mode.py to use non-existent prompt refs (test-impl.md etc.) so _resolve_prompt fallback path is exercised in those tests. Added loop prompts to self-consistency test exclusion set.

Added — loop engineering v1, blast-radius scheduler (task add-blast-radius-scheduler)

  • scripts/loop-runner.py extended: _ensure_worktree(state, cfg, loop_path, project_dir) creates a per-loop git worktree at <loop>/worktree on branch loop/<name> when blast_radius.use_worktree is true (default) and no worktree exists yet. Records worktree_path and worktree_branch in .state.loop atomically. Reuses existing worktree on subsequent ticks. Falls back to project root with a WARNING log when: not a git repo, git binary missing, or git worktree add fails. Handles branch-already-exists by retrying without -b. Stale worktree_path (directory deleted) is cleared and worktree recreated.
  • New tests: tests/test_blast_radius.py -- 15 tests covering worktree creation, reuse, fallback, branch-exists retry, state consistency, tick integration, platform paths, and backward compat. Full suite: 369 passed (was 354 + 15 new).

Added — loop engineering v1, goal-mode (task add-goal-mode)

  • scripts/loop-runner.py extended: _find_work(state, cfg, loop_path, project_dir) replaces the inline single-only block, dispatching on loop.json work_source.kind:
    • "single": unchanged behavior (uses .state.loop current_task).
    • "audit": runs status.py --audit --json --project <p>, picks the highest-severity unresolved violation. If the violation has a task, uses it; otherwise slugifies the message and calls --create-task, then sets the new task as current_task. When no unresolved violations → SKIP no_work (clean scheduler exit; does not increment iteration_count). Optional work_source.project overrides the audit target.
    • "backlog": reads <root>/design/<area>/BACKLOG.md (work_source.area defaults to "loops"); picks the topmost - [ ] item. Empty backlog → SKIP no_work.
  • Goal-oriented substitution tokens: three new tokens available in harness.command:
    • {task_brief} -- from <task>/RESEARCH.md or DESIGN.md or SPEC.md (first hit), capped at 4k tokens.
    • {acceptance_criteria} -- from loop.json acceptance_criteria (string OR list joined by newlines), capped at 2k tokens.
    • {next_hint} -- from state.last_verdict.next_hint (empty on first tick / after --approve), capped at 1k tokens.
    • Closes the next_hint feedback loop: tick N's verifier hint becomes tick N+1's Implement/Verify context (D7 graded verifier feedback).
  • _truncate_tokens(text, max_tokens) stdlib-only approximate cap (4-chars-per-token heuristic, …[truncated] marker). No tokenizer dependency.
  • scripts/status.py --audit --json: machine-readable audit mode. Emits a single JSON line on stdout: {"violations": [...], "loops": [...], "total_tasks": N, "untracked_tasks": M}. Each violation carries category, severity (high/med/low), task, message, resolved: false. Existing human-readable --audit output is unchanged when --json is absent. Backed by new _audit_collect and _audit_category3_paths helpers.
  • templates/loops/ci-triage/loop.json: added "work_source": {"kind": "single"} and an "acceptance_criteria" example so the template is self-documenting.
  • Backward compat: missing work_source or unknown kind falls back to "single" with a .state.log WARNING entry. Existing task-3 fixtures tick identically.
  • New tests: tests/test_goal_mode.py -- 26 tests covering R1-R8 and one regression. All subprocess calls stubbed via monkeypatch; no live LLM in CI. Full suite: 354 passed (was 328 + 26 new).
  • Doc updates: design/loops/technical.md §9 self-improvement template now shows acceptance_criteria; design/loops/functional.md §9 loop.json fields list documents work_source and acceptance_criteria.

Added — loop engineering v1, loop runner (task add-loop-runner)

  • scripts/loop-runner.py: per-tick engine. --mode tick runs one tick (gate -> find work -> spawn Implement -> spawn Verify -> parse graded JSON verdict -> spawn Orchestrate -> atomic state write -> log); --mode daemon runs a sleep loop bounded by --max-iterations. Calls status.py --check-gate first; any non-ok gate exits 0 (clean scheduler exit). Calls vram_detect.py --loop-mode --json before any harness subprocess; refuses below the 16k floor (D13) with human_intervention. Idempotent in failure -- pre-step-10 crashes do not corrupt .state.loop or advance iteration_count.
  • Graded verifier protocol: parse_verdict accepts raw JSON, ```json fenced blocks, or JSON with ///# line comments. Required keys: pass (bool), score (float). Optional: reasons (list[str]), next_hint (str). Parse failure halts as verifier_failed without advancing state.
  • Harness command substitution: loop.json harness.command (list of strings) with tokens {prompt}, {cwd}, {output}, {artifact}, {verdict}, {current_task}, {current_phase}. Default command (v1.1, see fix-harness-command-template above): ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"] (old default with --prompt-file/--cwd was non-functional; fixed in v1.1). Stdlib only, no new pip deps.
  • Score history capping: per-tick score appended to score_history, capped at brakes.score_plateau_window (oldest dropped). Score plateau halts via the next tick's --check-gate.
  • Tick log: every state-changing op appends an ISO-timestamped TICK pass=<bool> score=<f> iter=<N> line. Daemon mode appends DAEMON_STOPPED on KeyboardInterrupt.
  • New tests: tests/test_loop_runner.py -- 18 tests across 7 classes cover R1--R8. All subprocess calls stubbed via monkeypatch (no live LLM in CI). Full suite: 328 passed (was 310 + 18 new).
  • Doc updates: AGENTS.md "Loop runner" bullet under State Enforcement -- Loops (v1); README.md loop-runner one-liner in the Loop Engineering (beta) section.

Added — loop engineering v1, brakes layer (task add-status-brakes)

  • .state.loop runtime state: single source of truth per loop at {project}/.automaton/loops/<name>/.state.loop (framework-internal at ~/.automaton/loops/<name>/). Schema v1 with 13 fields (status, halt_reason, iteration_count, resumed_count, last_tick_at, last_verdict, score_history, current_task, worktree_branch, worktree_path). Atomic tmp+rename writes.
  • --create-loop NAME [--from-template T]: only way to bootstrap a loop dir + .state.loop + loop.json. Refuses non-kebab names, duplicates, unknown templates; patches name into the copied loop.json.
  • --check-gate NAME [--json]: runs 6 brake gates in order (status, iterations, budget, task phase, worktree drift, score plateau). First failure halts the loop, best-effort disables the OS schedule unit, emits structured JSON verdict.
  • --can-continue NAME [--json]: cheap pre-tick probe — ok := status == "running".
  • --approve --loop NAME: the only way to clear a halt. Increments resumed_count. No auto-approve in v1 (D4).
  • --pause-loop / --resume-loop: user-controlled soft stop; cannot clear halts.
  • --install-schedule NAME [--interval S]: triple-dispatched (Darwin launchd plist / Linux crontab block / Windows schtasks) native schedule unit generator per platform.system(). Generates automaton-loop-tick.sh / automaton-loop-tick.bat stub (self-documenting: when the scheduler fires it, the filename alone says what it does).
  • --can-edit --loop NAME [--loop-worktree] --file P: worktree scope check — is the file inside this loop's declared blast_radius.file_scope?
  • --transition halt refusal: refuses to transition any task owned by a HALTED loop until --approve --loop clears the halt.
  • --audit Cat-6 Loops block + --loop-list: flags halted / untracked / stale-running loops and loops whose current_task no longer exists. _audit_loops_block runs even when no tasks are present.
  • --version: prints framework version parsed from config.md's ## Framework Version section.
  • .state.log tick trail: every state-changing loop op appends an ISO-timestamped line (PAUSED / RESUMED / APPROVED / HALT).
  • New loop template: templates/loops/ci-triage/loop.json (minimal; full template expansion lands in task add-loop-templates-onboarding).
  • New tests: tests/test_status_brakes.py — 46 tests across 10 classes cover R1–R10. Full suite: 310 passed (was 264 + 46 new).
  • Doc updates: AGENTS.md Harness Integration modes block extended with the --loop worktree-scope mode; new "State Enforcement — Loops (v1)" section in AGENTS.md.

Added — loop engineering v1 (design only; implementation pending bootstrap tasks)

  • Loop system design: design/loops/{functional,technical,README,BACKLOG}.md — locked v1 design for an unattended, state-enforced loop runner built on top of the existing status.py phase machine. No second enforcement surface.
  • Five deaths halt model: iterations_exhausted, budget_exhausted, verifier_failed, drift_detected, human_intervention. All halts require human --approve --loop to resume; no auto-approve path (D4).
  • Tiered role budgets: three session roles (Implement:, Verify:, Orchestrate:) with explicit context tiers. 16k floor hard refuse (D13). Session divergence mandatory; model divergence only when Verify: != Implement: (D12).
  • Per-loop git worktree as blast radius default (D2); --no-worktree opt-out.
  • Native scheduler generator: platform.system()-dispatched launchd / crontab / schtasks unit generation via status.py --install-schedule (D1); --daemon opt-in fallback.
  • Self-improvement loop template at templates/loops/self-improvement/, default-on at install (D21). Ticks against status.py --audit on the framework's own repo — the literal seed of self-management.
  • Bootstrap task plan (8 tasks) committed to design/loops/README.md — fix-context-sizing, add-status-brakes, add-loop-runner, add-goal-mode, add-blast-radius-scheduler, add-loop-templates-onboarding, add-self-improvement-loop, fix-install-update-flow. These are the last tasks a human creates by hand; after task 7 lands, loops create subsequent tasks from BACKLOG.md and --audit output (D24, D25).
  • Scoped deferrals recorded in BACKLOG.md: Scope 2 (design-update loop) → v1.1; Scope 3 (self-designing loops) + parallel-mode-default + auto-approve-relax → deferred indefinitely (D20).

Added

  • Added: requirements.txt pinning pytest==7.4.4 for reproducible test runs.
  • Added: scripts/install.sh now creates .venv/ and installs pytest into it.
  • Harness pre-edit hook: --can-edit now supports project-level checks without --task, file scope checks with --file, and --json output for machine-readable harness integration
  • opencode plugin: plugins/automaton-guard/plugin.ts — intercepts edit and write tool calls, calls --can-edit before allowing modifications
  • Git pre-commit hook: scripts/git-hooks/pre-commit — blocks commits when no task is in an edit-allowed phase (universal safety net for all harnesses)
  • Pre-v2.0 task enforcement: Tasks without .state files are UNTRACKED — --transition, --can-edit, --task, and --approve all refuse to operate on them
  • New --upgrade command: Bootstraps .state files for pre-v2.0 tasks (single task with --task or all tasks at once)
  • Untracked task reporting: --list shows UNTRACKED (no .state) for tasks without .state files instead of silently bootstrapping
  • Project scoping fix: status.py errors when no project is detected instead of silently falling back to framework directory
  • Scope check fix: --scope-check marks framework files as OUT_OF_SCOPE when working on a project
  • Dashboard scope fix: Handler methods use stored project_root instead of re-detecting from CWD on every request
  • --project flag: Added to all status.py command invocations across 16+ prompt and config files
  • _infer_state_from_artifacts locked to --upgrade: Removed as silent fallback from all operational commands
  • Phase approval gates: Research, Decomposition, Design, and Test Design phases now require explicit user approval (:awaiting_approval → :approved) before proceeding
  • status.py script: Comprehensive enforcement and status tool with --task, --list, --create-task, --transition, --approve, --validate-folder, --audit, --claim, --release, --next-available, --available, --can-edit, --scope-check, --same-session, --upgrade
  • Untracked task enforcement: Tasks without .state files are UNTRACKED — --transition, --can-edit, --task, --approve all refuse to operate on them. Run --upgrade to bootstrap .state files
  • --project flag: All status.py commands now support --project for explicit project scoping when multiple projects exist on the same machine
  • --upgrade command: Bootstraps .state files for pre-v2.0 tasks that lack them (single task with --task or all tasks at once)
  • Project scoping: status.py now errors when not in a project directory and --project is not specified, instead of silently falling back to ~/.automaton/
  • Scope check fix: --scope-check now correctly marks framework files as OUT_OF_SCOPE when working on a project (was incorrectly always IN_SCOPE)
  • Dashboard scope fix: Dashboard handler methods now use stored project_root and scope instead of re-detecting from CWD on every request
  • Phase-scoped prompts: All phase prompts now include ALLOWED ACTIONS, FORBIDDEN ACTIONS, approval gates (where applicable), pre-work validation, and .state precondition checks
  • Orchestrator restructuring: Reduced from 493 lines to 143 lines; sub-task management extracted to subtask_management.md; state machine reference moved to workflow.md
  • ALLOWED/FORBIDDEN enforcement: Each phase prompt explicitly defines what agents can and cannot do, with user override resistance instructions
  • Workflow enforcement: --transition refuses illegal phase transitions; --validate-folder detects out-of-order artifacts; --audit checks all tasks for violations
  • Task creation gate: status.py --create-task is the only valid way to create tasks; --audit flags manually created folders
  • Approval log: .state.approvals file records all user approvals with timestamp and approver
  • Multi-agent support: Optional Agent Configuration section in .agent.md enables task claiming, role binding, and work discovery for multi-agent setups
  • Tool integration hooks: --can-edit, --scope-check, --same-session for agent tool integrations (optional, not called by prompts)
  • upgrade.sh script: Bootstraps .state files for existing tasks from artifact heuristic
  • Framework version marker: config.md now includes version 2.0 with state enforcement indicator

Changed

  • Changed: All documented python invocations now read python3 (stock macOS / Windows Python ship as python3).
  • orchestrate.md: Reduced from 493 to 143 lines; gate-check loop replaces soft advisory approach; approval gates enforced at research, decomposition, design, and test_design
  • workflow.md: Rewritten to reference .state as canonical phase indicator; approval sub-states documented; enforcement via status.py documented
  • All phase prompts: Added .state precondition check, pre-work validation, ALLOWED/FORBIDDEN sections, handling user overrides
  • research.md, design.md, decompose.md, test_design.md: Added approval gate sections with --transition {phase}:awaiting_approval and --approve
  • implement.md, bug_finder.md, adversarial_bug_find.md, doc_review.md, referee.md: Added no-approval-gate notes with direct --transition instructions
  • status_reason property on Task model showing human-readable explanation for each state (#task-status-reason)
  • Revoke buttons for approved/changes_requested reviews — replaces approve/request-changes with a single revoke option (#task-status-reason)
  • pytest test suite covering dashboard core, app security, and VRAM detection (#add-pytest-test-suite)
  • Structured verdict parsing: parse_verdict_status() uses ## Status: line before substring fallback, preventing false-BLOCKED classification (#fix-verdict-parsing)
  • State machine alignment: IMPLEMENTATION.md alone → Bug Find, ADVERSARIAL_BUG_REPORT alone → Bug Find (matching orchestrator spec) (#fix-verdict-parsing)
  • Filesystem task name validation: discover_tasks() and parse_sub_tasks() skip directories with invalid characters (#fix-verdict-parsing)
  • Added CORS headers, do_OPTIONS handler, X-Content-Type-Options to all dashboard API responses (#harden-dashboard-security)
  • Added POST content-length bounds (64KB) and review comment length limits (4096 chars) (#harden-dashboard-security)
  • Replaced inline onclick review handlers with data-* attributes and event delegation (#harden-dashboard-security)
  • Applied escapeHtml() to task display_name in dashboard card rendering (#harden-dashboard-security)
  • GET /api/config and PUT /api/config endpoints for reading and persisting dashboard configuration (#wire-dashboard-config)
  • Server-side task cache with 1s TTL to eliminate redundant disk I/O on every polling request (#wire-dashboard-config)
  • Dashboard JS applies config on init: theme, default_view, auto_refresh_interval, column_width, show_timelines (#wire-dashboard-config)
  • Review POSTinvalidates task cache so next poll picks up changes (#wire-dashboard-config)
  • decomposition_content, parent_spec_content, vram_config_content fields on Task model (#add-decomposition-content)
  • WaveGroup dataclass and parse_waves() for extracting wave structure from DECOMPOSITION.md (#add-decomposition-content)
  • parse_vram_config() for reading VRAM_CONFIG.md (#add-decomposition-content)
  • Dashboard JS wave statistics use parsed wave data instead of 50/50 heuristic (#add-decomposition-content)
  • Detail panel shows Decomposition, Parent Context, and VRAM Configuration sections (#add-decomposition-content)

Changed

  • Removed stale dashboard = ["inotify>=0.2"] optional dependency from pyproject.toml (#cleanup-cruft)
  • Deleted debug_root.py stray development script (#cleanup-cruft)
  • Deleted empty automaton/dashboard/ui/widgets/ directory (#cleanup-cruft)
  • Fixed config.md RAM detection description (was "via free", now "via /proc/meminfo or sysctl") (#cleanup-cruft)
  • _find_tasks_dir() returns Path instead of Path | None, removed tautological condition (#cleanup-cruft)
  • Removed sys.path.insert hack from __main__.py (#cleanup-cruft)
  • Documented scripts/dashboard.sh convenience wrapper in README.md (#cleanup-cruft)
  • Framework self-consistency test suite: 17 tests covering prompt stop conditions, hardcoded URLs, canonical paths, .rules.md sections, stale dependencies, CSS theme parity, verdict regression, and CI validation (#framework-self-consistency-tests)

Fixed

  • REFEREE state was never produced by state machine — verdict with unparseable status now correctly shows as REFEREE instead of silently falling through to earlier states (#task-status-reason)
  • Pending review count in header now excludes done/blocked tasks (#task-status-reason)
  • Critical: PASS verdicts mentioning FAIL/NEEDS_REVIEW in body text were falsely classified as BLOCKED (#fix-verdict-parsing)
  • State divergence: IMPLEMENTATION.md alone showed "Implement" instead of "Bug Find" (#fix-verdict-parsing)
  • Added mandatory stop conditions to bug_finder.md and adversarial_bug_find.md (#fix-prompt-consistency)
  • Fixed deprecated {project}/tasks/ path in onboarding.md (#fix-prompt-consistency)
  • Expanded prompt path test to catch concrete deprecated path patterns (#fix-prompt-consistency)
  • Root pyproject.toml with optional test/dashboard dependency groups (#add-pytest-test-suite)
  • AGENTS.md with build/test commands and conventions (#developer-experience-gitea-ci)
  • .gitea/workflows/ci.yml running py_compile, pytest, and shell script syntax checks (#developer-experience-gitea-ci)
  • templates/README.md documenting the task template examples (#developer-experience-gitea-ci)
  • Blocked phase column between Verification and Resolution on dashboard (#additive-extension-model)
  • Framework self-enforcement rules in .rules.md and system-prompt.md (#framework-self-enforcement)
  • Additive extension model: projects extend via extensions/ dir, never copy framework files (#additive-extension-model)
  • CHANGELOG.md for release notes tracking (#changelog)
  • Framework audit: comprehensive self-consistency check with RESEARCH.md (#framework-audit)
  • Audit Bug 1: --audit category 3 now checks .automaton/tasks/ paths (was only checking tasks/) (#fix-cat3-audit-paths)
  • Audit Bug 2: migrate-project.sh find command now has parentheses around -name group for correct -prune binding (#fix-migrate-find-precedence)
  • Audit Bug 3: _lookup_model_context() no longer false-matches model prefixes (e.g. phi-4 matching phi-4-mini) — uses three-tier matching with known suffix whitelist (#fix-vram-model-prefix-match)
  • Audit Bug 4: Verdict PASS/FAIL inference uses structured ## Status: line parsing instead of fragile substring search (#fix-verdict-pass-inference)
  • Audit Bug 5: register-guards.sh now checks both .json/.jsonc, writes to plugin (singular) key, and strips // comments before json.loads() (#fix-register-guards)
  • Audit Bug 6: Dashboard determine_task_state() now reads .state file (source of truth) before falling back to artifact heuristic (#fix-dashboard-read-state)
  • Audit Bug 7: --can-edit and --scope-check path prefix matching uses os.sep boundary to prevent sibling directory false matches (#fix-can-edit-path-prefix)
  • Audit Bug 8: Removed wildcard CORS Access-Control-Allow-Origin: * from dashboard — replaced with security headers (X-Content-Type-Options, X-Frame-Options) (#fix-dashboard-cors-origin)
  • Audit Bug 9: Stale-task detection uses .state.lastedit timestamp (touched on actual edit activity) instead of .state mtime (which only reflects phase transitions) (#fix-stale-task-mtime-proxy)
  • Audit Bug 10: TEST_PLAN.md now correctly maps to test_design phase (was mapping to implement) in both status.py and dashboard task.py (#fix-test-plan-phase-mapping)

Changed

  • All prompts now use the canonical task path {project}/.automaton/tasks/{task-name}/ (#standardize-task-path-conventions)
  • scripts/vram_detect.sh rewritten as scripts/vram_detect.py for testability and correctness (#rewrite-vram-detection-python)
  • tasks/dashboard-spec.md reconciled with the implemented web dashboard (#reconcile-dashboard-spec)
  • automaton/dashboard/README.md and help modal shortcuts now match the web UI (#reconcile-dashboard-spec)
  • prompts/orchestrate.md: always reads prompts/contracts/scripts from global, project extensions are additive (#additive-extension-model)
  • prompts/onboarding.md: removed diff/merge upgrade, replaced with migration check (#additive-extension-model)
  • README.md: updated upgrade docs for new additive model (#additive-extension-model); added Dashboard section (#dashboard-task-review)
  • scripts/update.sh: simplified to plain git pull (#additive-extension-model)
  • .rules.md: converted from template to concrete rules with Task-Driven Development, VRAM-aware sizing, Changelog, and Self-Improvement sections (#framework-self-enforcement)
  • system-prompt.md: added instruction to read global .rules.md (#framework-self-enforcement)
  • automaton/dashboard/ui/app.py: added review API endpoints (GET/POST /api/task/{name}/review), spec_content in responses, unquote() for URL-encoded task names, path traversal fix (#dashboard-task-review, #spec-in-detail)
  • automaton/dashboard/html/dashboard.js: review UI (badges, buttons, filter), artifact badges, specification display, modal conversion, textarea replacement, display group for approved planning tasks (#dashboard-task-review, #artifact-badges, #spec-in-detail, #task-detail-modal, #review-textarea)
  • automaton/dashboard/html/styles.css: review components, artifact badges, modal layout, textarea styles (#dashboard-task-review, #artifact-badges, #task-detail-modal, #review-textarea)
  • automaton/dashboard/html/index.html: review filter, pending count, modal overlay (#dashboard-task-review, #task-detail-modal)
  • automaton/dashboard/core/task.py: fixed state machine priority — IMPLEMENTATION.md now correctly detected, DOC_REVIEW checked before BUG_REPORT (#implement-task)
  • automaton/dashboard/core/board.py: fixed KanbanBoard — added missing COLUMNS and init (#implement-task)
  • automaton/dashboard/core/refresh.py: improved inotify error handling with explicit fallback messages (#implement-task)

Fixed

  • VRAM detection: undefined headroom, hardcoded JSON headroom, and code-block config parsing (#rewrite-vram-detection-python)
  • VRAM detection: 10KB file-read limit now enforced for API config files (#rewrite-vram-detection-python)
  • Dashboard static file serving: replaced string-prefix path traversal check with Path.relative_to() (#harden-dashboard-security-scripts)
  • Dashboard task name validation: restricted to [A-Za-z0-9_-]+ (#harden-dashboard-security-scripts)
  • scripts/update.sh: now warns and aborts on uncommitted changes before pulling (#harden-dashboard-security-scripts)
  • README/install.sh: replaced placeholder repository URL with real Gitea URL (#harden-dashboard-security-scripts)
  • State machine: IMPLEMENTATION.md was never checked in determine_task_state(), tasks showed as RESEARCH (#implement-task)
  • State machine: DOC_REVIEW checked after BUG_REPORT — wrong priority order (#implement-task)
  • Path traversal: review API accepted task names with ../ allowing writes outside tasks directory (#dashboard-task-review)
  • URL encoding: task names with spaces in API paths were not decoded (#implement-task)
  • Review parsing: comment extraction used fragile conditional, falsy comments (e.g., "0") skipped (#implement-task)
  • Board display: approved planning tasks stayed in Planning column instead of advancing to Design (#dashboard-task-review)

Removed

  • automaton/dashboard/themes.py (vestigial ANSI theme stub) (#reconcile-dashboard-spec)
  • automaton/dashboard/core/refresh.py (half-implemented file watcher; dashboard uses JS polling) (#remove-file-system-watcher)
  • templates/contract-template.md (unused) (#developer-experience-gitea-ci)
  • automaton/dashboard/pyproject.toml (consolidated into root pyproject.toml) (#add-pytest-test-suite)

Migration

  • Project migration script for old-model projects: scripts/migrate-project.sh (#project-migration)
  • Project migration detection in onboarding.md (#project-migration)