Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

343 lines
52 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Changelog
## [unreleased]
### Fixed — dashboard scroll-reset on auto-refresh
- **`automaton/dashboard/html/dashboard.js`** (`renderBoard`): auto-refresh rebuilt the board via `board.innerHTML = html` every tick (default 2s), destroying each `.column-body`'s `scrollTop` and snapping it back to 0 — so users couldn't scroll the Done group down to review older tasks. Now snapshots each column-body's `scrollTop` (plus the board's `scrollLeft` and the active view's `scrollTop`) before the rebuild and restores them after, matched by index (PHASE_GROUPS order is stable).
- **New tests**: `tests/test_dashboard_ui.py` — Playwright browser smoke test (board renders tasks; column scroll survives an auto-refresh tick). Skipped via `importorskip` when playwright/chromium is absent so CI without a browser stays green. Verified the test fails without the fix (scrollTop resets to 0) and passes with it.
### Fixed — failing plist-isolation test (host bleed false positive)
- **`tests/test_cleanup_done.py`** (`TestInstallCleanupScheduleIsolation.test_plist_written_to_override_dir_not_host`): asserted `not host.exists()`, but the host `~/Library/LaunchAgents/com.automaton.cleanup.plist` legitimately exists from a real `--install-cleanup-schedule` run, causing a false failure. Now snapshots the host plist's `st_mtime_ns` (or absence) before the test run and asserts it's unchanged after — a pre-existing real install no longer fails the test; only an actual write during the run would.
### Changed — bind ornith as the Implement model
- **`config.md`** (Model Configuration): set `Model: omlx/Ornith-1.0-35B-4bit-mlx` and `Override context window: 32768` (matches the opencode.json limit for the local LLM). Interactive autopilot already used ornith via opencode's default model; this makes it explicit so auto-detection can't pick another model. Loop ticks still use the single `harness.command` for all roles — per-role model binding (`{model}` substitution in loop-runner.py) is **not** implemented yet (see model-divergence gap below).
### Changed — README: document the self-improvement loop's scope for new projects
- **`README.md`** (Project Setup): added "The Self-Improvement Loop is framework-scoped" note — the default SI loop targets `~/.automaton/` (the framework), not your project, by design. Documents the leave-running / pause / create-a-project-loop paths.
### Known gap — model-divergence was marked complete but unimplemented
- The `model-divergence-enforcement` parent task and its 3 subtasks (`mde-manifest-detection`, `mde-interactive-enforcement`, `mde-loop-enforcement`) are `.state = complete` but contain only `SPEC.md`/`DECOMPOSITION.md` — no `IMPLEMENTATION.md`, no `VERDICT.md`. The promised code (`status.py` model_divergence audit category, `--transition --model`, `.state.models`, `loop.json` per-role `model` + `{model}` substitution in loop-runner.py, dashboard badges) was never written. Consequence: loop roles (implement/verify/orchestrate) all run the same model, so the D12 conflict-of-interest rule (Verify ≠ Implement session/model) is unenforced. Interactive autopilot is unaffected.
### Added — framework agent features design docs
- **New `design/framework/`** directory: design index, functional design, technical design, and backlog for three framework-level agent features:
- **Model-Divergence Enforcement** — `models.json` manifest, single vs multi-LLM mode detection, conflict matrix (`code_review≠implement`, `bug_find≠implement`, `adversarial_bug_find≠implement+bug_find`, `referee≠implement+bug_find+adversarial_bug_find`, `loop-verify≠loop-implement`), `--transition --model`, `.state.models`, `loop.json` per-role model binding + `{model}` substitution, `--audit` model_divergence category, dashboard badges. Shipped as a manual task (3 subtasks), not a backlog item.
- **Rule Agents** — Rule Proposer (daily scheduled standalone agent, scans completed tasks' `BUG_REPORT.md`/`ADVERSARIAL_BUG_REPORT.md`/`VERDICT.md`, proposes rules to `RULE_PROPOSALS.md`) and Rule Reviewer (monthly scheduled standalone agent, consolidates `.rules.md` → `RULE_REVIEW.md`). Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers. Conflict-of-interest: Reviewer must differ from Proposer; both must differ from tasks they review. Enforcement deferred until model-divergence ships.
- **Agent Tab Redesign** — replace 4 fake `AGENT_TYPE_META` types with two sections: Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) + Scheduled Jobs (real `job.kind`: cleanup, loop, rule-scan, rule-review). New `/api/phase-roles` endpoint.
- **`design/framework/BACKLOG.md`**: 3 v1 items (`agent-tab-real-roles`, `rule-proposer-agent`, `rule-reviewer-agent`) + 3 deferred items. Consumable by the self-improvement loop via `work_source.area = "framework"`.
- **Cross-references**: `design/loops/README.md` and `design/loops/BACKLOG.md` now point to the sibling `design/framework/` area. `AGENTS.md` repo layout updated.
### Added — cross-loop task claim (task `add-claim-loop-task`)
- **`scripts/status.py`**: New `--claim-loop-task <name> --task <taskname> [--project P]` command. Exit 0 = claimed (or already self-claimed, idempotent). Exit 2 = already claimed by another running/paused loop (`task_already_claimed:{other}` on stderr) or untracked loop (`loop_untracked`). Uses `_loop_lock` to serialize writes; cross-loop scan is advisory (self-healing on next tick).
- **`scripts/loop-runner.py`**: `cmd_tick` step 3.5: after `_find_work` returns a candidate different from `current_task`, spawn `status.py --claim-loop-task` as subprocess with `$AUTOMATON_NO_LOOP_LOCK=1` (same bypass pattern as `_gate`). Step 9.5: after orchestrator, re-read `.state`; if phase is `complete` or `human_intervention`, clear `current_task` to None. Closes `add-status-brakes/ADVERSARIAL_BUG_REPORT.md` A2 (cross-loop task race).
- **New tests**: `tests/test_claim_loop_task.py` — 10 tests covering claim success, refusal, idempotency, untracked loop, missing task, paused-loop ownership, self-healing race, and release on terminal phases.
- **Full suite**: **518 passed** (was 508; +10 new; 0 regressions).
### Added — Linux schedule parity (task `linux-schedule-parity`)
- **`scripts/status.py`**: New `_install_cron_block(name, project, interval_seconds)` for Linux cron support. Writes `# automaton-loop:<name>` / `# end automaton-loop:<name>` blocks into crontab via `crontab -`. Interval rounded to full minutes, minimum 1. Strips prior block before insert (idempotent).
- New `_enable_schedule(name, project)`: platform dispatch — Linux (cron insert via `_install_cron_block`), Darwin (rename `.plist.disabled` back), Windows (no-op). Reads `loop.json > schedule.interval_seconds`; non-int falls back to 3600.
- `_disable_schedule` already existed; verified correct.
- **New tests**: `tests/test_linux_schedule_parity.py` — 13 tests covering cron block writes, error handling, platform dispatch, and interval edge cases.
- **Full suite**: **508 passed** (was 495; +13 new; 0 regressions).
### Fixed — parametrize-base-branch (task `parametrize-base-branch`)
- **`scripts/status.py`**: Replaced hardcoded `"main"` in `_gate_worktree_drift`'s `git diff` call with `_base_branch(cfg) -> str`. New helper returns `blast_radius.base_branch` if configured (default `"main"`). Empty string and non-string types produce WARNING and fall back to `"main"`. Closes `add-status-brakes/BUG_REPORT.md` O3: projects on `master`/`trunk`/`develop` no longer have a silently-disabled drift gate — operator sets `"base_branch": "master"` in `loop.json`.
- **`templates/loops/self-improvement/loop.json`**: Added `"base_branch": "main"` to `blast_radius`.
- **`design/loops/functional.md` §9**: Updated `blast_radius` field list to include `base_branch`.
- **New tests**: `tests/test_base_branch.py` — 13 tests covering `_base_branch` helper (5) and drift-gate branch usage (8, with mocked `subprocess.run`).
- **Full suite**: **495 passed** (was 482; +13 new; 0 regressions).
### Added — outputs retention GC (task `add-outputs-retention`)
- **`scripts/loop-runner.py`**: Added `_get_retention(cfg)` and `_gc_outputs(loop_path, retention)` to bound `outputs/` directory growth. Every tick (`cmd_tick` step 10.5, inside `_loop_lock`), older tick groups are deleted, keeping only the last N (default 20). Closes `add-loop-runner/BUG_REPORT.md` O5 (tick dirs accumulate without bound).
- `_get_retention` reads `cfg.get("outputs", {}).get("retention", 20)`. Non-int types fall back to 20 with WARNING. Negative values coerce to 0 (unlimited) with WARNING. 0 = no GC (v1 behavior).
- `_gc_outputs` lists `outputs/`, parses `tickNN` indices via `^tick(\d+)-` regex, computes `cutoff = max_seen - retention + 1`, deletes files with tick index < cutoff. Non-tick files (e.g. `README.txt`, `tick-foo.md`) are preserved. Errors (permission, missing file) are logged as WARNING and swallowed — GC failure never crashes the tick.
- Off-by-one bug found and fixed inline: initial formula `cutoff = max_seen - retention` kept `retention+1` groups; caught by `test_gc_keeps_recent_deletes_old` length assertion.
- **`templates/loops/self-improvement/loop.json`**: Added `"outputs": {"retention": 20}` to the template schema.
- **New tests**: `tests/test_outputs_retention.py` — 13 tests across `TestGetRetention` (5) and `TestGcOutputs` (8). Pure-file-system; no subprocess, no live LLM.
- **Full suite**: **482 passed** (was 469; +13 new; 0 regressions).
### Fixed — `parse_verdict` defensive coercion (task `harden-parse-verdict`)
- **`scripts/loop-runner.py`**: Hardened `parse_verdict` against malformed verifier output. Closes `add-loop-runner/BUG_REPORT.md` O6 (string-typed `pass`) and the un-noted sibling issue (unclamped `score`).
- `pass` field: now accepts bool OR the strings `"true"`/`"false"` (case-insensitive, whitespace-stripped). The pre-fix `bool(data.get("pass"))` returned `True` for `"false"` (non-empty string is truthy) — a verifier emitting `{"pass": "false", "score": 0.1}` was recorded as `pass=True`, advancing the loop on a failed verdict. Non-`"true"`/`"false"` strings fall through to `bool(raw_pass.strip())` for backwards compat (`"yes"` stays truthy; `""` stays falsy; `None`/`0`/`[]`/`{}` keep v1's `bool(...)` semantics).
- `score` field: now clamped to `[0, 1]` via `max(0.0, min(1.0, score))`. `NaN`, `Infinity`, and `-Infinity` (which `json.loads` accepts as bare tokens because `float(...)` happily returns `math.nan`/`inf`) default to `0.5` (neutral midpoint) via a `math.isfinite` guard. Non-numeric types (`None`, `[]`, `"high"`, etc.) and non-numeric strings default to `0.5` via a `try/except (TypeError, ValueError)` around `float(...)`.
- Pure-function change; no CLI surface; no schema change; no new deps (`math` is stdlib). Strict emitters (those already emitting `bool pass` and `0.0 ≤ score ≤ 1.0`) are unaffected.
- **`design/loops/technical.md` §7**: Documented the new coercion contract inline in the tick-flow step 7 (parse verdict).
- **`design/loops/functional.md` §10**: Updated the Verifier Contract section to note the runner's defensive coercion for `pass` strings and `score` clamping.
- **New tests**: `tests/test_parse_verdict.py` — 22 tests across 5 classes (`TestStrictBaseline`, `TestPassStringCoercion`, `TestScoreClamping`, `TestFenceBlockStillWorks`, `TestOptionalKeysPreserved`). Pure-functional; no subprocess, no live LLM.
- **Adversarial probe**: 12-row sweep against `pass` as non-string non-bool types (`None`/`[]`/`{}`/`0`/`1`/`-1`/`1.5`/`[False]`/`[True]`), `score` as `[]`/`{}`/`[1, 2]`/`"high"`/numeric strings/JSON literal `NaN`/`Infinity`/`-Infinity`/`true`/`false`. All inputs yield deterministic, documented results; no crashes; no silent truthy/coercion regressions vs v1.
- **Inline bug found and fixed during implementation**: the first iteration of the three-way branch set `verdict_pass = raw_pass.strip().lower() == "true"` — which mapped EVERY non-`"true"` string to False, breaking the SPEC R1 `bool(...)` fallback clause (a `"yes"` string would have become False, silently regressing v1). Fixed to the explicit `if "true" / elif "false" / else: bool(...)` form; test `test_other_truthy_string_pass` was added immediately to lock the contract.
- **Full suite**: **469 passed** (was 447; +22 new; 0 regressions).
### Added — `.state.loop` file lock (task `add-state-loop-lock`)
- **`scripts/status.py`**: New `_loop_lock(loop_path, exclusive=True)` context manager wrapping the read-modify-write cycle of `.state.loop` with a cross-process file lock (POSIX `fcntl.flock(LOCK_EX)`, Windows `msvcrt.locking(LK_LOCK, 1)`) on `<loop_path>/.state.lock`. Closes the TOCTOU races flagged in `add-status-brakes/ADVERSARIAL_BUG_REPORT.md` (A6) and `add-loop-runner/ADVERSARIAL_BUG_REPORT.md` (A2, A7) where concurrent ticks (two scheduler firings on the same loop) or a concurrent `--pause-loop` / `--approve --loop` write could overwrite a tick's iteration increment or lose an approve's `resumed_count` increment. Per-loop granularity; blocking acquire; no timeout in v1.1 (operators notice a wedged tick via `--loop-list` stale `last_tick_at`).
- **`scripts/status.py` callsites wrapped**: `cmd_pause_loop`, `cmd_resume_loop`, `cmd_approve_loop`, `cmd_check_gate` now acquire `_loop_lock` around their read-modify-write blocks. `--create-loop` intentionally unwrapped (no prior state to race against; create is name-unique-refused). `_disable_schedule` / `_enable_schedule` side-effect toggles moved OUTSIDE the lock to keep the critical section tight; the per-tick `--check-gate` self-skip on non-`running` state makes the transient ~100ms window harmless.
- **`scripts/loop-runner.py`**: Mirrored `_loop_lock` helper (no env bypass — the runner is the lock holder). Wraps the entire `cmd_tick` critical section from the `_gate` subprocess call through step-10's state write. `_read_state_loop` is invoked twice: once unlocked (fast-fail untracked) and once inside the lock (re-read fresh to capture any concurrent mutate).
- **Subprocess-deadlock avoidance (D-L6)**: When the runner spawns `status.py --check-gate` as a subprocess inside its held lock, status.py's own `cmd_check_gate` would otherwise deadlock waiting on the parent's held flock. Resolved by introducing the `$AUTOMATON_NO_LOOP_LOCK=1` env var, scoped ONLY to the `--check-gate` subprocess's env (set via `subprocess.run`'s `env` kwarg in `_gate`). status.py's `_loop_lock` checks this env var; if set, it yields without flocking (trusting the caller's outer lock). Harness subprocesses (Implement/Verify/Orchestrate) do NOT inherit the env var, so any `status.py --transition` calls the harness transitively invokes lock normally and serialize correctly.
- **Scope of the env var**: `$AUTOMATON_NO_LOOP_LOCK` is read-only-internal; the runner never sets it in `os.environ` globally, only in the `_gate` subprocess's explicit env dict. Manual CLI users who set it in their shell bypass locking (documented escape hatch; same trust boundary as "shell user can kill the runner").
- **`design/loops/technical.md` §7**: Added a "Lock serialization" subsection documenting the lock shape, the env-bypass mechanism, the re-entry forbidding contract, and the harness-contract implication (loop-control commands deadlethal inside a tick; task commands are safe).
- **`AGENTS.md`**: Updated "State Enforcement — Loops (v1)" section to mention `_loop_lock` and the env var.
- **`README.md`**: Added a row in the loop engineering monitoring table mentioning `.state.lock` per-loop serialization.
- **Stdlib-only**: `fcntl` (POSIX) and `msvcrt` (Windows) are stdlib. No `filelock` package. No new pip deps.
- **Backwards-compatible**: existing scripts that don't invoke `--pause-loop` / `--resume-loop` / `--approve --loop` / `--check-gate` concurrently are unaffected. Primary lock surface is per-tick blocking when schedulers fire on the same loop concurrently.
- **New tests**: `tests/test_state_loop_lock.py` -- 7 tests covering (1) concurrent-acquire serialization, (2) clean-release + re-acquire, (3) release-on-exception, (4) per-loop granularity, (5) `--create-loop` doesn't create `.state.lock` (D-L3) but `--check-gate` does, (6) `--pause-loop` blocks under concurrent lock holder, (7) `cmd_tick` holds the lock across the harness subprocess while `--approve --loop` blocks. All use `tmp_path`, stdlib only, no live LLM.
- **Adversarial findings**: A1 (concurrent approve: race closed, only 1 of 5 won, `resumed_count` incremented once), A2 (concurrent pause: idempotent but consistent), A3 (lock releases on mid-tick exception: verified), A4 (harness calls loop-control inside tick: theoretical deadlock, LOW severity, deferred to harness-integration contract docs follow-up), A5 (manual env var escape hatch: documented), A6 (orphaned `.state.lock` after crash: benign; POSIX auto-releases flock on process exit), A7 (NFS: documented assumption), A8 (cyclomatic complexity: acceptable). No blockers.
- **Full suite**: **447 passed** (was 440; +7 new; 0 regressions).
### Fixed — harness command template (task `fix-harness-command-template`)
- **`scripts/loop-runner.py`**: Fixed the broken default `harness.command`. The old default `["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]` used flags that do not exist in `opencode run` (`--prompt-file`, `--cwd`) — the v1 runner has only been exercised via mocked subprocess tests, so the bug was never caught. New default is `["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]`. Introduces a new `{prompt_content}` substitution token that carries the resolved prompt's text as a single argv element under `subprocess.run` list mode (no shell expansion, safe for prompts containing quotes/special chars).
- **`{prompt}` and `{cwd}` tokens retained**: backwards-compatible. Users with custom `harness.command` in `loop.json` using these tokens are unaffected.
- **`UnicodeDecodeError` now caught** alongside `OSError` when reading the resolved prompt file (defensive; falls back to empty prompt content rather than crashing mid-tick).
- **Harness-agnostic contract preserved and extended**: runner core has zero harness awareness; the new `{prompt_content}` token covers harnesses that prefer a message argument (Pi Dev, aider, any CLI taking a prompt as a positional). Non-opencode users override `harness.command` in `loop.json` (e.g. `["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]`). D8 (no model/provider inspection) intact.
- **`design/loops/technical.md` §8 and §9**: Updated default command shape; documented the new `{prompt_content}` token alongside existing `{prompt}`/`{cwd}`/`{output}`/`{artifact}` tokens; added Pi Dev, aider, and generic shell-wrapper examples in `loop.json` form.
- **`templates/loops/self-improvement/loop.json`**: Updated `harness.command` to the new default.
- **Test scaffolding updates**: `_make_loop` helpers in `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py`, `tests/test_loop_templates.py` now write loop-local prompt stubs (containing role-marker content `prompt: <ref>`) when the framework prompt at `~/.automaton/prompts/<ref>` does not already exist. This preserves the substring-matcher strategy used by tick-flow tests while not overriding real framework prompts (which contain `{current_task}` etc. substitution tokens). Custom-`{prompt}`-command tests in `test_goal_mode.py` updated their matcher substrings from `test-impl`/`test-verify`/`test-orch` to `implement-prompt`/`verify-prompt`/`orchestrate-prompt` (the resolved temp-file path contains those markers).
- **New tests**: `tests/test_harness_command.py` -- 7 tests covering: default uses `--dir` not `--cwd`/`--prompt-file`; default passes prompt content as last argv element; prompt content preserves single quotes/double quotes/dollar signs as a single argv element; `{prompt}` token still available for custom commands; `{cwd}` token still works in custom commands; empty `command` falls back to the new default; Pi Dev-shaped command substitution (proves the substitution mechanism is harness-agnostic — any binary works).
- **Backwards-compat correction**: Updated the v1 `add-loop-runner` CHANGELOG bullet that documented the old default — no retroactive edit to the released entry shape, but the new `### Fixed` entry supersedes the default-command wording. The default command works end-to-end now.
- Full suite: **440 passed** (was 433 baseline; +7 new).
### Changed — completed tasks moved to `tasks/complete/` (task `move-completed-tasks-to-complete-folder`)
- **`scripts/status.py`**: `--transition complete` now moves the task directory from `tasks/<name>/` to `tasks/complete/<name>/`. `_task_dir` has a fallback to find completed tasks. `--list` and `--audit` exclude completed tasks (visible only via `--task <name>` fallback).
- **New tests**: `tests/test_move_completed.py` -- 9 tests covering directory move, fallback, listing exclusion, create-task refusal, and transition refusal from complete. Full suite: 433 passed.
### Fixed — install/update flow (task `fix-install-update-flow`)
- **`scripts/install.sh`**: Replaced hardcoded private git URL with user-supplied `GIT_URL="${1:-}"` argument (D11). Refuses with usage and irreversibility warning if absent. Fixed `.venv` cwd bug: venv now created in `$FRAMEWORK_DIR/.venv` instead of CWD. Added Windows venv path support (`.venv/Scripts/python.exe`). Uses `"$VENV_PY" -m pip` for cross-platform pip invocation. Added `status.py --version` smoke test after install.
- **`scripts/update.sh`**: Changed hook installation from `ln -sf` (symlink) to `cp` + `chmod +x` (copy), matching `install-hooks.sh`.
- **`scripts/upgrade.sh`**: Replaced all `ln -sf` and symlink-checking logic (`readlink`, `-L`) with `cp` + `chmod +x`. Simplified hook-exists warning to point to `install-hooks.sh`.
- **New tests**: `tests/test_install_update_flow.py` -- 15 tests covering git URL, venv paths, version check, and hook consistency. Full suite: 424 passed.
### Added — loop engineering v1, self-improvement loop default-on (task `add-self-improvement-loop`)
- **`scripts/install.sh`**: after clone and guard registration, creates and schedules the self-improvement loop (`--create-loop self-improvement --from-template self-improvement --project "$FRAMEWORK_DIR"` + `--install-schedule self-improvement --interval 3600`). Both commands use `|| true` so the framework continues to work even if loop creation fails. Prints a user-facing message with opt-out instructions (`--pause-loop self-improvement`).
- **`scripts/update.sh`**: idempotent bootstrap for existing users. Checks `if [ ! -d "$FRAMEWORK_DIR/loops/self-improvement" ]` before creating. Same `|| true` non-fatal behavior.
- **New tests**: `tests/test_self_improvement_loop.py` -- 16 tests covering install.sh wiring (5), update.sh wiring (4), template fields (5), and loop creation from template (2). Full suite: 409 passed.
- **Doc updates**: `README.md` Loop Engineering section notes default-on; `design/loops/technical.md` section 9 notes install.sh creates it.
### Added — loop engineering v1, templates and onboarding (task `add-loop-templates-onboarding`)
- **Prompt-file token substitution**: `_resolve_prompt(prompt_ref, extras, loop_path, tick_num, role)` in `scripts/loop-runner.py` resolves prompt refs to full paths (searches `<loop_path>/<ref>` then `~/.automaton/prompts/<ref>`), reads the file, substitutes content-level tokens (`{task_brief}`, `{acceptance_criteria}`, `{next_hint}`, `{current_task}`, `{current_phase}`, `{verdict}`, `{artifact_content}`), writes to `<loop_path>/outputs/tickN-<role>-prompt.md`, and returns the temp path. Falls back to raw `prompt_ref` when file not found (backward compat). `{artifact_content}` reads the file at `extras["artifact"]` and substitutes its content (empty string if missing).
- **`_invoke_harness` extended**: new optional `loop_path` and `tick_num` params. When `loop_path` is provided, calls `_resolve_prompt` before building the harness command. All three call sites in `cmd_tick` (implement, verify, orchestrate) now pass `loop_path` and `tick_num`.
- **New prompt files**: `prompts/loop-implement.md` (Implement role with task_brief/criteria/hint tokens, ALLOWED/FORBIDDEN sections), `prompts/loop-verifier.md` (Verify role with strict JSON output, score rubric 0.0-1.0, artifact_content token), `prompts/loop-orchestrate.md` (Orchestrate role with verdict token, phase transition logic, no-edit/no-auto-approve rules).
- **`templates/loops/ci-triage/loop.json`**: `roles` filled from `null` to `{"prompt": "loop-implement.md"}` etc.
- **`templates/loops/self-improvement/loop.json`**: new template with `work_source: audit`, `blast_radius.use_worktree: true`, `file_scope: ["scripts/", "prompts/", "tests/", "design/"]`, `brakes.max_iterations: 10`, `score_plateau_window: 3`.
- **README.md**: new "Loop Engineering" onboarding section (quick start, tick cycle, configuration, monitoring, halt/resume).
- **`design/loops/technical.md` section 8**: documented prompt resolution and token substitution flow.
- **New tests**: `tests/test_loop_templates.py` -- 18 tests covering R1-R6 (resolve_prompt, prompt file content, ci-triage template, self-improvement template, tick integration). Full suite: 393 passed.
- **Test infrastructure**: updated `tests/test_loop_runner.py`, `tests/test_blast_radius.py`, `tests/test_goal_mode.py` to use non-existent prompt refs (`test-impl.md` etc.) so `_resolve_prompt` fallback path is exercised in those tests. Added loop prompts to self-consistency test exclusion set.
### Added — loop engineering v1, blast-radius scheduler (task `add-blast-radius-scheduler`)
- **`scripts/loop-runner.py` extended**: `_ensure_worktree(state, cfg, loop_path, project_dir)` creates a per-loop git worktree at `<loop>/worktree` on branch `loop/<name>` when `blast_radius.use_worktree` is true (default) and no worktree exists yet. Records `worktree_path` and `worktree_branch` in `.state.loop` atomically. Reuses existing worktree on subsequent ticks. Falls back to project root with a WARNING log when: not a git repo, git binary missing, or `git worktree add` fails. Handles branch-already-exists by retrying without `-b`. Stale `worktree_path` (directory deleted) is cleared and worktree recreated.
- **New tests**: `tests/test_blast_radius.py` -- 15 tests covering worktree creation, reuse, fallback, branch-exists retry, state consistency, tick integration, platform paths, and backward compat. Full suite: 369 passed (was 354 + 15 new).
### Added — loop engineering v1, goal-mode (task `add-goal-mode`)
- **`scripts/loop-runner.py` extended**: `_find_work(state, cfg, loop_path, project_dir)` replaces the inline `single`-only block, dispatching on `loop.json` `work_source.kind`:
- `"single"`: unchanged behavior (uses `.state.loop` `current_task`).
- `"audit"`: runs `status.py --audit --json --project <p>`, picks the highest-severity unresolved violation. If the violation has a task, uses it; otherwise slugifies the message and calls `--create-task`, then sets the new task as `current_task`. When no unresolved violations → SKIP `no_work` (clean scheduler exit; does not increment `iteration_count`). Optional `work_source.project` overrides the audit target.
- `"backlog"`: reads `<root>/design/<area>/BACKLOG.md` (`work_source.area` defaults to `"loops"`); picks the topmost `- [ ]` item. Empty backlog → SKIP `no_work`.
- **Goal-oriented substitution tokens**: three new tokens available in `harness.command`:
- `{task_brief}` -- from `<task>/RESEARCH.md` or `DESIGN.md` or `SPEC.md` (first hit), capped at 4k tokens.
- `{acceptance_criteria}` -- from `loop.json` `acceptance_criteria` (string OR list joined by newlines), capped at 2k tokens.
- `{next_hint}` -- from `state.last_verdict.next_hint` (empty on first tick / after `--approve`), capped at 1k tokens.
- Closes the `next_hint` feedback loop: tick N's verifier hint becomes tick N+1's Implement/Verify context (D7 graded verifier feedback).
- **`_truncate_tokens(text, max_tokens)`** stdlib-only approximate cap (4-chars-per-token heuristic, `…[truncated]` marker). No tokenizer dependency.
- **`scripts/status.py --audit --json`**: machine-readable audit mode. Emits a single JSON line on stdout: `{"violations": [...], "loops": [...], "total_tasks": N, "untracked_tasks": M}`. Each violation carries `category`, `severity` (`high`/`med`/`low`), `task`, `message`, `resolved: false`. Existing human-readable `--audit` output is unchanged when `--json` is absent. Backed by new `_audit_collect` and `_audit_category3_paths` helpers.
- **`templates/loops/ci-triage/loop.json`**: added `"work_source": {"kind": "single"}` and an `"acceptance_criteria"` example so the template is self-documenting.
- **Backward compat**: missing `work_source` or unknown `kind` falls back to `"single"` with a `.state.log` WARNING entry. Existing task-3 fixtures tick identically.
- **New tests**: `tests/test_goal_mode.py` -- 26 tests covering R1-R8 and one regression. All subprocess calls stubbed via `monkeypatch`; no live LLM in CI. Full suite: 354 passed (was 328 + 26 new).
- **Doc updates**: `design/loops/technical.md` §9 self-improvement template now shows `acceptance_criteria`; `design/loops/functional.md` §9 loop.json fields list documents `work_source` and `acceptance_criteria`.
### Added — loop engineering v1, loop runner (task `add-loop-runner`)
- **`scripts/loop-runner.py`**: per-tick engine. `--mode tick` runs one tick (gate -> find work -> spawn Implement -> spawn Verify -> parse graded JSON verdict -> spawn Orchestrate -> atomic state write -> log); `--mode daemon` runs a sleep loop bounded by `--max-iterations`. Calls `status.py --check-gate` first; any non-ok gate exits 0 (clean scheduler exit). Calls `vram_detect.py --loop-mode --json` before any harness subprocess; refuses below the 16k floor (D13) with `human_intervention`. Idempotent in failure -- pre-step-10 crashes do not corrupt `.state.loop` or advance `iteration_count`.
- **Graded verifier protocol**: `parse_verdict` accepts raw JSON, ```json fenced blocks, or JSON with `//`/`#` line comments. Required keys: `pass` (bool), `score` (float). Optional: `reasons` (list[str]), `next_hint` (str). Parse failure halts as `verifier_failed` without advancing state.
- **Harness command substitution**: `loop.json` `harness.command` (list of strings) with tokens `{prompt}`, `{cwd}`, `{output}`, `{artifact}`, `{verdict}`, `{current_task}`, `{current_phase}`. Default command (v1.1, see `fix-harness-command-template` above): `["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]` (old default with `--prompt-file`/`--cwd` was non-functional; fixed in v1.1). Stdlib only, no new pip deps.
- **Score history capping**: per-tick score appended to `score_history`, capped at `brakes.score_plateau_window` (oldest dropped). Score plateau halts via the next tick's `--check-gate`.
- **Tick log**: every state-changing op appends an ISO-timestamped `TICK pass=<bool> score=<f> iter=<N>` line. Daemon mode appends `DAEMON_STOPPED` on KeyboardInterrupt.
- **New tests**: `tests/test_loop_runner.py` -- 18 tests across 7 classes cover R1--R8. All subprocess calls stubbed via monkeypatch (no live LLM in CI). Full suite: 328 passed (was 310 + 18 new).
- **Doc updates**: `AGENTS.md` "Loop runner" bullet under State Enforcement -- Loops (v1); `README.md` loop-runner one-liner in the Loop Engineering (beta) section.
### Added — loop engineering v1, brakes layer (task `add-status-brakes`)
- **`.state.loop` runtime state**: single source of truth per loop at `{project}/.automaton/loops/<name>/.state.loop` (framework-internal at `~/.automaton/loops/<name>/`). Schema v1 with 13 fields (`status`, `halt_reason`, `iteration_count`, `resumed_count`, `last_tick_at`, `last_verdict`, `score_history`, `current_task`, `worktree_branch`, `worktree_path`). Atomic tmp+rename writes.
- **`--create-loop NAME [--from-template T]`**: only way to bootstrap a loop dir + `.state.loop` + `loop.json`. Refuses non-kebab names, duplicates, unknown templates; patches `name` into the copied `loop.json`.
- **`--check-gate NAME [--json]`**: runs 6 brake gates in order (status, iterations, budget, task phase, worktree drift, score plateau). First failure halts the loop, best-effort disables the OS schedule unit, emits structured JSON verdict.
- **`--can-continue NAME [--json]`**: cheap pre-tick probe — `ok := status == "running"`.
- **`--approve --loop NAME`**: the only way to clear a halt. Increments `resumed_count`. No auto-approve in v1 (D4).
- **`--pause-loop` / `--resume-loop`**: user-controlled soft stop; cannot clear halts.
- **`--install-schedule NAME [--interval S]`**: triple-dispatched (Darwin launchd plist / Linux crontab block / Windows schtasks) native schedule unit generator per `platform.system()`. Generates `automaton-loop-tick.sh` / `automaton-loop-tick.bat` stub (self-documenting: when the scheduler fires it, the filename alone says what it does).
- **`--can-edit --loop NAME [--loop-worktree] --file P`**: worktree scope check — is the file inside this loop's declared `blast_radius.file_scope`?
- **`--transition` halt refusal**: refuses to transition any task owned by a HALTED loop until `--approve --loop` clears the halt.
- **`--audit` Cat-6 Loops block + `--loop-list`**: flags halted / untracked / stale-running loops and loops whose `current_task` no longer exists. `_audit_loops_block` runs even when no tasks are present.
- **`--version`**: prints framework version parsed from `config.md`'s `## Framework Version` section.
- **`.state.log` tick trail**: every state-changing loop op appends an ISO-timestamped line (PAUSED / RESUMED / APPROVED / HALT).
- **New loop template**: `templates/loops/ci-triage/loop.json` (minimal; full template expansion lands in task `add-loop-templates-onboarding`).
- **New tests**: `tests/test_status_brakes.py` — 46 tests across 10 classes cover R1–R10. Full suite: 310 passed (was 264 + 46 new).
- **Doc updates**: `AGENTS.md` Harness Integration modes block extended with the `--loop` worktree-scope mode; new "State Enforcement — Loops (v1)" section in `AGENTS.md`.
### Added — loop engineering v1 (design only; implementation pending bootstrap tasks)
- **Loop system design**: `design/loops/{functional,technical,README,BACKLOG}.md` — locked v1 design for an unattended, state-enforced loop runner built on top of the existing `status.py` phase machine. No second enforcement surface.
- **Five deaths halt model**: `iterations_exhausted`, `budget_exhausted`, `verifier_failed`, `drift_detected`, `human_intervention`. All halts require human `--approve --loop` to resume; no auto-approve path (D4).
- **Tiered role budgets**: three session roles (`Implement:`, `Verify:`, `Orchestrate:`) with explicit context tiers. 16k floor hard refuse (D13). Session divergence mandatory; model divergence only when `Verify:` != `Implement:` (D12).
- **Per-loop git worktree** as blast radius default (D2); `--no-worktree` opt-out.
- **Native scheduler generator**: `platform.system()`-dispatched `launchd` / `crontab` / `schtasks` unit generation via `status.py --install-schedule` (D1); `--daemon` opt-in fallback.
- **Self-improvement loop template** at `templates/loops/self-improvement/`, **default-on at install** (D21). Ticks against `status.py --audit` on the framework's own repo — the literal seed of self-management.
- **Bootstrap task plan** (8 tasks) committed to `design/loops/README.md` — `fix-context-sizing`, `add-status-brakes`, `add-loop-runner`, `add-goal-mode`, `add-blast-radius-scheduler`, `add-loop-templates-onboarding`, `add-self-improvement-loop`, `fix-install-update-flow`. These are the **last tasks a human creates by hand**; after task 7 lands, loops create subsequent tasks from `BACKLOG.md` and `--audit` output (D24, D25).
- **Scoped deferrals** recorded in `BACKLOG.md`: Scope 2 (design-update loop) → v1.1; Scope 3 (self-designing loops) + parallel-mode-default + auto-approve-relax → deferred indefinitely (D20).
### Added
- **Added:** `requirements.txt` pinning `pytest==7.4.4` for reproducible test runs.
- **Added:** `scripts/install.sh` now creates `.venv/` and installs pytest into it.
- **Harness pre-edit hook**: `--can-edit` now supports project-level checks without `--task`, file scope checks with `--file`, and `--json` output for machine-readable harness integration
- **opencode plugin**: `plugins/automaton-guard/plugin.ts` — intercepts `edit` and `write` tool calls, calls `--can-edit` before allowing modifications
- **Git pre-commit hook**: `scripts/git-hooks/pre-commit` — blocks commits when no task is in an edit-allowed phase (universal safety net for all harnesses)
- **Pre-v2.0 task enforcement**: Tasks without `.state` files are UNTRACKED — `--transition`, `--can-edit`, `--task`, and `--approve` all refuse to operate on them
- **New `--upgrade` command**: Bootstraps `.state` files for pre-v2.0 tasks (single task with `--task` or all tasks at once)
- **Untracked task reporting**: `--list` shows `UNTRACKED (no .state)` for tasks without `.state` files instead of silently bootstrapping
- **Project scoping fix**: `status.py` errors when no project is detected instead of silently falling back to framework directory
- **Scope check fix**: `--scope-check` marks framework files as OUT_OF_SCOPE when working on a project
- **Dashboard scope fix**: Handler methods use stored `project_root` instead of re-detecting from CWD on every request
- **`--project` flag**: Added to all status.py command invocations across 16+ prompt and config files
- **`_infer_state_from_artifacts` locked to `--upgrade`**: Removed as silent fallback from all operational commands
- **Phase approval gates**: Research, Decomposition, Design, and Test Design phases now require explicit user approval (`:awaiting_approval` → `:approved`) before proceeding
- **status.py script**: Comprehensive enforcement and status tool with `--task`, `--list`, `--create-task`, `--transition`, `--approve`, `--validate-folder`, `--audit`, `--claim`, `--release`, `--next-available`, `--available`, `--can-edit`, `--scope-check`, `--same-session`, `--upgrade`
- **Untracked task enforcement**: Tasks without `.state` files are UNTRACKED — `--transition`, `--can-edit`, `--task`, `--approve` all refuse to operate on them. Run `--upgrade` to bootstrap `.state` files
- **`--project` flag**: All `status.py` commands now support `--project` for explicit project scoping when multiple projects exist on the same machine
- **`--upgrade` command**: Bootstraps `.state` files for pre-v2.0 tasks that lack them (single task with `--task` or all tasks at once)
- **Project scoping**: `status.py` now errors when not in a project directory and `--project` is not specified, instead of silently falling back to `~/.automaton/`
- **Scope check fix**: `--scope-check` now correctly marks framework files as OUT_OF_SCOPE when working on a project (was incorrectly always IN_SCOPE)
- **Dashboard scope fix**: Dashboard handler methods now use stored `project_root` and `scope` instead of re-detecting from CWD on every request
- **Phase-scoped prompts**: All phase prompts now include ALLOWED ACTIONS, FORBIDDEN ACTIONS, approval gates (where applicable), pre-work validation, and `.state` precondition checks
- **Orchestrator restructuring**: Reduced from 493 lines to 143 lines; sub-task management extracted to `subtask_management.md`; state machine reference moved to `workflow.md`
- **ALLOWED/FORBIDDEN enforcement**: Each phase prompt explicitly defines what agents can and cannot do, with user override resistance instructions
- **Workflow enforcement**: `--transition` refuses illegal phase transitions; `--validate-folder` detects out-of-order artifacts; `--audit` checks all tasks for violations
- **Task creation gate**: `status.py --create-task` is the only valid way to create tasks; `--audit` flags manually created folders
- **Approval log**: `.state.approvals` file records all user approvals with timestamp and approver
- **Multi-agent support**: Optional `Agent Configuration` section in `.agent.md` enables task claiming, role binding, and work discovery for multi-agent setups
- **Tool integration hooks**: `--can-edit`, `--scope-check`, `--same-session` for agent tool integrations (optional, not called by prompts)
- **upgrade.sh script**: Bootstraps `.state` files for existing tasks from artifact heuristic
- **Framework version marker**: `config.md` now includes version 2.0 with state enforcement indicator
### Changed
- **Changed:** All documented `python` invocations now read `python3` (stock macOS / Windows Python ship as `python3`).
- **orchestrate.md**: Reduced from 493 to 143 lines; gate-check loop replaces soft advisory approach; approval gates enforced at research, decomposition, design, and test_design
- **workflow.md**: Rewritten to reference `.state` as canonical phase indicator; approval sub-states documented; enforcement via `status.py` documented
- **All phase prompts**: Added `.state` precondition check, pre-work validation, ALLOWED/FORBIDDEN sections, handling user overrides
- **research.md, design.md, decompose.md, test_design.md**: Added approval gate sections with `--transition {phase}:awaiting_approval` and `--approve`
- **implement.md, bug_finder.md, adversarial_bug_find.md, doc_review.md, referee.md**: Added no-approval-gate notes with direct `--transition` instructions
- `status_reason` property on Task model showing human-readable explanation for each state (#task-status-reason)
- Revoke buttons for approved/changes_requested reviews — replaces approve/request-changes with a single revoke option (#task-status-reason)
- pytest test suite covering dashboard core, app security, and VRAM detection (#add-pytest-test-suite)
- Structured verdict parsing: `parse_verdict_status()` uses `## Status:` line before substring fallback, preventing false-BLOCKED classification (#fix-verdict-parsing)
- State machine alignment: IMPLEMENTATION.md alone → Bug Find, ADVERSARIAL_BUG_REPORT alone → Bug Find (matching orchestrator spec) (#fix-verdict-parsing)
- Filesystem task name validation: `discover_tasks()` and `parse_sub_tasks()` skip directories with invalid characters (#fix-verdict-parsing)
- Added CORS headers, `do_OPTIONS` handler, `X-Content-Type-Options` to all dashboard API responses (#harden-dashboard-security)
- Added POST content-length bounds (64KB) and review comment length limits (4096 chars) (#harden-dashboard-security)
- Replaced inline `onclick` review handlers with `data-*` attributes and event delegation (#harden-dashboard-security)
- Applied `escapeHtml()` to task `display_name` in dashboard card rendering (#harden-dashboard-security)
- `GET /api/config` and `PUT /api/config` endpoints for reading and persisting dashboard configuration (#wire-dashboard-config)
- Server-side task cache with 1s TTL to eliminate redundant disk I/O on every polling request (#wire-dashboard-config)
- Dashboard JS applies config on init: theme, default_view, auto_refresh_interval, column_width, show_timelines (#wire-dashboard-config)
- Review POSTinvalidates task cache so next poll picks up changes (#wire-dashboard-config)
- `decomposition_content`, `parent_spec_content`, `vram_config_content` fields on `Task` model (#add-decomposition-content)
- `WaveGroup` dataclass and `parse_waves()` for extracting wave structure from DECOMPOSITION.md (#add-decomposition-content)
- `parse_vram_config()` for reading VRAM_CONFIG.md (#add-decomposition-content)
- Dashboard JS wave statistics use parsed wave data instead of 50/50 heuristic (#add-decomposition-content)
- Detail panel shows Decomposition, Parent Context, and VRAM Configuration sections (#add-decomposition-content)
### Changed
- Removed stale `dashboard = ["inotify>=0.2"]` optional dependency from pyproject.toml (#cleanup-cruft)
- Deleted `debug_root.py` stray development script (#cleanup-cruft)
- Deleted empty `automaton/dashboard/ui/widgets/` directory (#cleanup-cruft)
- Fixed `config.md` RAM detection description (was "via `free`", now "via `/proc/meminfo` or `sysctl`") (#cleanup-cruft)
- `_find_tasks_dir()` returns `Path` instead of `Path | None`, removed tautological condition (#cleanup-cruft)
- Removed `sys.path.insert` hack from `__main__.py` (#cleanup-cruft)
- Documented `scripts/dashboard.sh` convenience wrapper in README.md (#cleanup-cruft)
- Framework self-consistency test suite: 17 tests covering prompt stop conditions, hardcoded URLs, canonical paths, .rules.md sections, stale dependencies, CSS theme parity, verdict regression, and CI validation (#framework-self-consistency-tests)
### Fixed
- REFEREE state was never produced by state machine — verdict with unparseable status now correctly shows as REFEREE instead of silently falling through to earlier states (#task-status-reason)
- Pending review count in header now excludes done/blocked tasks (#task-status-reason)
- Critical: PASS verdicts mentioning FAIL/NEEDS_REVIEW in body text were falsely classified as BLOCKED (#fix-verdict-parsing)
- State divergence: IMPLEMENTATION.md alone showed "Implement" instead of "Bug Find" (#fix-verdict-parsing)
- Added mandatory stop conditions to `bug_finder.md` and `adversarial_bug_find.md` (#fix-prompt-consistency)
- Fixed deprecated `{project}/tasks/` path in `onboarding.md` (#fix-prompt-consistency)
- Expanded prompt path test to catch concrete deprecated path patterns (#fix-prompt-consistency)
- Root `pyproject.toml` with optional test/dashboard dependency groups (#add-pytest-test-suite)
- `AGENTS.md` with build/test commands and conventions (#developer-experience-gitea-ci)
- `.gitea/workflows/ci.yml` running py_compile, pytest, and shell script syntax checks (#developer-experience-gitea-ci)
- `templates/README.md` documenting the task template examples (#developer-experience-gitea-ci)
- Blocked phase column between Verification and Resolution on dashboard (#additive-extension-model)
- Framework self-enforcement rules in .rules.md and system-prompt.md (#framework-self-enforcement)
- Additive extension model: projects extend via extensions/ dir, never copy framework files (#additive-extension-model)
- CHANGELOG.md for release notes tracking (#changelog)
- Framework audit: comprehensive self-consistency check with RESEARCH.md (#framework-audit)
- Audit Bug 1: `--audit` category 3 now checks `.automaton/tasks/` paths (was only checking `tasks/`) (#fix-cat3-audit-paths)
- Audit Bug 2: `migrate-project.sh` find command now has parentheses around `-name` group for correct `-prune` binding (#fix-migrate-find-precedence)
- Audit Bug 3: `_lookup_model_context()` no longer false-matches model prefixes (e.g. `phi-4` matching `phi-4-mini`) — uses three-tier matching with known suffix whitelist (#fix-vram-model-prefix-match)
- Audit Bug 4: Verdict PASS/FAIL inference uses structured `## Status:` line parsing instead of fragile substring search (#fix-verdict-pass-inference)
- Audit Bug 5: `register-guards.sh` now checks both `.json`/`.jsonc`, writes to `plugin` (singular) key, and strips `//` comments before `json.loads()` (#fix-register-guards)
- Audit Bug 6: Dashboard `determine_task_state()` now reads `.state` file (source of truth) before falling back to artifact heuristic (#fix-dashboard-read-state)
- Audit Bug 7: `--can-edit` and `--scope-check` path prefix matching uses `os.sep` boundary to prevent sibling directory false matches (#fix-can-edit-path-prefix)
- Audit Bug 8: Removed wildcard CORS `Access-Control-Allow-Origin: *` from dashboard — replaced with security headers (`X-Content-Type-Options`, `X-Frame-Options`) (#fix-dashboard-cors-origin)
- Audit Bug 9: Stale-task detection uses `.state.lastedit` timestamp (touched on actual edit activity) instead of `.state` mtime (which only reflects phase transitions) (#fix-stale-task-mtime-proxy)
- Audit Bug 10: TEST_PLAN.md now correctly maps to `test_design` phase (was mapping to `implement`) in both `status.py` and dashboard `task.py` (#fix-test-plan-phase-mapping)
### Changed
- All prompts now use the canonical task path `{project}/.automaton/tasks/{task-name}/` (#standardize-task-path-conventions)
- `scripts/vram_detect.sh` rewritten as `scripts/vram_detect.py` for testability and correctness (#rewrite-vram-detection-python)
- `tasks/dashboard-spec.md` reconciled with the implemented web dashboard (#reconcile-dashboard-spec)
- `automaton/dashboard/README.md` and help modal shortcuts now match the web UI (#reconcile-dashboard-spec)
- prompts/orchestrate.md: always reads prompts/contracts/scripts from global, project extensions are additive (#additive-extension-model)
- prompts/onboarding.md: removed diff/merge upgrade, replaced with migration check (#additive-extension-model)
- README.md: updated upgrade docs for new additive model (#additive-extension-model); added Dashboard section (#dashboard-task-review)
- scripts/update.sh: simplified to plain git pull (#additive-extension-model)
- .rules.md: converted from template to concrete rules with Task-Driven Development, VRAM-aware sizing, Changelog, and Self-Improvement sections (#framework-self-enforcement)
- system-prompt.md: added instruction to read global .rules.md (#framework-self-enforcement)
- automaton/dashboard/ui/app.py: added review API endpoints (GET/POST /api/task/{name}/review), spec_content in responses, unquote() for URL-encoded task names, path traversal fix (#dashboard-task-review, #spec-in-detail)
- automaton/dashboard/html/dashboard.js: review UI (badges, buttons, filter), artifact badges, specification display, modal conversion, textarea replacement, display group for approved planning tasks (#dashboard-task-review, #artifact-badges, #spec-in-detail, #task-detail-modal, #review-textarea)
- automaton/dashboard/html/styles.css: review components, artifact badges, modal layout, textarea styles (#dashboard-task-review, #artifact-badges, #task-detail-modal, #review-textarea)
- automaton/dashboard/html/index.html: review filter, pending count, modal overlay (#dashboard-task-review, #task-detail-modal)
- automaton/dashboard/core/task.py: fixed state machine priority — IMPLEMENTATION.md now correctly detected, DOC_REVIEW checked before BUG_REPORT (#implement-task)
- automaton/dashboard/core/board.py: fixed KanbanBoard — added missing COLUMNS and __init__ (#implement-task)
- automaton/dashboard/core/refresh.py: improved inotify error handling with explicit fallback messages (#implement-task)
### Fixed
- VRAM detection: undefined headroom, hardcoded JSON headroom, and code-block config parsing (#rewrite-vram-detection-python)
- VRAM detection: 10KB file-read limit now enforced for API config files (#rewrite-vram-detection-python)
- Dashboard static file serving: replaced string-prefix path traversal check with `Path.relative_to()` (#harden-dashboard-security-scripts)
- Dashboard task name validation: restricted to `[A-Za-z0-9_-]+` (#harden-dashboard-security-scripts)
- `scripts/update.sh`: now warns and aborts on uncommitted changes before pulling (#harden-dashboard-security-scripts)
- README/install.sh: replaced placeholder repository URL with real Gitea URL (#harden-dashboard-security-scripts)
- State machine: IMPLEMENTATION.md was never checked in determine_task_state(), tasks showed as RESEARCH (#implement-task)
- State machine: DOC_REVIEW checked after BUG_REPORT — wrong priority order (#implement-task)
- Path traversal: review API accepted task names with ../ allowing writes outside tasks directory (#dashboard-task-review)
- URL encoding: task names with spaces in API paths were not decoded (#implement-task)
- Review parsing: comment extraction used fragile conditional, falsy comments (e.g., "0") skipped (#implement-task)
- Board display: approved planning tasks stayed in Planning column instead of advancing to Design (#dashboard-task-review)
### Removed
- `automaton/dashboard/themes.py` (vestigial ANSI theme stub) (#reconcile-dashboard-spec)
- `automaton/dashboard/core/refresh.py` (half-implemented file watcher; dashboard uses JS polling) (#remove-file-system-watcher)
- `templates/contract-template.md` (unused) (#developer-experience-gitea-ci)
- `automaton/dashboard/pyproject.toml` (consolidated into root `pyproject.toml`) (#add-pytest-test-suite)
### Migration
- Project migration script for old-model projects: scripts/migrate-project.sh (#project-migration)
- Project migration detection in onboarding.md (#project-migration)