- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
5.2 KiB
ADVERSARIAL_BUG_REPORT: add-loop-runner
Attack the runner as a hostile cyclic process would: find ways to escape the brakes, inflate iteration counts, burn tokens through the harness, or hide halt state.
Attack vectors tried
A1 — Can a tick fraudulently increment iteration_count by writing a bogus verdict?
No — parse_verdict requires pass and score keys; if missing, returns None and the tick halts verifier_failed without writing state. A hostile verifier subprocess that returns malformed JSON halts the loop, not advances it. ✅ Defended.
A2 — Can the runner be coerced into running past max_iterations?
_gate_iterations (status.py, called via --check-gate at step 2) refuses when iteration_count >= max_iterations. The runner's step 10 increments iteration_count only after the gate has passed. If the gate passed at iteration N (count=N), the tick runs, count becomes N+1. Next tick the gate sees count=N+1; if that exceeds max, it halts. So the runner increments AT MOST once per gate-pass. ✅ Defended.
BUT: there's a TOCTOU window. Between --check-gate returning ok at iteration 9 (max=10) and the runner's step 10 writing iteration_count=10, a concurrent second tick could also pass the gate at count=9 and increment to 10. Both ticks would then run. Loop max_iterations=10 could fire 11 ticks. Window: the OS scheduler interval typically >> tick duration, so this requires two schedulers firing the same loop concurrently (user error). Halt is not corrupted; the next gate halts iterations_exhausted correctly.
Mitigation: documented runbook says "one scheduler per loop". Adding a file-lock would close this (status.py's .state.loop write is atomic but doesn't lock). Logged to BACKLOG (same item A6 from add-status-brakes). Not blocking.
A3 — Can the orchestrator role itself escape enforcement?
The orchestrator role is spawned with full substitution and runs in the same cwd as the verifier. It's expected to call status.py --transition itself. If a hostile orchestrator calls status.py --transition on a task the loop doesn't own, status.py's R8 halt-refusal only kicks in if the owning loop is halted — a non-halted loop's orchestrator could transition arbitrary tasks. This is the runner contract: the orchestrator's prompt (task 6) must restrict it to current_task. The runner does not enforce orchestrator behavior beyond passing it the verdict and current_task. The defense is prompt-based + the file_scope gate (--can-edit --loop --file) for any file edits. ✅ Defended by layered enforcement (prompt + can-edit + R8).
A4 — Can a hostile harness command execute shell injection?
subprocess.run(final_argv, ...) uses list argv (no shell). Tokens are substituted as raw strings, but no shell=True. A malicious harness.command in loop.json could include "rm -rf /" as a literal argv element — but that's a config-trust issue, not a runtime escape. The loop.json is controlled by the human operator who created the loop. ✅ Accepted threat model.
A5 — Can the runner be pointed at a different project via --project to escape scope?
cmd_tick resolves project_dir from args.project and uses it for _loop_dir and cwd. If a hostile caller passes --project /etc, the runner will look for .automaton/loops/<name> under /etc — which won't exist — and skip untracked. No escape. ✅ Defended.
A6 — Verdict score outside [0, 1]?
parse_verdict does float(data.get("score", 0.0)). A hostile verifier returning score: 99999 would inflate score_history. The score-plateau gate checks "flat or non-increasing" so inflation actually breaks a plateau (good for the attacker — loop continues). No hard cap on score. Acceptable for v1: the score is informational; verifier-prompt contract (task 6) will say "score in [0, 1]". Could clamp in parse_verdict for safety; noted for v1.1. Not blocking.
A7 — Can the OS scheduler fire a tick while the runner is mid-tick?
OS unit fires automaton-loop-tick.sh which invokes loop-runner.py --mode tick. If the previous tick is still running, two cmd_tick instances run concurrently. Both might pass --check-gate, both might invoke harness subprocesses, both might write state (atomic last-writer-wins). Result: double-spent tokens for one iteration count increment. Mitigation: scheduler interval should exceed tick duration; lock-file in v1.1. Same TOCTOU as A2; same BACKLOG item.
A8 — Can a corrupt loop.json crash the runner?
_read_loop_config returns None on JSON parse failure. cmd_tick calls (cfg or {}) for all .get() accesses. No crash. ✅ Defended.
Hardening recommendations (for BACKLOG)
- fcntl lock on
.state.loopwould close A2/A7 TOCTOU (same item asadd-status-brakesA6). parse_verdictshould clampscoreto[0, 1]and reject non-boolpassstrings (O6 + A6).outputs.retentioninloop.json(O5) + automatic pruning in the runner.
All three are explicit follow-ups; none block task 3.
Verdict
PASS — no exploitable escape. The runner enforces the contract; remaining race windows are bounded by the scheduler interval and accept-rate; mitigations are explicit v1.1 hardening.