Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

5.2 KiB

ADVERSARIAL_BUG_REPORT: add-loop-runner

Attack the runner as a hostile cyclic process would: find ways to escape the brakes, inflate iteration counts, burn tokens through the harness, or hide halt state.

Attack vectors tried

A1 — Can a tick fraudulently increment iteration_count by writing a bogus verdict?

No — parse_verdict requires pass and score keys; if missing, returns None and the tick halts verifier_failed without writing state. A hostile verifier subprocess that returns malformed JSON halts the loop, not advances it. ✅ Defended.

A2 — Can the runner be coerced into running past max_iterations?

_gate_iterations (status.py, called via --check-gate at step 2) refuses when iteration_count >= max_iterations. The runner's step 10 increments iteration_count only after the gate has passed. If the gate passed at iteration N (count=N), the tick runs, count becomes N+1. Next tick the gate sees count=N+1; if that exceeds max, it halts. So the runner increments AT MOST once per gate-pass. ✅ Defended.

BUT: there's a TOCTOU window. Between --check-gate returning ok at iteration 9 (max=10) and the runner's step 10 writing iteration_count=10, a concurrent second tick could also pass the gate at count=9 and increment to 10. Both ticks would then run. Loop max_iterations=10 could fire 11 ticks. Window: the OS scheduler interval typically >> tick duration, so this requires two schedulers firing the same loop concurrently (user error). Halt is not corrupted; the next gate halts iterations_exhausted correctly.

Mitigation: documented runbook says "one scheduler per loop". Adding a file-lock would close this (status.py's .state.loop write is atomic but doesn't lock). Logged to BACKLOG (same item A6 from add-status-brakes). Not blocking.

A3 — Can the orchestrator role itself escape enforcement?

The orchestrator role is spawned with full substitution and runs in the same cwd as the verifier. It's expected to call status.py --transition itself. If a hostile orchestrator calls status.py --transition on a task the loop doesn't own, status.py's R8 halt-refusal only kicks in if the owning loop is halted — a non-halted loop's orchestrator could transition arbitrary tasks. This is the runner contract: the orchestrator's prompt (task 6) must restrict it to current_task. The runner does not enforce orchestrator behavior beyond passing it the verdict and current_task. The defense is prompt-based + the file_scope gate (--can-edit --loop --file) for any file edits. ✅ Defended by layered enforcement (prompt + can-edit + R8).

A4 — Can a hostile harness command execute shell injection?

subprocess.run(final_argv, ...) uses list argv (no shell). Tokens are substituted as raw strings, but no shell=True. A malicious harness.command in loop.json could include "rm -rf /" as a literal argv element — but that's a config-trust issue, not a runtime escape. The loop.json is controlled by the human operator who created the loop. ✅ Accepted threat model.

A5 — Can the runner be pointed at a different project via --project to escape scope?

cmd_tick resolves project_dir from args.project and uses it for _loop_dir and cwd. If a hostile caller passes --project /etc, the runner will look for .automaton/loops/<name> under /etc — which won't exist — and skip untracked. No escape. ✅ Defended.

A6 — Verdict score outside [0, 1]?

parse_verdict does float(data.get("score", 0.0)). A hostile verifier returning score: 99999 would inflate score_history. The score-plateau gate checks "flat or non-increasing" so inflation actually breaks a plateau (good for the attacker — loop continues). No hard cap on score. Acceptable for v1: the score is informational; verifier-prompt contract (task 6) will say "score in [0, 1]". Could clamp in parse_verdict for safety; noted for v1.1. Not blocking.

A7 — Can the OS scheduler fire a tick while the runner is mid-tick?

OS unit fires automaton-loop-tick.sh which invokes loop-runner.py --mode tick. If the previous tick is still running, two cmd_tick instances run concurrently. Both might pass --check-gate, both might invoke harness subprocesses, both might write state (atomic last-writer-wins). Result: double-spent tokens for one iteration count increment. Mitigation: scheduler interval should exceed tick duration; lock-file in v1.1. Same TOCTOU as A2; same BACKLOG item.

A8 — Can a corrupt loop.json crash the runner?

_read_loop_config returns None on JSON parse failure. cmd_tick calls (cfg or {}) for all .get() accesses. No crash. ✅ Defended.

Hardening recommendations (for BACKLOG)

  1. fcntl lock on .state.loop would close A2/A7 TOCTOU (same item as add-status-brakes A6).
  2. parse_verdict should clamp score to [0, 1] and reject non-bool pass strings (O6 + A6).
  3. outputs.retention in loop.json (O5) + automatic pruning in the runner.

All three are explicit follow-ups; none block task 3.

Verdict

PASS — no exploitable escape. The runner enforces the contract; remaining race windows are bounded by the scheduler interval and accept-rate; mitigations are explicit v1.1 hardening.