Files
automaton/tasks/complete/add-loop-runner/ADVERSARIAL_BUG_REPORT.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

5.2 KiB

ADVERSARIAL_BUG_REPORT: add-loop-runner

Attack the runner as a hostile cyclic process would: find ways to escape the brakes, inflate iteration counts, burn tokens through the harness, or hide halt state.

Attack vectors tried

A1 — Can a tick fraudulently increment iteration_count by writing a bogus verdict?

No — parse_verdict requires pass and score keys; if missing, returns None and the tick halts verifier_failed without writing state. A hostile verifier subprocess that returns malformed JSON halts the loop, not advances it. ✅ Defended.

A2 — Can the runner be coerced into running past max_iterations?

_gate_iterations (status.py, called via --check-gate at step 2) refuses when iteration_count >= max_iterations. The runner's step 10 increments iteration_count only after the gate has passed. If the gate passed at iteration N (count=N), the tick runs, count becomes N+1. Next tick the gate sees count=N+1; if that exceeds max, it halts. So the runner increments AT MOST once per gate-pass. ✅ Defended.

BUT: there's a TOCTOU window. Between --check-gate returning ok at iteration 9 (max=10) and the runner's step 10 writing iteration_count=10, a concurrent second tick could also pass the gate at count=9 and increment to 10. Both ticks would then run. Loop max_iterations=10 could fire 11 ticks. Window: the OS scheduler interval typically >> tick duration, so this requires two schedulers firing the same loop concurrently (user error). Halt is not corrupted; the next gate halts iterations_exhausted correctly.

Mitigation: documented runbook says "one scheduler per loop". Adding a file-lock would close this (status.py's .state.loop write is atomic but doesn't lock). Logged to BACKLOG (same item A6 from add-status-brakes). Not blocking.

A3 — Can the orchestrator role itself escape enforcement?

The orchestrator role is spawned with full substitution and runs in the same cwd as the verifier. It's expected to call status.py --transition itself. If a hostile orchestrator calls status.py --transition on a task the loop doesn't own, status.py's R8 halt-refusal only kicks in if the owning loop is halted — a non-halted loop's orchestrator could transition arbitrary tasks. This is the runner contract: the orchestrator's prompt (task 6) must restrict it to current_task. The runner does not enforce orchestrator behavior beyond passing it the verdict and current_task. The defense is prompt-based + the file_scope gate (--can-edit --loop --file) for any file edits. ✅ Defended by layered enforcement (prompt + can-edit + R8).

A4 — Can a hostile harness command execute shell injection?

subprocess.run(final_argv, ...) uses list argv (no shell). Tokens are substituted as raw strings, but no shell=True. A malicious harness.command in loop.json could include "rm -rf /" as a literal argv element — but that's a config-trust issue, not a runtime escape. The loop.json is controlled by the human operator who created the loop. ✅ Accepted threat model.

A5 — Can the runner be pointed at a different project via --project to escape scope?

cmd_tick resolves project_dir from args.project and uses it for _loop_dir and cwd. If a hostile caller passes --project /etc, the runner will look for .automaton/loops/<name> under /etc — which won't exist — and skip untracked. No escape. ✅ Defended.

A6 — Verdict score outside [0, 1]?

parse_verdict does float(data.get("score", 0.0)). A hostile verifier returning score: 99999 would inflate score_history. The score-plateau gate checks "flat or non-increasing" so inflation actually breaks a plateau (good for the attacker — loop continues). No hard cap on score. Acceptable for v1: the score is informational; verifier-prompt contract (task 6) will say "score in [0, 1]". Could clamp in parse_verdict for safety; noted for v1.1. Not blocking.

A7 — Can the OS scheduler fire a tick while the runner is mid-tick?

OS unit fires automaton-loop-tick.sh which invokes loop-runner.py --mode tick. If the previous tick is still running, two cmd_tick instances run concurrently. Both might pass --check-gate, both might invoke harness subprocesses, both might write state (atomic last-writer-wins). Result: double-spent tokens for one iteration count increment. Mitigation: scheduler interval should exceed tick duration; lock-file in v1.1. Same TOCTOU as A2; same BACKLOG item.

A8 — Can a corrupt loop.json crash the runner?

_read_loop_config returns None on JSON parse failure. cmd_tick calls (cfg or {}) for all .get() accesses. No crash. ✅ Defended.

Hardening recommendations (for BACKLOG)

  1. fcntl lock on .state.loop would close A2/A7 TOCTOU (same item as add-status-brakes A6).
  2. parse_verdict should clamp score to [0, 1] and reject non-bool pass strings (O6 + A6).
  3. outputs.retention in loop.json (O5) + automatic pruning in the runner.

All three are explicit follow-ups; none block task 3.

Verdict

PASS — no exploitable escape. The runner enforces the contract; remaining race windows are bounded by the scheduler interval and accept-rate; mitigations are explicit v1.1 hardening.