4.0 KiB
Verdict: add-state-loop-lock
Status: PASS
Summary
Closed the TOCTOU read-modify-write race flagged in
add-status-brakes/ADVERSARIAL_BUG_REPORT.md (A6) and
add-loop-runner/ADVERSARIAL_BUG_REPORT.md (A2, A7) by wrapping the
read-modify-write cycles on .state.loop in a cross-process file lock
(_loop_lock). POSIX fcntl.flock(LOCK_EX), Windows
msvcrt.locking(LK_LOCK, 1), per-loop granularity, blocking acquire,
no timeout in v1.1. Stdlib only.
SPEC compliance
| Requirement | Status |
|---|---|
R1 — _loop_lock context manager (POSIX/Windows, blocking, finally-safe FD lifecycle) |
✓ |
R2 — Wrap status.py cmd_pause_loop, cmd_resume_loop, cmd_approve_loop, cmd_check_gate (NOT --create-loop) |
✓ |
R3 — Wrap cmd_tick's step-10 state write; lock covers _gate subprocess + state write |
✓ |
R4 — Idempotence + early-return inside with releases cleanly (try/finally in the context manager, not caller) |
✓ |
R5 — Lock file <loop_path>/.state.lock (per-loop granularity) |
✓ |
R6 — Stdlib only (fcntl POSIX, msvcrt Windows, contextlib, sys, os) |
✓ |
R7 — Existing atomic write (_write_state_loop tmp-then-replace) retained |
✓ |
| D-L1 — Blocking acquire, no timeout | ✓ |
D-L2 — .state.lock per-loop, not GC'd |
✓ |
D-L3 — --create-loop unwrapped |
✓ |
D-L4 — Stdlib only, no filelock package |
✓ |
| D-L5 — Existing atomic write retained (defense-in-depth) | ✓ |
D-L6 (new, necessary for SPEC R2+R3 consistency) — Env-var bypass ($AUTOMATON_NO_LOOP_LOCK=1) avoids self-deadlock when the runner spawns the --check-gate subprocess inside its held lock |
✓ |
Bug reports
- BUG_REPORT: 5 non-blocking observations (O1-O5), all documented behaviors.
- ADVERSARIAL_BUG_REPORT: 8 attack vectors probed (A1-A8). One LOW finding (A4: harness calling loop-control command inside a tick would deadlock; deferred to harness-integration contract docs follow-up). All others verified safe.
Test results
python3 -m py_compile scripts/status.py scripts/loop-runner.py✓python3 -m pytest tests/test_state_loop_lock.py -v— 7 passedpython3 -m pytest tests/ -q— 447 passed (was 440; +7 new; 0 regressions)- Manual:
--create-loopdoes NOT leave.state.lock(D-L3 ✓); first--check-gatedoes;AUTOMATON_NO_LOOP_LOCK=1 status.py --check-gateworks (env var bypass exercised). - Manual adversarial: 5 concurrent
--approve --loopon a halted loop — only one wins (code 0, resumed_count=1); 4 re-read inside the lock and see status=running, exit 1. Race closed. - Manual adversarial: 5 concurrent
--pause-loop— all succeed (idempotent; pause is well-defined on already-paused); final state consistent. - Manual adversarial: monkeypatch
_gateto raise → exit exception → follow-up_loop_lockacquires immediately (lock released in finally).
D-items applied
- D-L1 to D-L6 all locked (see SPEC compliance table).
Subprocess-deadlock avoidance
The original SPEC's R2 and R3 contradict each other (both list --check-gate
to acquire _loop_lock AND the runner to acquire the same lock across the
--check-gate subprocess — would deadlock). Resolved via D-L6 (env-var
bypass). The runner sets $AUTOMATON_NO_LOOP_LOCK=1 in the --check-gate
subprocess's env ONLY (scoped via _run_json's env kwarg, propagated to
_gate's subprocess.run). Harness subprocesses inherit os.environ
unchanged (no env var) so their nested status.py calls lock normally and
serialize against the runner's outer lock (intended for tasks; harness
typically does status.py --transition only which doesn't touch
.state.lock). Documented in _loop_lock's docstring + design doc.
Pipeline
research → research:awaiting_approval → research:approved → implement → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete
Pipeline driven end-to-end. Ready for --transition complete (relocates
to tasks/complete/).