Files
automaton/tasks/complete/add-state-loop-lock/BUG_REPORT.md
T

80 lines
3.9 KiB
Markdown

# Bug Report: add-state-loop-lock
Bug_find phase observations. Each observation is non-blocking unless marked BLOCKER.
## O1 — `print(... state['resumed_count'] ...)` after `with _loop_lock` exits, status.py:cmd_approve_loop
`cmd_approve_loop` references `state['resumed_count']` AFTER the `with`
block exits. `state` is in function scope and was assigned inside the with
block; the value is the post-mutation dict. **Not a bug** — confirmed by
tracing the variable lifecycle. Safe.
## O2 — `_disable_schedule` / `_enable_schedule` left OUTSIDE the lock for pause/resume/approve; INSIDE for halt
For `cmd_pause_loop` / `cmd_resume_loop` / `cmd_approve_loop`,
`_disable_schedule` / `_enable_schedule` is called AFTER the `with
_loop_lock` block exits (line ~1950 area, after the lock releases).
For `cmd_check_gate`'s `_halt_loop` call, `_disable_schedule` is called
INSIDE the lock (since `_halt_loop` couples the halt-write with the
schedule disable).
**Transient**: between the loop's `.state.loop` write (inside the lock)
and the subsequent OS schedule unit disable (outside the lock), the OS
scheduler could fire another tick. That tick's `_gate` subprocess reads
`status=paused` and exits 0 (clean scheduler self-skip). So no real
over-tick — just a no-op tick for ~100ms. Same for resume/approve.
**Not a bug** — documented behavior; matches SPEC R4 (idempotence inside
the lock scope; OS-level schedule toggles are out-of-band best-effort).
The transient inconsistency is harmless because `--check-gate` already
self-skips on non-running.
## O3 — `_gate` subprocess acquires status.py's `_loop_lock`, honors env-var bypass
If a future caller of `status.py --check-gate` manually sets
`$AUTOMATON_NO_LOOP_LOCK=1` in their shell, `_loop_lock` becomes a no-op
even when invoked standalone. **Not a bug**: the env var is a documented
escape hatch; a manual user who sets it accepts that the lock is bypassed.
The runner's own subprocess env is private to the subprocess (passed via
the `env` kwarg to `subprocess.run` in `_run_json` invoked from `_gate`).
The harness subprocesses do NOT inherit the var (verified: `subprocess.run`
without `env` inherits `os.environ`, which is unmodified at runner top
level).
Risk assessment: HIGH only if a user wraps `status.py` invocations with
`AUTOMATON_NO_LOOP_LOCK=1` AND expects pause-loop / approve-loop /
check-gate invocations to serialize. Documented in `_loop_lock`'s
docstring. **Not a bug** — escape hatch has explicit semver-stable
contract.
## O4 — `_loop_lock` is non-re-entrant across processes
POSIX `flock` is per-fd-per-process: a second process blocks cleanly
waiting for the first to release. POSIX `flock` IS re-entrant within a
single process on a single fd. Windows `msvcrt.locking` is NOT re-entrant
within a single process (would deadlock on re-acquire). Documented in
`_loop_lock`'s docstring.
Audit shows no nested `_loop_lock` callsites. **Not a bug** — explicitly
forbidden by the SPEC ("Audit every callsite to ensure no nested
`_loop_lock` within the same `with` block"). Audited in
IMPLEMENTATION.md's "NESTED-LOCK AUDIT" section.
## O5 — `cmd_tick`'s lock scope includes the entire harness subprocess run
The runner holds `_loop_lock` across the long-running
Implement/Verify/Orchestrate harness subprocesses. A concurrent
`--pause-loop` invoked by an operator will block for the WHOLE tick
duration (potentially minutes). The harness is unaware of `_loop_lock`
and cannot signal the operator to wait gracefully.
**Documented behavior** per SPEC D-L1: "ticks short; operator notices
via `--loop-list` stale `last_tick_at`". If ticks grow long, future
work could split the lock into a short gate-decision lock and a longer
state-mutation lock. **Not a bug** — explicit v1.1 scope per
`design/loops/BACKLOG.md` (out of scope for this task).
## Verdict
No BLOCKERS. All observations are documented behaviors per SPEC + D-L6.
Recommend proceeding to adversarial_bug_find.