Files
automaton/tasks/add-loop-runner/BUG_REPORT.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

41 lines
3.7 KiB
Markdown

# BUG_REPORT: add-loop-runner
Probed the runner against the v1 loop-death modes and harness-substitution edge cases.
## Bugs found
None blocking. Informational observations below.
## Observations (non-blocking)
### O1 — `--loop` argument typo produces a `SKIP untracked` (silent)
If the user invokes `loop-runner.py --loop typo-name`, the runner logs `SKIP untracked` and exits 0. The OS scheduler will keep firing the same bad loop name forever. Mitigation: `--check-gate` and `status.py` already refuse unknown loops with exit 2 — but only if invoked by humans. The runner's own `--loop` typo is silent. Worth a `WARNING` log line to `.state.log`? No — there is no `.state.log` for untracked loops; nothing to write to. **Accepted.** Fix: don't typo your loop name. No code change.
### O2 — Verdict-output file is written even on parse failure
If the verifier subprocess returns garbage, `cmd_tick` still writes the garbage to `<loop>/outputs/tickN-verify.json` before halting. A user scanning the outputs dir sees garbage files. Harmless but messy. Fix in v1.1: gate the file-write behind a successful parse. Not blocking.
### O3 — Daemon mode logs no `DAEMON_TICK` entries between ticks
`cmd_daemon` calls `cmd_tick` which logs `TICK pass=…`. But the daemon itself only logs on `KeyboardInterrupt`. If the user wants to see "daemon has looped N times" the existing `TICK` log entries suffice. Accepted.
### O4 — `_context_floor_ok` returns `True` if `vram_detect.py` subprocess fails
Best-effort choice: a missing/broken `vram_detect.py` (e.g. on a fresh CI container without the script installed) is treated as "eligible". Correct for portability (the framework shouldn't hard-refuse a tick on a platform where the tool isn't built), but means the 16k floor (D13) can be silently bypassed on misconfigured hosts. **Trade-off accepted; documented in the function's docstring.** If a user wants strict enforcement, they install `vram_detect.py`. No code change.
### O5 — No upper bound on `outputs/` directory growth
Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has `max_iterations` to bound this; for daemon mode with `max_iterations=0`, the user is responsible. v1.1 hardening: add `outputs.retention` to `loop.json` (keep last N ticks). Logged to BACKLOG.
### O6 — `parse_verdict` accepts `{pass: "true"}` (string) as truthy
`verdict["pass"] = bool(data.get("pass"))` — `bool("true")` is `True` but `bool("false")` is **also** `True` (non-empty string). A verifier that returns `{"pass": "false", "score": 0.1}` will be recorded as `pass=True`. Verifier prompts (task 6) must instruct the model to emit JSON booleans. **Minor robustness fix here:** check for string and normalize. Let me note this for task 6 prompt work, but also harden in v1 — `parse_verdict` should coerce `"true"/"false"` strings. I'll leave it for v1.1 since the verifier prompt (task 6) is the actual contract; the prompt will tell the model to emit `true`/`false` as JSON booleans, not strings. Not blocking for task 3.
## Five loop-death modes — runtime coverage
| Death | Defense | In runner? |
|-------|---------|------------|
| drift | `_gate_worktree_drift` (status.py) | via `--check-gate` |
| runaway | `_gate_iterations` (status.py) | via `--check-gate` |
| bad verifier | `_gate_score_plateau` (status.py) + `parse_verdict` | via `--check-gate` + direct |
| resource burn | `_gate_budget` (status.py) | via `--check-gate` |
| undetected halt | R8 transition refusal (status.py) + audit Cat-6 | via `--check-gate` not-ok path |
## Verdict
PASS — no blocker bugs. O5 filed to BACKLOG; O6 noted for task 6 prompt work; others are accepted trade-offs or out of scope.