Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

3.7 KiB

BUG_REPORT: add-loop-runner

Probed the runner against the v1 loop-death modes and harness-substitution edge cases.

Bugs found

None blocking. Informational observations below.

Observations (non-blocking)

O1 — --loop argument typo produces a SKIP untracked (silent)

If the user invokes loop-runner.py --loop typo-name, the runner logs SKIP untracked and exits 0. The OS scheduler will keep firing the same bad loop name forever. Mitigation: --check-gate and status.py already refuse unknown loops with exit 2 — but only if invoked by humans. The runner's own --loop typo is silent. Worth a WARNING log line to .state.log? No — there is no .state.log for untracked loops; nothing to write to. Accepted. Fix: don't typo your loop name. No code change.

O2 — Verdict-output file is written even on parse failure

If the verifier subprocess returns garbage, cmd_tick still writes the garbage to <loop>/outputs/tickN-verify.json before halting. A user scanning the outputs dir sees garbage files. Harmless but messy. Fix in v1.1: gate the file-write behind a successful parse. Not blocking.

O3 — Daemon mode logs no DAEMON_TICK entries between ticks

cmd_daemon calls cmd_tick which logs TICK pass=…. But the daemon itself only logs on KeyboardInterrupt. If the user wants to see "daemon has looped N times" the existing TICK log entries suffice. Accepted.

O4 — _context_floor_ok returns True if vram_detect.py subprocess fails

Best-effort choice: a missing/broken vram_detect.py (e.g. on a fresh CI container without the script installed) is treated as "eligible". Correct for portability (the framework shouldn't hard-refuse a tick on a platform where the tool isn't built), but means the 16k floor (D13) can be silently bypassed on misconfigured hosts. Trade-off accepted; documented in the function's docstring. If a user wants strict enforcement, they install vram_detect.py. No code change.

O5 — No upper bound on outputs/ directory growth

Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has max_iterations to bound this; for daemon mode with max_iterations=0, the user is responsible. v1.1 hardening: add outputs.retention to loop.json (keep last N ticks). Logged to BACKLOG.

O6 — parse_verdict accepts {pass: "true"} (string) as truthy

verdict["pass"] = bool(data.get("pass")) — bool("true") is True but bool("false") is also True (non-empty string). A verifier that returns {"pass": "false", "score": 0.1} will be recorded as pass=True. Verifier prompts (task 6) must instruct the model to emit JSON booleans. Minor robustness fix here: check for string and normalize. Let me note this for task 6 prompt work, but also harden in v1 — parse_verdict should coerce "true"/"false" strings. I'll leave it for v1.1 since the verifier prompt (task 6) is the actual contract; the prompt will tell the model to emit true/false as JSON booleans, not strings. Not blocking for task 3.

Five loop-death modes — runtime coverage

Death Defense In runner?
drift _gate_worktree_drift (status.py) via --check-gate
runaway _gate_iterations (status.py) via --check-gate
bad verifier _gate_score_plateau (status.py) + parse_verdict via --check-gate + direct
resource burn _gate_budget (status.py) via --check-gate
undetected halt R8 transition refusal (status.py) + audit Cat-6 via --check-gate not-ok path

Verdict

PASS — no blocker bugs. O5 filed to BACKLOG; O6 noted for task 6 prompt work; others are accepted trade-offs or out of scope.