- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
3.7 KiB
BUG_REPORT: add-loop-runner
Probed the runner against the v1 loop-death modes and harness-substitution edge cases.
Bugs found
None blocking. Informational observations below.
Observations (non-blocking)
O1 — --loop argument typo produces a SKIP untracked (silent)
If the user invokes loop-runner.py --loop typo-name, the runner logs SKIP untracked and exits 0. The OS scheduler will keep firing the same bad loop name forever. Mitigation: --check-gate and status.py already refuse unknown loops with exit 2 — but only if invoked by humans. The runner's own --loop typo is silent. Worth a WARNING log line to .state.log? No — there is no .state.log for untracked loops; nothing to write to. Accepted. Fix: don't typo your loop name. No code change.
O2 — Verdict-output file is written even on parse failure
If the verifier subprocess returns garbage, cmd_tick still writes the garbage to <loop>/outputs/tickN-verify.json before halting. A user scanning the outputs dir sees garbage files. Harmless but messy. Fix in v1.1: gate the file-write behind a successful parse. Not blocking.
O3 — Daemon mode logs no DAEMON_TICK entries between ticks
cmd_daemon calls cmd_tick which logs TICK pass=…. But the daemon itself only logs on KeyboardInterrupt. If the user wants to see "daemon has looped N times" the existing TICK log entries suffice. Accepted.
O4 — _context_floor_ok returns True if vram_detect.py subprocess fails
Best-effort choice: a missing/broken vram_detect.py (e.g. on a fresh CI container without the script installed) is treated as "eligible". Correct for portability (the framework shouldn't hard-refuse a tick on a platform where the tool isn't built), but means the 16k floor (D13) can be silently bypassed on misconfigured hosts. Trade-off accepted; documented in the function's docstring. If a user wants strict enforcement, they install vram_detect.py. No code change.
O5 — No upper bound on outputs/ directory growth
Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has max_iterations to bound this; for daemon mode with max_iterations=0, the user is responsible. v1.1 hardening: add outputs.retention to loop.json (keep last N ticks). Logged to BACKLOG.
O6 — parse_verdict accepts {pass: "true"} (string) as truthy
verdict["pass"] = bool(data.get("pass")) — bool("true") is True but bool("false") is also True (non-empty string). A verifier that returns {"pass": "false", "score": 0.1} will be recorded as pass=True. Verifier prompts (task 6) must instruct the model to emit JSON booleans. Minor robustness fix here: check for string and normalize. Let me note this for task 6 prompt work, but also harden in v1 — parse_verdict should coerce "true"/"false" strings. I'll leave it for v1.1 since the verifier prompt (task 6) is the actual contract; the prompt will tell the model to emit true/false as JSON booleans, not strings. Not blocking for task 3.
Five loop-death modes — runtime coverage
| Death | Defense | In runner? |
|---|---|---|
| drift | _gate_worktree_drift (status.py) |
via --check-gate |
| runaway | _gate_iterations (status.py) |
via --check-gate |
| bad verifier | _gate_score_plateau (status.py) + parse_verdict |
via --check-gate + direct |
| resource burn | _gate_budget (status.py) |
via --check-gate |
| undetected halt | R8 transition refusal (status.py) + audit Cat-6 | via --check-gate not-ok path |
Verdict
PASS — no blocker bugs. O5 filed to BACKLOG; O6 noted for task 6 prompt work; others are accepted trade-offs or out of scope.