- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
2.6 KiB
2.6 KiB
Implementation: add-outputs-retention
SCOPE
Add loop.json outputs.retention field (default 20) to bound growth of the outputs/ directory. GC runs after step 10 inside _loop_lock, deleting tick groups older than the retention window. Source: add-loop-runner/BUG_REPORT.md O5.
FILES TOUCHED
scripts/loop-runner.py- Added
_get_retention(cfg) -> int: readscfg.get("outputs", {}).get("retention", 20). Non-int types fall back to 20 with WARNING. Negative values are coerced to 0 (unlimited) with WARNING. - Added
_gc_outputs(loop_path, retention): listsoutputs/, finds max tick index from filenames matching^tick(\d+)-, computescutoff = max_seen - retention + 1, deletes files with tick index < cutoff. Non-tick files (README.txt, etc.) are preserved. Errors logged as WARNING via_append_tick_logand swallowed. - Modified
cmd_tick: calls_get_retention(cfg)+_gc_outputs(loop_path, retention)after step 10 (_write_state_loop) and before step 11 (tick log), inside the_loop_lockblock. - Updated docstring step list: added
10.5. GC outputs/....
- Added
templates/loops/self-improvement/loop.json- Added
"outputs": {"retention": 20}block.
- Added
BUG FOUND AND FIXED INLINE
Off-by-one in GC formula: the initial implementation used cutoff = max_seen - retention, which kept retention + 1 tick groups (21 instead of 20 for retention=20). Fixed to cutoff = max_seen - retention + 1. Test test_gc_keeps_recent_deletes_old caught this (expected 20 kept, got 21 remaining → obvious failure when the remaining-count length check triggered).
DECISIONS LOCKED
- D-O1: retention counts tick GROUPS (all
tick{N}-*files), not individual files. - D-O2: GC runs INSIDE
_loop_lockcritical section (after state write, before tick log). - D-O3: Default 20.
- D-O4: 0 = unlimited (no GC). Negative coerces to 0 with WARNING.
- D-O5: GC based on
outputs/filenames (max_seen), NOTstate.iteration_count. - D-O6: Regex
^tick(\d+)-. Non-matching files preserved. - D-O7: GC failure → WARNING log + swallow.
TESTS
New file tests/test_outputs_retention.py — 13 tests across 2 classes:
TestGetRetention(5): defaultmain, explicit value, negative→0, non-int→20, None cfg→20.TestGcOutputs(8): deletes old keeps recent, retention=0 skip, retention>count, missing dir, non-tick files preserved, unrelatedtick-fooprefix preserved, single tick, error path.
TEST COUNT
- Baseline: 469 passed (post-
harden-parse-verdict). - New: +13 in
tests/test_outputs_retention.py. - Final: 482 passed, 0 regressions.