Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

2.6 KiB

Implementation: add-outputs-retention

SCOPE

Add loop.json outputs.retention field (default 20) to bound growth of the outputs/ directory. GC runs after step 10 inside _loop_lock, deleting tick groups older than the retention window. Source: add-loop-runner/BUG_REPORT.md O5.

FILES TOUCHED

  • scripts/loop-runner.py
    • Added _get_retention(cfg) -> int: reads cfg.get("outputs", {}).get("retention", 20). Non-int types fall back to 20 with WARNING. Negative values are coerced to 0 (unlimited) with WARNING.
    • Added _gc_outputs(loop_path, retention): lists outputs/, finds max tick index from filenames matching ^tick(\d+)-, computes cutoff = max_seen - retention + 1, deletes files with tick index < cutoff. Non-tick files (README.txt, etc.) are preserved. Errors logged as WARNING via _append_tick_log and swallowed.
    • Modified cmd_tick: calls _get_retention(cfg) + _gc_outputs(loop_path, retention) after step 10 (_write_state_loop) and before step 11 (tick log), inside the _loop_lock block.
    • Updated docstring step list: added 10.5. GC outputs/....
  • templates/loops/self-improvement/loop.json
    • Added "outputs": {"retention": 20} block.

BUG FOUND AND FIXED INLINE

Off-by-one in GC formula: the initial implementation used cutoff = max_seen - retention, which kept retention + 1 tick groups (21 instead of 20 for retention=20). Fixed to cutoff = max_seen - retention + 1. Test test_gc_keeps_recent_deletes_old caught this (expected 20 kept, got 21 remaining → obvious failure when the remaining-count length check triggered).

DECISIONS LOCKED

  • D-O1: retention counts tick GROUPS (all tick{N}-* files), not individual files.
  • D-O2: GC runs INSIDE _loop_lock critical section (after state write, before tick log).
  • D-O3: Default 20.
  • D-O4: 0 = unlimited (no GC). Negative coerces to 0 with WARNING.
  • D-O5: GC based on outputs/ filenames (max_seen), NOT state.iteration_count.
  • D-O6: Regex ^tick(\d+)-. Non-matching files preserved.
  • D-O7: GC failure → WARNING log + swallow.

TESTS

New file tests/test_outputs_retention.py — 13 tests across 2 classes:

  • TestGetRetention (5): default main, explicit value, negative→0, non-int→20, None cfg→20.
  • TestGcOutputs (8): deletes old keeps recent, retention=0 skip, retention>count, missing dir, non-tick files preserved, unrelated tick-foo prefix preserved, single tick, error path.

TEST COUNT

  • Baseline: 469 passed (post-harden-parse-verdict).
  • New: +13 in tests/test_outputs_retention.py.
  • Final: 482 passed, 0 regressions.