Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

8.5 KiB
Raw Permalink Blame History

SPEC: add-outputs-retention

Problem

tasks/add-loop-runner/BUG_REPORT.md O5:

Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has max_iterations to bound this; for daemon mode with max_iterations=0, the user is responsible.

Actual count is 6 files per tick (each role: a tick{N}-<role>-prompt.md written by _resolve_prompt, plus a tick{N}-<role>.json written by cmd_tick). With max_iterations=0 (daemon, unbounded), the outputs/ directory grows without bound.

Goal

Bound outputs/ directory growth by retaining only the last N tick groups. A "tick group" = all files with the tick{N}- prefix for a single tick index N. Older tick groups are garbage-collected on every tick.

Non-goals

  • Per-role retention (e.g. keep verify-outputs longer than implement-outputs). Out of scope; would complicate the schema.
  • Compression / archival of old tick dirs to a tarball. Out of scope.
  • Cross-loop retention. Each loop's outputs/ is independent.
  • Retention of .state.log (tick log). That file is append-only and grows linearly; separate concern.

Schema addition (loop.json)

Add an optional outputs object:

"outputs": {
  "retention": 20
}
  • outputs.retention (int, optional, default 20): keep the last N tick groups. Older tick groups are deleted on every tick. 0 = unlimited (no GC; v1 behavior). Negative values are rejected at --create-loop.

Requirements

R1 — retention config plumbing

  • status.py --create-loop accepts outputs.retention in the loop.json template.
  • The runner reads cfg.get("outputs", {}).get("retention", 20).
  • Validation on read: if retention is < 0, log WARNING and treat as 0 (unlimited). Non-int types coerce via int(...); on TypeError/ValueError fall back to default 20.

R2 — GC executes on every tick (post-write)

  • After step 10 (_write_state_loop) and before step 11 (tick log), the runner invokes _gc_outputs(loop_path, state, retention).
  • GC iterates outputs/ directory, parses tickNN- prefixes, computes the cutoff = iteration_count - retention + 1 (kept range: [cutoff, iteration_count] inclusive).
  • Any file whose tick-index prefix is < cutoff is deleted. Files without a tickN- prefix are left alone (forward-compat; user may place other files in outputs/).
  • GC errors (file in use, permission) are logged via _append_tick_log WARNING and swallowed — GC failure must not crash the tick.

R3 — Retention = 0 means no GC

  • 0 skips the GC step entirely (cheapest path for max_iterations users who prefer manual cleanup).

R4 — Atomicity / failure isolation

  • GC failures (permission, file not found mid-iteration) don't roll back the tick. State has already advanced; losing a GC pass is benign (next tick re-attempts).
  • Missing outputs/ (loop never ticked) — GC no-ops, no error.

R5 — No new pip deps

  • Pure stdlib: os.listdir, os.remove, re.match. No shutil.rmtree (we delete individual files; a tick group is not a directory).

Detailed semantics

Tick-index extraction

Filenames follow the pattern tick<int>-<remainder> where <int> is the 1-based tick index. Examples:

  • tick1-implement.json, tick1-verify.json, tick1-orchestrate.json, tick1-implement-prompt.md, tick1-verify-prompt.md, tick1-orchestrate-prompt.md

Regex: ^tick(\d+)-. Tick indices are extracted into a set, the maximum tick index (max_seen) is computed, and the cutoff floor is max_seen - retention + 1. Files with tick index < floor get deleted.

Why max_seen - retention + 1 instead of state.iteration_count?

State could lag (e.g. concurrent ticks), but the on-disk filenames ARE ground truth. Using max filename keeps GC self-contained.

Default retention choice

Default = 20. Rationale:

  • Score-plateau window default is often 5-10; keeping 2x that covers debugging.
  • 20 ticks × 6 files = 120 files max — comfortably under any filesystem degradation threshold.
  • Operators who need longer history (audit use cases) override upward in loop.json.

Where GC runs in the tick flow

... step 10: _write_state_loop(state)
  # NEW: step 10.5
  _gc_outputs(loop_path, state, retention)
  # step 11
  _append_tick_log(...)

GC runs INSIDE the _loop_lock critical section, so a concurrent --pause-loop / --approve --loop can't be mid-write and observe a missing tick dir. GC's filesystem delete ops are independent of .state.loop.

Test plan

Pure-function + filesystem tests (no subprocess, no live LLM):

  1. GC deletes old tick groups, keeps recent N: write 30 tick groups (6 files each), retention=20, expect last 20 kept, oldest 10 deleted, all 6 files per kept tick are present.
  2. Retention = 0 skips GC entirely: 30 tick groups, retention=0, expect no deletion, all files present.
  3. Retention > file count (no-op): 5 tick groups, retention=20, expect no deletion.
  4. Missing outputs/ dir (no-op, no error): fresh loop, no outputs/, GC returns cleanly.
  5. Non-tick files in outputs/ are preserved: write 30 tick groups + a README.txt and loop-info.md, retention=20, expect tick groups deleted but README.txt and loop-info.md intact.
  6. Negative retention coerces to 0 (no GC): retention=-5 in loop.json, expect WARNING + no deletion.
  7. Non-int retention coerces to default 20: retention="twenty", expect WARNING + default 20 used (deletes oldest 10 of 30).
  8. Tick-index regex preserves unrelated tick-foo files (defensive): tick-foo.md (no number) does NOT match ^tick(\d+)-; expect preserved.
  9. GC error swallowed (permission-denied file): chmod 000 a stale tick file (or use a non-existent mock that raises PermissionError); expect GC logs WARNING and continues; tick proceeds.
  10. Concurrent with state write (lock interaction): GC runs inside the lock; no separate test needed (the test_state_loop_lock.py suite already covers lock integrity).
  11. Config plumbing: --create-loop writes outputs.retention: 20 into generated loop.json (if --outputs-retention not provided; or honors override).
  12. Default getter: _get_retention(cfg) returns 20 for missing outputs, 0 when {"outputs": {"retention": 0}}, 20 for {"outputs": {"retention": "garbage"}} (post-WARNING).

Decisions (locked)

  • D-O1: retention counts tick GROUPS not individual files. A tick group = all tick{N}-* files. Keeps the mental model aligned with "ticks as the atomic unit".
  • D-O2: GC runs INSIDE _loop_lock critical section (after state write, before tick log). Cheapest correct placement — no separate lock, no concurrent --pause-loop / --approve --loop mid-GC race. Filesystem delete ops are independent of .state.loop but the lock keeps the loop's externally-observable state consistent.
  • D-O3: Default 20 (covers debugging; 120 files max comfortably under fs degradation).
  • D-O4: 0 = unlimited (no GC). Negative coerces to 0 with WARNING.
  • D-O5: GC based on outputs/ filenames (max_seen), NOT state.iteration_count. Self-contained; robust to state lag.
  • D-O6: Regex ^tick(\d+)-. Files not matching are preserved (forward-compat for helper docs, scratch notes, etc.).
  • D-O7: GC failure (PermissionError, FileNotFoundError mid-iteration) → WARNING log + swallow. Tick not affected.

Out of scope (filed BACKLOG.md)

  • outputs.retention_bytes (磁盘 budget cap). Future.
  • Tarball archival of GC'd tick groups. Future.
  • Cross-loop retention aggregation. Future.
  • GC .state.log rotation. Separate task (add-state-log-rotation).

Files touched

  • scripts/loop-runner.py — add _get_retention(cfg) + _gc_outputs(loop_path, state, retention); call after step 10 inside _loop_lock.
  • scripts/status.py — --create-loop writes outputs.retention default 20 into generated loop.json template; validates non-negative.
  • templates/loops/self-improvement/loop.json — add "outputs": {"retention": 20} to template.
  • design/loops/technical.md — new subsection §7b "Outputs retention (v1.1 — add-outputs-retention)".
  • design/loops/functional.md — note outputs.retention field in the schema enum.
  • CHANGELOG.md — new entry under [unreleased].
  • tests/test_outputs_retention.py (NEW) — 12 tests per plan above.

Pipeline plan

research → research:awaiting_approval → research:approved → implement → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete.