- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
8.5 KiB
SPEC: add-outputs-retention
Problem
tasks/add-loop-runner/BUG_REPORT.md O5:
Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has
max_iterationsto bound this; for daemon mode withmax_iterations=0, the user is responsible.
Actual count is 6 files per tick (each role: a tick{N}-<role>-prompt.md written by _resolve_prompt, plus a tick{N}-<role>.json written by cmd_tick). With max_iterations=0 (daemon, unbounded), the outputs/ directory grows without bound.
Goal
Bound outputs/ directory growth by retaining only the last N tick groups. A "tick group" = all files with the tick{N}- prefix for a single tick index N. Older tick groups are garbage-collected on every tick.
Non-goals
- Per-role retention (e.g. keep verify-outputs longer than implement-outputs). Out of scope; would complicate the schema.
- Compression / archival of old tick dirs to a tarball. Out of scope.
- Cross-loop retention. Each loop's
outputs/is independent. - Retention of
.state.log(tick log). That file is append-only and grows linearly; separate concern.
Schema addition (loop.json)
Add an optional outputs object:
"outputs": {
"retention": 20
}
outputs.retention(int, optional, default 20): keep the last N tick groups. Older tick groups are deleted on every tick.0= unlimited (no GC; v1 behavior). Negative values are rejected at--create-loop.
Requirements
R1 — retention config plumbing
status.py --create-loopacceptsoutputs.retentionin theloop.jsontemplate.- The runner reads
cfg.get("outputs", {}).get("retention", 20). - Validation on read: if
retentionis < 0, log WARNING and treat as0(unlimited). Non-int types coerce viaint(...); onTypeError/ValueErrorfall back to default20.
R2 — GC executes on every tick (post-write)
- After step 10 (
_write_state_loop) and before step 11 (tick log), the runner invokes_gc_outputs(loop_path, state, retention). - GC iterates
outputs/directory, parsestickNN-prefixes, computes the cutoff =iteration_count - retention + 1(kept range:[cutoff, iteration_count]inclusive). - Any file whose tick-index prefix is
< cutoffis deleted. Files without atickN-prefix are left alone (forward-compat; user may place other files inoutputs/). - GC errors (file in use, permission) are logged via
_append_tick_logWARNING and swallowed — GC failure must not crash the tick.
R3 — Retention = 0 means no GC
0skips the GC step entirely (cheapest path formax_iterationsusers who prefer manual cleanup).
R4 — Atomicity / failure isolation
- GC failures (permission, file not found mid-iteration) don't roll back the tick. State has already advanced; losing a GC pass is benign (next tick re-attempts).
- Missing
outputs/(loop never ticked) — GC no-ops, no error.
R5 — No new pip deps
- Pure stdlib:
os.listdir,os.remove,re.match. Noshutil.rmtree(we delete individual files; a tick group is not a directory).
Detailed semantics
Tick-index extraction
Filenames follow the pattern tick<int>-<remainder> where <int> is the 1-based tick index. Examples:
tick1-implement.json,tick1-verify.json,tick1-orchestrate.json,tick1-implement-prompt.md,tick1-verify-prompt.md,tick1-orchestrate-prompt.md
Regex: ^tick(\d+)-. Tick indices are extracted into a set, the maximum tick index (max_seen) is computed, and the cutoff floor is max_seen - retention + 1. Files with tick index < floor get deleted.
Why max_seen - retention + 1 instead of state.iteration_count?
State could lag (e.g. concurrent ticks), but the on-disk filenames ARE ground truth. Using max filename keeps GC self-contained.
Default retention choice
Default = 20. Rationale:
- Score-plateau window default is often 5-10; keeping 2x that covers debugging.
- 20 ticks × 6 files = 120 files max — comfortably under any filesystem degradation threshold.
- Operators who need longer history (
audituse cases) override upward inloop.json.
Where GC runs in the tick flow
... step 10: _write_state_loop(state)
# NEW: step 10.5
_gc_outputs(loop_path, state, retention)
# step 11
_append_tick_log(...)
GC runs INSIDE the _loop_lock critical section, so a concurrent --pause-loop / --approve --loop can't be mid-write and observe a missing tick dir. GC's filesystem delete ops are independent of .state.loop.
Test plan
Pure-function + filesystem tests (no subprocess, no live LLM):
- GC deletes old tick groups, keeps recent N: write 30 tick groups (6 files each), retention=20, expect last 20 kept, oldest 10 deleted, all 6 files per kept tick are present.
- Retention = 0 skips GC entirely: 30 tick groups, retention=0, expect no deletion, all files present.
- Retention > file count (no-op): 5 tick groups, retention=20, expect no deletion.
- Missing
outputs/dir (no-op, no error): fresh loop, nooutputs/, GC returns cleanly. - Non-tick files in
outputs/are preserved: write 30 tick groups + aREADME.txtandloop-info.md, retention=20, expect tick groups deleted butREADME.txtandloop-info.mdintact. - Negative retention coerces to 0 (no GC): retention=-5 in
loop.json, expect WARNING + no deletion. - Non-int retention coerces to default 20: retention="twenty", expect WARNING + default 20 used (deletes oldest 10 of 30).
- Tick-index regex preserves unrelated
tick-foofiles (defensive):tick-foo.md(no number) does NOT match^tick(\d+)-; expect preserved. - GC error swallowed (permission-denied file): chmod 000 a stale tick file (or use a non-existent mock that raises
PermissionError); expect GC logs WARNING and continues; tick proceeds. - Concurrent with state write (lock interaction): GC runs inside the lock; no separate test needed (the
test_state_loop_lock.pysuite already covers lock integrity). - Config plumbing:
--create-loopwritesoutputs.retention: 20into generatedloop.json(if--outputs-retentionnot provided; or honors override). - Default getter:
_get_retention(cfg)returns 20 for missingoutputs, 0 when{"outputs": {"retention": 0}}, 20 for{"outputs": {"retention": "garbage"}}(post-WARNING).
Decisions (locked)
- D-O1: retention counts tick GROUPS not individual files. A tick group = all
tick{N}-*files. Keeps the mental model aligned with "ticks as the atomic unit". - D-O2: GC runs INSIDE
_loop_lockcritical section (after state write, before tick log). Cheapest correct placement — no separate lock, no concurrent--pause-loop/--approve --loopmid-GC race. Filesystem delete ops are independent of.state.loopbut the lock keeps the loop's externally-observable state consistent. - D-O3: Default 20 (covers debugging; 120 files max comfortably under fs degradation).
- D-O4:
0= unlimited (no GC). Negative coerces to 0 with WARNING. - D-O5: GC based on
outputs/filenames (max_seen), NOTstate.iteration_count. Self-contained; robust to state lag. - D-O6: Regex
^tick(\d+)-. Files not matching are preserved (forward-compat for helper docs, scratch notes, etc.). - D-O7: GC failure (PermissionError, FileNotFoundError mid-iteration) → WARNING log + swallow. Tick not affected.
Out of scope (filed BACKLOG.md)
outputs.retention_bytes(磁盘 budget cap). Future.- Tarball archival of GC'd tick groups. Future.
- Cross-loop retention aggregation. Future.
- GC
.state.logrotation. Separate task (add-state-log-rotation).
Files touched
scripts/loop-runner.py— add_get_retention(cfg)+_gc_outputs(loop_path, state, retention); call after step 10 inside_loop_lock.scripts/status.py—--create-loopwritesoutputs.retentiondefault 20 into generatedloop.jsontemplate; validates non-negative.templates/loops/self-improvement/loop.json— add"outputs": {"retention": 20}to template.design/loops/technical.md— new subsection §7b "Outputs retention (v1.1 —add-outputs-retention)".design/loops/functional.md— noteoutputs.retentionfield in the schema enum.CHANGELOG.md— new entry under[unreleased].tests/test_outputs_retention.py(NEW) — 12 tests per plan above.
Pipeline plan
research → research:awaiting_approval → research:approved → implement → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete.