Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

140 lines
8.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPEC: add-outputs-retention
## Problem
`tasks/add-loop-runner/BUG_REPORT.md` O5:
> Every tick writes 3 files (implement, verify, orchestrate). Over 100 ticks that's 300 files. Trees on some filesystems (HFS+, ext4 default) degrade past 10k entries per dir. v1 has `max_iterations` to bound this; for daemon mode with `max_iterations=0`, the user is responsible.
Actual count is **6 files per tick** (each role: a `tick{N}-<role>-prompt.md` written by `_resolve_prompt`, plus a `tick{N}-<role>.json` written by `cmd_tick`). With `max_iterations=0` (daemon, unbounded), the `outputs/` directory grows without bound.
## Goal
Bound `outputs/` directory growth by retaining only the **last N tick groups**. A "tick group" = all files with the `tick{N}-` prefix for a single tick index N. Older tick groups are garbage-collected on every tick.
## Non-goals
- Per-role retention (e.g. keep verify-outputs longer than implement-outputs). Out of scope; would complicate the schema.
- Compression / archival of old tick dirs to a tarball. Out of scope.
- Cross-loop retention. Each loop's `outputs/` is independent.
- Retention of `.state.log` (tick log). That file is append-only and grows linearly; separate concern.
## Schema addition (`loop.json`)
Add an optional `outputs` object:
```json
"outputs": {
"retention": 20
}
```
- **`outputs.retention`** (int, optional, default **20**): keep the last N tick groups. Older tick groups are deleted on every tick. `0` = unlimited (no GC; v1 behavior). Negative values are rejected at `--create-loop`.
## Requirements
### R1 — retention config plumbing
- `status.py --create-loop` accepts `outputs.retention` in the `loop.json` template.
- The runner reads `cfg.get("outputs", {}).get("retention", 20)`.
- Validation on read: if `retention` is < 0, log WARNING and treat as `0` (unlimited). Non-int types coerce via `int(...)`; on `TypeError`/`ValueError` fall back to default `20`.
### R2 — GC executes on every tick (post-write)
- After step 10 (`_write_state_loop`) and before step 11 (tick log), the runner invokes `_gc_outputs(loop_path, state, retention)`.
- GC iterates `outputs/` directory, parses `tickNN-` prefixes, computes the cutoff = `iteration_count - retention + 1` (kept range: `[cutoff, iteration_count]` inclusive).
- Any file whose tick-index prefix is `< cutoff` is deleted. Files without a `tickN-` prefix are left alone (forward-compat; user may place other files in `outputs/`).
- GC errors (file in use, permission) are logged via `_append_tick_log` WARNING and swallowed — GC failure must not crash the tick.
### R3 — Retention = 0 means no GC
- `0` skips the GC step entirely (cheapest path for `max_iterations` users who prefer manual cleanup).
### R4 — Atomicity / failure isolation
- GC failures (permission, file not found mid-iteration) don't roll back the tick. State has already advanced; losing a GC pass is benign (next tick re-attempts).
- Missing `outputs/` (loop never ticked) — GC no-ops, no error.
### R5 — No new pip deps
- Pure stdlib: `os.listdir`, `os.remove`, `re.match`. No `shutil.rmtree` (we delete individual files; a tick group is not a directory).
## Detailed semantics
### Tick-index extraction
Filenames follow the pattern `tick<int>-<remainder>` where `<int>` is the 1-based tick index. Examples:
- `tick1-implement.json`, `tick1-verify.json`, `tick1-orchestrate.json`, `tick1-implement-prompt.md`, `tick1-verify-prompt.md`, `tick1-orchestrate-prompt.md`
Regex: `^tick(\d+)-`. Tick indices are extracted into a set, the maximum tick index (`max_seen`) is computed, and the cutoff floor is `max_seen - retention + 1`. Files with tick index `< floor` get deleted.
**Why `max_seen - retention + 1` instead of `state.iteration_count`?**
State could lag (e.g. concurrent ticks), but the on-disk filenames ARE ground truth. Using max filename keeps GC self-contained.
### Default retention choice
Default = **20**. Rationale:
- Score-plateau window default is often 5-10; keeping 2x that covers debugging.
- 20 ticks × 6 files = 120 files max — comfortably under any filesystem degradation threshold.
- Operators who need longer history (`audit` use cases) override upward in `loop.json`.
### Where GC runs in the tick flow
```
... step 10: _write_state_loop(state)
# NEW: step 10.5
_gc_outputs(loop_path, state, retention)
# step 11
_append_tick_log(...)
```
GC runs INSIDE the `_loop_lock` critical section, so a concurrent `--pause-loop` / `--approve --loop` can't be mid-write and observe a missing tick dir. GC's filesystem delete ops are independent of `.state.loop`.
## Test plan
Pure-function + filesystem tests (no subprocess, no live LLM):
1. **GC deletes old tick groups, keeps recent N**: write 30 tick groups (6 files each), retention=20, expect last 20 kept, oldest 10 deleted, all 6 files per kept tick are present.
2. **Retention = 0 skips GC entirely**: 30 tick groups, retention=0, expect no deletion, all files present.
3. **Retention > file count** (no-op): 5 tick groups, retention=20, expect no deletion.
4. **Missing `outputs/` dir** (no-op, no error): fresh loop, no `outputs/`, GC returns cleanly.
5. **Non-tick files in `outputs/` are preserved**: write 30 tick groups + a `README.txt` and `loop-info.md`, retention=20, expect tick groups deleted but `README.txt` and `loop-info.md` intact.
6. **Negative retention coerces to 0 (no GC)**: retention=-5 in `loop.json`, expect WARNING + no deletion.
7. **Non-int retention coerces to default 20**: retention="twenty", expect WARNING + default 20 used (deletes oldest 10 of 30).
8. **Tick-index regex preserves unrelated `tick-foo` files** (defensive): `tick-foo.md` (no number) does NOT match `^tick(\d+)-`; expect preserved.
9. **GC error swallowed (permission-denied file)**: chmod 000 a stale tick file (or use a non-existent mock that raises `PermissionError`); expect GC logs WARNING and continues; tick proceeds.
10. **Concurrent with state write** (lock interaction): GC runs inside the lock; no separate test needed (the `test_state_loop_lock.py` suite already covers lock integrity).
11. **Config plumbing**: `--create-loop` writes `outputs.retention: 20` into generated `loop.json` (if `--outputs-retention` not provided; or honors override).
12. **Default getter**: `_get_retention(cfg)` returns 20 for missing `outputs`, 0 when `{"outputs": {"retention": 0}}`, 20 for `{"outputs": {"retention": "garbage"}}` (post-WARNING).
## Decisions (locked)
- **D-O1**: retention counts tick GROUPS not individual files. A tick group = all `tick{N}-*` files. Keeps the mental model aligned with "ticks as the atomic unit".
- **D-O2**: GC runs INSIDE `_loop_lock` critical section (after state write, before tick log). Cheapest correct placement — no separate lock, no concurrent `--pause-loop` / `--approve --loop` mid-GC race. Filesystem delete ops are independent of `.state.loop` but the lock keeps the loop's externally-observable state consistent.
- **D-O3**: Default 20 (covers debugging; 120 files max comfortably under fs degradation).
- **D-O4**: `0` = unlimited (no GC). Negative coerces to 0 with WARNING.
- **D-O5**: GC based on `outputs/` filenames (`max_seen`), NOT `state.iteration_count`. Self-contained; robust to state lag.
- **D-O6**: Regex `^tick(\d+)-`. Files not matching are preserved (forward-compat for helper docs, scratch notes, etc.).
- **D-O7**: GC failure (PermissionError, FileNotFoundError mid-iteration) → WARNING log + swallow. Tick not affected.
## Out of scope (filed BACKLOG.md)
- `outputs.retention_bytes` (磁盘 budget cap). Future.
- Tarball archival of GC'd tick groups. Future.
- Cross-loop retention aggregation. Future.
- GC `.state.log` rotation. Separate task (`add-state-log-rotation`).
## Files touched
- `scripts/loop-runner.py` — add `_get_retention(cfg)` + `_gc_outputs(loop_path, state, retention)`; call after step 10 inside `_loop_lock`.
- `scripts/status.py` — `--create-loop` writes `outputs.retention` default 20 into generated `loop.json` template; validates non-negative.
- `templates/loops/self-improvement/loop.json` — add `"outputs": {"retention": 20}` to template.
- `design/loops/technical.md` — new subsection §7b "Outputs retention (v1.1 — `add-outputs-retention`)".
- `design/loops/functional.md` — note `outputs.retention` field in the schema enum.
- `CHANGELOG.md` — new entry under `[unreleased]`.
- `tests/test_outputs_retention.py` (NEW) — 12 tests per plan above.
## Pipeline plan
research → research:awaiting_approval → research:approved → implement → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete.