Files
automaton/tasks/runnable-test-suite/subtasks/make-tests-runnable/PARENT_SPEC.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

60 lines
3.0 KiB
Markdown

# Parent Task: runnable-test-suite
This is the parent SPEC for the runnable-test-suite task. Each sub-task references
this for context, scope boundaries, and the parent acceptance contract.
## Parent Goal
Make `python3 -m pytest tests/ -v` pass from a clean checkout of `~/.automaton`,
with deterministic Python deps pinned and docs that reflect the actual interpreter
on stock macOS/Windows/Linux. In the same wave, make `scripts/vram_detect.py`
cross-platform (macOS, Windows, Linux) — the existing version was developed on
Cachyos and only fully works on Linux.
## Parent Acceptance Contract
1. `pip3 install -r requirements.txt` succeeds on stock macOS CommandLineTools + pip3.
2. `python3 -m pytest tests/ -v` exits 0 from a clean clone, zero `error` lines.
3. **Streak verifier:** 10 consecutive clean `python3 -m pytest tests/ -v` runs with
no edits between runs. A single failure resets the count. Cap: 5 attempts.
4. `python3 -m py_compile automaton/**/*.py automaton/dashboard/**/*.py scripts/*.py` exits 0.
5. `bash -n scripts/*.sh` exits 0.
6. `rg "^python " AGENTS.md README.md automaton/dashboard/README.md prompts/orchestrate.md`
returns zero matches for a bare `python ` command.
7. `python3 scripts/vram_detect.py` on Darwin prints `gpu_vram_gb > 0` (was 0 before).
8. `python3 scripts/vram_detect.py` JSON shape identical to before on Linux/Cachyos.
## Sub-tasks
- `make-tests-runnable` — Wave 1, parallel-ok
- `vram-detect-cross-platform` — Wave 1, parallel-ok
- `vram-detect-cross-platform-tests` — Wave 2, depends on `vram-detect-cross-platform`
Parent is complete ONLY when ALL three sub-tasks pass and the streak verifier above
runs 10 consecutive clean passes.
## Anti-spin rails (from the source article)
- The streak verifier IS the independent checker model from Boris's loop. The
worker (local LLM) does not grade its own homework.
- If a test is genuinely broken (not just import-failing due to missing pytest),
STOP and report. Do not patch the test to make it pass. An agent that grades
itself will delete the failing test and call it done.
- Unknown model name → fail open with `0`. Wrong-context detection is worse than none.
- No new third-party deps beyond `pytest`. Pure stdlib for `vram_detect.py`.
## Hardware/VRAM context
- Detected by `vram_detect.py` on this box: **32GB RAM, no GPU, model unknown**
(because `vram_detect.py` is broken on macOS — subtask-2 fixes that)
- Target context: 12k tokens, headroom 25%, max peak per sub-task: 9k tokens.
- Sub-task peak estimates all fit within 9k. No further decomposition.
## Constraints / non-goals (parent)
- No `psutil`, `wmi`, `pywin32`, `tox`, `pytest-cov`, or other third-party deps.
- No removal of existing OpenAI/Anthropic entries in `MODEL_CONTEXT_WINDOWS`.
- No changes to `status.py`, `autopilot.py`, or any other runtime script's behavior.
- No touching files under `tasks/` (those are state, not source).
- No Docker, no conda, no `pyenv`. Stock `python3` + `pip3` only.
- VRAM detection is additive — Linux/Cachyos output must NOT regress.