Files
automaton/tasks/runnable-test-suite/SPEC.md
T
Lap Tran f32f98575b
CI / build (push) Has been cancelled
Make test suite runnable from clean checkout + cross-platform vram_detect
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00

59 lines
4.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPEC — runnable-test-suite
## Goal
Make `python3 -m pytest tests/ -v` pass from a clean checkout of `~/.automaton`, with deterministic Python deps pinned in the repo and documentation that reflects the actual interpreter that ships on the user's machine.
The prime symptom that proves nothing at all runs today: a fresh clone executes `python` (per `AGENTS.md`) and silently fails because macOS only ships `python3`, and even with the right interpreter the suite fails on `No module named pytest`.
## Requirements (numbered)
1. Add `requirements.txt` at the repo root pinning `pytest` (lowest version that supports the syntax used in `tests/`, which is plain fixtures and `tmp_path` — pytest ≥ 7.0). No other third-party deps may be added.
2. Provide a venv-based install path: a one-line install in `scripts/install.sh` (or a new snippet) that creates `.venv/` and `pip install -r requirements.txt`. Must not require sudo and must not pollute the system Python.
3. Make `python3 -m pytest tests/ -v` exit 0 from a clean checkout after `pip install -r requirements.txt` (no venv required — system `pip3 install -r requirements.txt` must also work).
4. Fix every Python file under the repo that fails `python3 -m py_compile` (currently clean, but must stay clean).
5. Replace every bare `python ` invocation in documentation and prompts with `python3 ` so the documented commands actually run on a stock macOS without a shim.
- `AGENTS.md` lines 58, 61, 64, 70
- `README.md` lines 203, 206, 209, 212, 213, 216, 219, 222, 225, 228, 229, 230, 242, 245, 324, 338, 353
- `automaton/dashboard/README.md` lines 11, 18, 21, 24
- `prompts/orchestrate.md` lines 28, 29, 30, 36
- Any other `python ` (bare) reference found by `rg` AFTER the first pass
6. Do NOT change `python` references inside shell scripts that already invoke `#!/usr/bin/env python3` shebangs or that explicitly resolve via `command -v`. Only fix bare `python ` commands that shell out (none expected in scripts/ after audit, but verify).
7. Add a CI step note to `CHANGELOG.md` under `[unreleased]` documenting the new `requirements.txt` and the `python3` requirement.
8. Update `AGENTS.md` "Build & Test Commands" section to reference `requirements.txt` and use `python3` consistently.
## Acceptance criteria
Each must pass from a **fresh clone** with only stock macOS CommandLineTools + pip3:
1. `pip3 install -r requirements.txt` succeeds.
2. `python3 -m pytest tests/ -v` exits 0 with `N passed` (N ≥ 1) and zero `error` lines.
3. `python3 -m py_compile automaton/**/*.py automaton/dashboard/**/*.py scripts/*.py` exits 0.
4. `bash -n scripts/*.sh` exits 0.
5. `rg -n "^python |\"python " AGENTS.md README.md automaton/dashboard/README.md prompts/orchestrate.md` returns zero matches for a bare `python ` command.
6. Following the install instructions in `AGENTS.md` verbatim, a new contributor can run the test suite within 60 seconds of clone.
## Success contract (streak)
Per the goal mode this task derives from, "done" requires **10 consecutive clean `python3 -m pytest tests/ -v` runs** in a row without any edit between runs. A single failure resets the counter. The cap on attempts is 5; on hitting the cap, stop and report.
## Constraints / non-goals
- No new dependencies beyond `pytest`. Do not add `pytest-cov`, `pytest-mock`, `tox`, etc.
- No virtualenv vendoring. The user creates `.venv` themselves if they want isolation; system `pip3 install -r requirements.txt` must also work.
- No changes to existing test logic. If a test is genuinely broken (not just import-failing because pytest is missing), STOP and report — do not patch the test to make it pass. That is the anti-spin rule from #9 in the source article.
- Do not touch any file under `tasks/` (per-framework tasks are state, not source).
- Do not modify `status.py`, `vram_detect.py`, or any other runtime script's behavior. Only documentation and config files change.
- No Docker, no conda, no `pyenv` requirements. Stock `python3` + `pip3` only.
- VRAM-aware scoping: this task fits in ONE sub-task (~9k peak context budget on this 32GB-RAM / no-GPU machine with `Model: auto`). **Do not decompose further.** Sub-tasks would exceed the budget on overhead alone.
## Recommended implementation approach (high-level)
1. Create `requirements.txt` with `pytest==7.4.4` (last 7.x; works on Python 3.9+).
2. `pip3 install -r requirements.txt` locally and run the suite; capture every failure.
3. For each failure, decide: import/install issue (fix dep) vs. real code bug (report, do not patch test).
4. Sweep `python ` → `python3 ` in docs/prompts with `edit` batching.
5. Add install snippet to `scripts/install.sh` (idempotent; only if `.venv` doesn't exist).
6. Add a one-line test smoke-check at the end of `install.sh`: `python3 -m pytest tests/ -q || echo "tests deferred"`.
7. Update `CHANGELOG.md` `[unreleased]`.
8. Run the streak verifier: 10× `python3 -m pytest tests/ -v`; stop at first clean streak or 5 attempts.