runnable-test-suite (parent) — complete. Three sub-tasks all complete: - make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships python3), add idempotent .venv install block to scripts/install.sh, and add 'from __future__ import annotations' to 3 dashboard modules using PEP 604 union syntax at definition time so they import on Python 3.9+. The PEP 604 bug was caught by the streak verifier itself during implementation. - vram-detect-cross-platform: scripts/vram_detect.py now branches on platform.system() for Linux/Darwin/Windows. macOS path uses system_profiler SPDisplaysDataType (Apple Silicon unified memory via sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic path win32_VideoController get AdapterRAM with PowerShell fallback. Linux /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5, mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited model cards. Added _probe_ollama_model() that runs 'ollama list' as a last-resort fallback. run_command() now wraps PowerShell cmdlets on Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]). - vram-detect-cross-platform-tests: 11 new monkeypatched tests in tests/test_vram_detect.py covering Linux/Darwin/Windows branches for detect_ram and detect_gpu_vram, prefix-match for unknown model names, ollama probe, Windows PowerShell wrapper, and a LOCKED regression test for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are mocked; no live hardware probes. Suite total: 235 passed, 0 errors. Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory), target context correctly jumped 12k -> 42k. Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak verifier ran as the independent checker model (article #2/#9/#13 in 'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced inline branches where subtask-3 tests expected private _detect_ram_linux() helper; extracted helper to match the test contract without weakening tests. Parent + all 3 subtasks complete. Prior opencode-subagent implementation of subtask-2 preserved in git stash for reference.
4.9 KiB
SPEC — runnable-test-suite
Goal
Make python3 -m pytest tests/ -v pass from a clean checkout of ~/.automaton, with deterministic Python deps pinned in the repo and documentation that reflects the actual interpreter that ships on the user's machine.
The prime symptom that proves nothing at all runs today: a fresh clone executes python (per AGENTS.md) and silently fails because macOS only ships python3, and even with the right interpreter the suite fails on No module named pytest.
Requirements (numbered)
- Add
requirements.txtat the repo root pinningpytest(lowest version that supports the syntax used intests/, which is plain fixtures andtmp_path— pytest ≥ 7.0). No other third-party deps may be added. - Provide a venv-based install path: a one-line install in
scripts/install.sh(or a new snippet) that creates.venv/andpip install -r requirements.txt. Must not require sudo and must not pollute the system Python. - Make
python3 -m pytest tests/ -vexit 0 from a clean checkout afterpip install -r requirements.txt(no venv required — systempip3 install -r requirements.txtmust also work). - Fix every Python file under the repo that fails
python3 -m py_compile(currently clean, but must stay clean). - Replace every bare
pythoninvocation in documentation and prompts withpython3so the documented commands actually run on a stock macOS without a shim.AGENTS.mdlines 58, 61, 64, 70README.mdlines 203, 206, 209, 212, 213, 216, 219, 222, 225, 228, 229, 230, 242, 245, 324, 338, 353automaton/dashboard/README.mdlines 11, 18, 21, 24prompts/orchestrate.mdlines 28, 29, 30, 36- Any other
python(bare) reference found byrgAFTER the first pass
- Do NOT change
pythonreferences inside shell scripts that already invoke#!/usr/bin/env python3shebangs or that explicitly resolve viacommand -v. Only fix barepythoncommands that shell out (none expected in scripts/ after audit, but verify). - Add a CI step note to
CHANGELOG.mdunder[unreleased]documenting the newrequirements.txtand thepython3requirement. - Update
AGENTS.md"Build & Test Commands" section to referencerequirements.txtand usepython3consistently.
Acceptance criteria
Each must pass from a fresh clone with only stock macOS CommandLineTools + pip3:
pip3 install -r requirements.txtsucceeds.python3 -m pytest tests/ -vexits 0 withN passed(N ≥ 1) and zeroerrorlines.python3 -m py_compile automaton/**/*.py automaton/dashboard/**/*.py scripts/*.pyexits 0.bash -n scripts/*.shexits 0.rg -n "^python |\"python " AGENTS.md README.md automaton/dashboard/README.md prompts/orchestrate.mdreturns zero matches for a barepythoncommand.- Following the install instructions in
AGENTS.mdverbatim, a new contributor can run the test suite within 60 seconds of clone.
Success contract (streak)
Per the goal mode this task derives from, "done" requires 10 consecutive clean python3 -m pytest tests/ -v runs in a row without any edit between runs. A single failure resets the counter. The cap on attempts is 5; on hitting the cap, stop and report.
Constraints / non-goals
- No new dependencies beyond
pytest. Do not addpytest-cov,pytest-mock,tox, etc. - No virtualenv vendoring. The user creates
.venvthemselves if they want isolation; systempip3 install -r requirements.txtmust also work. - No changes to existing test logic. If a test is genuinely broken (not just import-failing because pytest is missing), STOP and report — do not patch the test to make it pass. That is the anti-spin rule from #9 in the source article.
- Do not touch any file under
tasks/(per-framework tasks are state, not source). - Do not modify
status.py,vram_detect.py, or any other runtime script's behavior. Only documentation and config files change. - No Docker, no conda, no
pyenvrequirements. Stockpython3+pip3only. - VRAM-aware scoping: this task fits in ONE sub-task (~9k peak context budget on this 32GB-RAM / no-GPU machine with
Model: auto). Do not decompose further. Sub-tasks would exceed the budget on overhead alone.
Recommended implementation approach (high-level)
- Create
requirements.txtwithpytest==7.4.4(last 7.x; works on Python 3.9+). pip3 install -r requirements.txtlocally and run the suite; capture every failure.- For each failure, decide: import/install issue (fix dep) vs. real code bug (report, do not patch test).
- Sweep
python→python3in docs/prompts witheditbatching. - Add install snippet to
scripts/install.sh(idempotent; only if.venvdoesn't exist). - Add a one-line test smoke-check at the end of
install.sh:python3 -m pytest tests/ -q || echo "tests deferred". - Update
CHANGELOG.md[unreleased]. - Run the streak verifier: 10×
python3 -m pytest tests/ -v; stop at first clean streak or 5 attempts.