CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete: - make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships python3), add idempotent .venv install block to scripts/install.sh, and add 'from __future__ import annotations' to 3 dashboard modules using PEP 604 union syntax at definition time so they import on Python 3.9+. The PEP 604 bug was caught by the streak verifier itself during implementation. - vram-detect-cross-platform: scripts/vram_detect.py now branches on platform.system() for Linux/Darwin/Windows. macOS path uses system_profiler SPDisplaysDataType (Apple Silicon unified memory via sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic path win32_VideoController get AdapterRAM with PowerShell fallback. Linux /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5, mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited model cards. Added _probe_ollama_model() that runs 'ollama list' as a last-resort fallback. run_command() now wraps PowerShell cmdlets on Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]). - vram-detect-cross-platform-tests: 11 new monkeypatched tests in tests/test_vram_detect.py covering Linux/Darwin/Windows branches for detect_ram and detect_gpu_vram, prefix-match for unknown model names, ollama probe, Windows PowerShell wrapper, and a LOCKED regression test for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are mocked; no live hardware probes. Suite total: 235 passed, 0 errors. Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory), target context correctly jumped 12k -> 42k. Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak verifier ran as the independent checker model (article #2/#9/#13 in 'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced inline branches where subtask-3 tests expected private _detect_ram_linux() helper; extracted helper to match the test contract without weakening tests. Parent + all 3 subtasks complete. Prior opencode-subagent implementation of subtask-2 preserved in git stash for reference.
60 lines
3.0 KiB
Markdown
60 lines
3.0 KiB
Markdown
# Parent Task: runnable-test-suite
|
|
|
|
This is the parent SPEC for the runnable-test-suite task. Each sub-task references
|
|
this for context, scope boundaries, and the parent acceptance contract.
|
|
|
|
## Parent Goal
|
|
|
|
Make `python3 -m pytest tests/ -v` pass from a clean checkout of `~/.automaton`,
|
|
with deterministic Python deps pinned and docs that reflect the actual interpreter
|
|
on stock macOS/Windows/Linux. In the same wave, make `scripts/vram_detect.py`
|
|
cross-platform (macOS, Windows, Linux) — the existing version was developed on
|
|
Cachyos and only fully works on Linux.
|
|
|
|
## Parent Acceptance Contract
|
|
|
|
1. `pip3 install -r requirements.txt` succeeds on stock macOS CommandLineTools + pip3.
|
|
2. `python3 -m pytest tests/ -v` exits 0 from a clean clone, zero `error` lines.
|
|
3. **Streak verifier:** 10 consecutive clean `python3 -m pytest tests/ -v` runs with
|
|
no edits between runs. A single failure resets the count. Cap: 5 attempts.
|
|
4. `python3 -m py_compile automaton/**/*.py automaton/dashboard/**/*.py scripts/*.py` exits 0.
|
|
5. `bash -n scripts/*.sh` exits 0.
|
|
6. `rg "^python " AGENTS.md README.md automaton/dashboard/README.md prompts/orchestrate.md`
|
|
returns zero matches for a bare `python ` command.
|
|
7. `python3 scripts/vram_detect.py` on Darwin prints `gpu_vram_gb > 0` (was 0 before).
|
|
8. `python3 scripts/vram_detect.py` JSON shape identical to before on Linux/Cachyos.
|
|
|
|
## Sub-tasks
|
|
|
|
- `make-tests-runnable` — Wave 1, parallel-ok
|
|
- `vram-detect-cross-platform` — Wave 1, parallel-ok
|
|
- `vram-detect-cross-platform-tests` — Wave 2, depends on `vram-detect-cross-platform`
|
|
|
|
Parent is complete ONLY when ALL three sub-tasks pass and the streak verifier above
|
|
runs 10 consecutive clean passes.
|
|
|
|
## Anti-spin rails (from the source article)
|
|
|
|
- The streak verifier IS the independent checker model from Boris's loop. The
|
|
worker (local LLM) does not grade its own homework.
|
|
- If a test is genuinely broken (not just import-failing due to missing pytest),
|
|
STOP and report. Do not patch the test to make it pass. An agent that grades
|
|
itself will delete the failing test and call it done.
|
|
- Unknown model name → fail open with `0`. Wrong-context detection is worse than none.
|
|
- No new third-party deps beyond `pytest`. Pure stdlib for `vram_detect.py`.
|
|
|
|
## Hardware/VRAM context
|
|
|
|
- Detected by `vram_detect.py` on this box: **32GB RAM, no GPU, model unknown**
|
|
(because `vram_detect.py` is broken on macOS — subtask-2 fixes that)
|
|
- Target context: 12k tokens, headroom 25%, max peak per sub-task: 9k tokens.
|
|
- Sub-task peak estimates all fit within 9k. No further decomposition.
|
|
|
|
## Constraints / non-goals (parent)
|
|
|
|
- No `psutil`, `wmi`, `pywin32`, `tox`, `pytest-cov`, or other third-party deps.
|
|
- No removal of existing OpenAI/Anthropic entries in `MODEL_CONTEXT_WINDOWS`.
|
|
- No changes to `status.py`, `autopilot.py`, or any other runtime script's behavior.
|
|
- No touching files under `tasks/` (those are state, not source).
|
|
- No Docker, no conda, no `pyenv`. Stock `python3` + `pip3` only.
|
|
- VRAM detection is additive — Linux/Cachyos output must NOT regress. |