Make test suite runnable from clean checkout + cross-platform vram_detect
CI / build (push) Has been cancelled

runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
This commit is contained in:
Lap Tran
2026-06-21 18:28:15 -04:00
parent 1d36c0e4ad
commit f32f98575b
46 changed files with 1499 additions and 84 deletions
@@ -0,0 +1,126 @@
# SPEC — vram-detect-cross-platform-tests
Parent: `runnable-test-suite` (see PARENT_SPEC.md).
**Dependency:** This sub-task runs AFTER `vram-detect-cross-platform` is complete.
Its function signatures and detection logic are the contract you test against.
## Scope
pytest tests proving cross-platform branches of `scripts/vram_detect.py` work
WITHOUT hitting real hardware. Every `subprocess.run` / `Path.exists` / `sysctl`
call is mocked.
## Files this sub-task touches (and ONLY this)
- `tests/test_vram_detect.py` — extend or rewrite. Preserve any existing
passing test (Linux regression tests especially).
## MUST NOT touch
- `scripts/vram_detect.py` (if it has a bug, escalate back to Orchestrator —
subtask-2 owns it)
- Any documentation, prompt, install script, or other source file
- Any other test file
## Test patterns (parametrize + monkeypatch)
### 1. `test_detect_ram_linux` (regression — must already be there, keep it)
- Patches `Path("/proc/meminfo")` text with synthetic `MemTotal: 16384000 kB\nMemAvailable: 8192000 kB`.
- Asserts `(16_384_000, 8_192_000)` returned.
### 2. `test_detect_ram_macos`
- `monkeypatch.setattr(platform, "system", lambda: "Darwin")`.
- Patches `subprocess.run` so `sysctl -n hw.memsize` returns `"34359738368"` (32GB).
- Asserts total_kb correct (33_554_432), available == total.
### 3. `test_detect_ram_windows`
- `monkeypatch.setattr(platform, "system", lambda: "Windows")`.
- Patches subprocess to return wmic output `TotalPhysicalMemory=34359738368`.
- Asserts total_kb == 33_554_432.
### 4. `test_detect_gpu_vram_nvidia_linux`
- `monkeypatch.setattr(platform, "system", lambda: "Linux")`.
- Patches `nvidia-smi` output `"24576\n"`.
- Asserts `(25_165_824, 25_165_824, 1)` (24GB × 1024 = KB).
### 5. `test_detect_gpu_vram_apple_silicon`
- `monkeypatch.setattr(platform, "system", lambda: "Darwin")`.
- Patches `system_profiler SPDisplaysDataType` with a real Apple M-series snippet:
```
Graphics/Displays:
Apple M2:
Chipset Model: Apple M2
Type: GPU
Bus: Built-In
Total Number of Cores: 10
Vendor: Apple (0x106b)
Metal: Supported, version 2
```
- Patches `sysctl -n hw.memsize` with `"17179869184"` (16GB).
- Asserts gpu_vram_kb > 0, num_gpus >= 1.
### 6. `test_detect_gpu_vram_windows_wmic`
- `monkeypatch.setattr(platform, "system", lambda: "Windows")`.
- Patches wmic output:
```
AdapterRAM=8589934592
Name=NVIDIA GeForce RTX 3060
```
- Asserts total_vram_kb == 8_388_608 (8GB).
### 7. `test_lookup_model_context_unknown_returns_zero`
- `_lookup_model_context("completely-unknown-model")` returns `0`.
### 8. `test_lookup_model_context_prefix_match`
- `_lookup_model_context("deepseek-r1:7b")` matches `deepseek-r1` entry and returns its context.
- `_lookup_model_context("llama-3.1-8b-instruct")` matches `llama-3.1-8b`.
### 9. `test_detect_model_context_ollama_probe`
- No config file specifies a model.
- `ollama list` is on PATH (patch `shutil.which`).
- Patches subprocess to return:
```
NAME ID SIZE MODIFIED
llama-3.1-8b abc 4.7GB 2 days ago
```
- Asserts the returned context matches the `llama-3.1-8b` entry.
### 10. `test_run_command_windows_powershell_wrapper`
- Show that on Windows, a PowerShell cmdlet invocation goes through
`powershell -NoProfile -NoLogo -Command "..."` (assert argv[0] is powershell
when cmd is a PS string).
### 11. `test_detect_ram_linux_regression` (LOCKED — do not modify)
- If this test already exists, keep its byte-for-byte assertions. If not, add
an exact match on `(total_kb, available_kb)` for a specific `/proc/meminfo`
fixture to prevent subtask-2 from regressing Linux output.
## Acceptance criteria
1. `python3 -m pytest tests/test_vram_detect.py -v` exits 0 with all new tests passing.
2. Total `python3 -m pytest tests/ -v` still exits 0.
3. **Streak:** 10 consecutive clean `python3 -m pytest tests/ -v` runs.
4. No live `subprocess` against real `system_profiler`/`nvidia-smi`/`wmic` — every
external call goes through `monkeypatch.setattr`.
5. Every new test runs < 500ms (mock only, no I/O).
## Anti-spin rails
- If a subtask-2 function signature doesn't support what a test needs, STOP and
report. Do not edit `vram_detect.py` yourself and do not weaken the test to fit.
Escalate via Orchestrator.
- No `pytest.mark.skip` unless the platform genuinely doesn't support the feature
(e.g. skip a Windows test on Linux only if it can't be mocked — but mocking is
the whole point, so this should never happen).
- Do not add `pytest-cov` or any new deps.
## Recommended approach
1. Read `scripts/vram_detect.py` and confirm subtask-2's signatures.
2. For each test, write the fixture data as a module-level constant (multiline string).
3. Use `monkeypatch.setattr` for `platform.system`, `subprocess.run`, `shutil.which`,
`Path.exists`, `Path.read_text`. Never call the real thing.
4. Use `capfd` for stdout assertions where the SPEC calls for print messages.
5. Run `python3 -m pytest tests/test_vram_detect.py -v` until green.
6. Run the full suite streak verifier.