Files
automaton/tasks/runnable-test-suite/subtasks/vram-detect-cross-platform-tests/IMPLEMENTATION.md
T
Lap Tran f32f98575b
CI / build (push) Has been cancelled
Make test suite runnable from clean checkout + cross-platform vram_detect
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00

91 lines
4.7 KiB
Markdown

# IMPLEMENTATION — vram-detect-cross-platform-tests
Parent: `runnable-test-suite` (see PARENT_SPEC.md).
SPEC: `tasks/runnable-test-suite/subtasks/vram-detect-cross-platform-tests/SPEC.md`.
## File touched
- `tests/test_vram_detect.py` — extended with 11 new test functions + helpers.
No other file was modified.
## Test functions added (11)
| # | Function | What it verifies |
|---|----------|------------------|
| 1 | `test_detect_ram_linux` | `/proc/meminfo` parse → `(16_384_000, 8_192_000)` via `detect_ram()` with `platform.system()` patched to `Linux`. |
| 2 | `test_detect_ram_macos` | `sysctl -n hw.memsize` → `"34359738368"` → total_kb `33_554_432`, available == total. |
| 3 | `test_detect_ram_windows` | `wmic ComputerSystem` → `TotalPhysicalMemory=34359738368` → total_kb `33_554_432`. |
| 4 | `test_detect_gpu_vram_nvidia_linux` | `nvidia-smi` → `"24576\n"` → `(25_165_824, 25_165_824, 1)`. |
| 5 | `test_detect_gpu_vram_apple_silicon` | `system_profiler` Apple M2 snippet + `sysctl hw.memsize=17179869184` → total > 0, num >= 1. |
| 6 | `test_detect_gpu_vram_windows_wmic` | `wmic win32_VideoController` → `AdapterRAM=8589934592` → total_vram_kb `8_388_608`. |
| 7 | `test_lookup_model_context_unknown_returns_zero` | `_lookup_model_context("completely-unknown-model")` → `0`. |
| 8 | `test_lookup_model_context_prefix_match` | `deepseek-r1:7b` → `64_000`; `llama-3.1-8b-instruct` → `128_000`. |
| 9 | `test_detect_model_context_ollama_probe` | No config model; `ollama list` returns `llama-3.1-8b` → context `128_000`. |
| 10 | `test_run_command_windows_powershell_wrapper` | On Windows, `Get-CimInstance ...` is routed through `powershell -NoProfile -NoLogo -Command "..."`. |
| 11 | `test_detect_ram_linux_regression` (LOCKED) | Exact match `(32_768_000, 16_384_000)` for a 32GB `/proc/meminfo` fixture via `_detect_ram_linux()`. |
Counting parametrize cases: **11 new tests** (no parametrization used; each function
is a single case).
## Existing tests modified
None. All 9 pre-existing test functions (`test_lookup_model_context`,
`test_parse_token_value`, `test_extract_value`, `test_parse_config_model`,
`test_parse_config_model_skips_code_blocks`, `test_parse_vram_config_manual`,
`test_recommend_context_api_model`, `test_recommend_context_manual_mode`,
`test_extract_model_from_file_respects_10kb_limit`) were preserved byte-for-byte.
No diffs to existing code.
## Mocking strategy
Every external call is mocked via `monkeypatch.setattr` — no live subprocess,
`system_profiler`, `nvidia-smi`, or `wmic` invocation:
- `vram.platform.system` → lambda returning the target OS string.
- `vram.subprocess.run` → `_make_fake_run(responses)` mapping `cmd[0]` (or the
joined PowerShell string) to a canned `_FakeResult(stdout, returncode=0)`.
- `vram.shutil.which` → lambda returning a truthy name (or `None` for non-ollama
in the ollama probe test).
- `Path.exists` / `Path.read_text` → `_patch_meminfo` serves the fixture only
for `/proc/meminfo` and falls through to the original for any other path
(keeps `tmp_path` and pytest internals working during the test).
- `vram.Path.home` → `tmp_path` in the ollama probe test so the real
`~/.automaton/config.md` is never consulted.
Module-level multiline string fixtures: `MEMINFO_LINUX_16GB`,
`MEMINFO_LINUX_32GB`, `APPLE_M2_PROFILER`, `WMIC_VIDEOCONTROLLER`,
`WMIC_COMPUTERSYSTEM`, `OLLAMA_LIST`.
## Acceptance criteria
| Criterion | Result |
|-----------|--------|
| `pytest tests/test_vram_detect.py -v` exits 0, all new tests pass | PASS — 20/20 (9 existing + 11 new) |
| `pytest tests/ -v` exits 0 | PASS — 235 passed, 0 errors |
| Streak: 10 consecutive clean `pytest tests/ -v` runs | PASS — `STREAK_COMPLETE attempt=1 clean=10/10` |
| No live subprocess against real hardware | PASS — every `subprocess.run` / `shutil.which` / `Path` I/O patched |
| Every new test < 500ms | PASS — entire file 0.01s; slowest 60 durations < 0.005s |
| No `pytest.mark.skip` | PASS — none used |
| No new deps | PASS — stdlib + pytest only |
| Did NOT touch `scripts/vram_detect.py` | CONFIRMED — only `tests/test_vram_detect.py` edited |
## Final pass counts
- Baseline (`tests/test_vram_detect.py`): 9 passed.
- After implementation (`tests/test_vram_detect.py`): 20 passed (+11).
- Full suite baseline: 224 passed.
- Full suite after: **235 passed** (+11), 0 errors, 0 skipped.
## Streak result
```
STREAK_COMPLETE attempt=1 clean=10/10
```
## STOP-and-report triggers hit
None. All 11 SPEC-required tests passed against subtask-2's `vram_detect.py`
without any signature mismatch. No edits to `scripts/vram_detect.py` were
required or made. The `ollama list` probe (SPEC test 9) is present in
`vram_detect.py:525-539` and works as specified.