# SPEC — vram-detect-cross-platform Parent: `runnable-test-suite` (see PARENT_SPEC.md). ## Scope Make `scripts/vram_detect.py` work on macOS, Windows, and Linux without behavior change on Cachyos/Linux. The existing version was developed on Cachyos and only fully works on Linux. Pure detection logic — no CLI/JSON-shape changes. ## Files this sub-task touches (and ONLY this) - `scripts/vram_detect.py` — surgical edits to detection functions only. ## MUST NOT touch - `tests/test_vram_detect.py` (subtask-3 owns all vram_detect tests) - Any other test file - Any documentation, prompt, or install script (subtask-1 owns those) - CLI args, JSON output shape, main() flow — only detection internals ## Functions to refactor (by current line in `scripts/vram_detect.py`) ### `run_command()` (vram_detect.py:54-70) - On Windows, route PowerShell cmdlets via `powershell -NoProfile -NoLogo -Command "..."` wrapper. - Keep `shutil.which()` gating so missing tools return None on all OSes. - Do not break Linux path. ### `detect_gpu_vram()` (vram_detect.py:73-108) Branch on `platform.system()`: - **Linux**: keep `nvidia-smi` → `lspci -vnn` path exactly as-is. Regression guard. - **Darwin** (macOS): add `system_profiler SPDisplaysDataType` parser. Parse `VRAM (Total):` line for Intel Macs, and for Apple Silicon unified memory, detect `Chipset Model: Apple M*` and treat total RAM as shared VRAM (call `sysctl -n hw.memsize` once and use that number, since Apple Silicon has no dedicated VRAM). Print clearly which kind was detected. Fallback (if `system_profiler` missing): `ioreg -c IOPlatformDevice -r -d 1`. - **Windows**: add `wmic path win32_VideoController get AdapterRAM,Name /format:list` (deprecated but ubiquitous; works on Win10/11). Sum `AdapterRAM=` values across GPUs. PowerShell fallback: `powershell -NoProfile -Command "Get-CimInstance Win32_VideoController | Select-Object AdapterRAM"` ### `detect_ram()` (vram_detect.py:137-162) Branch on `platform.system()`: - **Linux**: keep `/proc/meminfo` path. Parse `MemTotal` + `MemAvailable`. Regression guard. - **Darwin**: keep `sysctl -n hw.memsize` path (currently the fallback; promote to the macOS branch as primary). Returns (total_kb, total_kb) because macOS doesn't expose "available RAM" via sysctl directly — leave available == total. Print clear "available RAM detection not supported on macOS, reporting total" message once. - **Windows**: add `wmic ComputerSystem get TotalPhysicalMemory /format:list`. Returns bytes — divide by 1024 for KB. PowerShell fallback: `powershell -NoProfile -Command "(Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory"` ### `MODEL_CONTEXT_WINDOWS` dict (vram_detect.py:24-47) Additive only. Add these local-LLM entries with context sizes from public model cards: - `llama-3.1-8b`: 128_000 - `llama-3.3-70b`: 128_000 - `qwen2.5-7b`: 128_000 (Qwen2.5 supports up to 128k per model card) - `qwen2.5-72b`: 128_000 - `mistral-7b`: 32_000 - `mistral-large`: 128_000 - `deepseek-r1`: 64_000 (DeepSeek-R1) - `deepseek-v3`: 64_000 - `glm-4`: 128_000 - `glm-4.5`: 128_000 - `gemma-2`: 8_000 - `gemma-2-27b`: 8_000 - `phi-3`: 128_000 - `phi-4`: 16_000 Add a docstring comment above each: `# Source: ` — never fabricate. If a number is uncertain, use the smaller conservative value and leave a comment noting the uncertainty. ### `detect_model_context()` (vram_detect.py:175-228) Add an `ollama list` probe when no config file names a model AND the system has `ollama` on PATH. Steps: 1. `ollama list` → parse first non-header row's NAME column (strip `:latest` tag). 2. Look up the cleaned name in `MODEL_CONTEXT_WINDOWS` via existing `_lookup_model_context`. 3. If matched, return that context. If not matched, fall through to fail-open `0`. Do NOT modify any other code path in `detect_model_context`. ## MUST NOT regress - The existing Linux output of `python3 vram_detect.py` must produce byte-identical stdout (after the equivalent hardware probe) on the original Cachyos box. Subtask-3 will write a Linux-fixture test to lock this in. - Do not remove or alter any existing OpenAI/Anthropic entry in `MODEL_CONTEXT_WINDOWS`. ## Acceptance criteria 1. `python3 scripts/vram_detect.py` on Darwin prints `gpu_vram_gb > 0` (currently prints 0). 2. `python3 scripts/vram_detect.py` JSON contains the same keys, same order, same types. 3. `python3 -m py_compile scripts/vram_detect.py` exits 0. 4. On a Linux fixture (simulated by subtask-3 tests with patched `platform.system`), stdout matches the pre-refactor output line-by-line for GPU/RAM sections. 5. `python3 scripts/vram_detect.py` does not crash on Windows stub (subtask-3 sets `monkeypatch.setattr(platform, "system", lambda: "Windows")` and mocks subprocess). 6. No new third-party imports. ## Anti-spin rails - Unknown model name → fail open with `0`. NEVER guess a context window. - If `platform.system()` returns an unexpected string (e.g. "AIX"), fall through to Linux path or print "Unsupported OS: X" and return 0s. Do not crash. - If `system_profiler` output format on the local M-series Mac is different from what you parsed, STOP and report. Don't patch a half-working parser. ## Hardware context (this box) - Darwin arm64, Python 3.9.6 (stock CommandLineTools). - `system_profiler SPDisplaysDataType` is the canonical probe. - You can iterate locally by running `python3 scripts/vram_detect.py` after each edit. ## Recommended approach 1. Add the local-LLM entries to `MODEL_CONTEXT_WINDOWS` first (mechanical). 2. Refactor `detect_ram()` with a `platform.system()` dispatch — easiest, lowest risk. 3. Refactor `detect_gpu_vram()` — hardest, leave for after RAM is green. 4. Add the `ollama list` probe. 5. Wrapping `run_command()` for Windows PowerShell — defer until last. 6. After each function, run `python3 scripts/vram_detect.py` and confirm no crash + correct output. 7. Do NOT touch `tests/test_vram_detect.py`. Subtask-3 will write tests against your function signatures.