CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete: - make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships python3), add idempotent .venv install block to scripts/install.sh, and add 'from __future__ import annotations' to 3 dashboard modules using PEP 604 union syntax at definition time so they import on Python 3.9+. The PEP 604 bug was caught by the streak verifier itself during implementation. - vram-detect-cross-platform: scripts/vram_detect.py now branches on platform.system() for Linux/Darwin/Windows. macOS path uses system_profiler SPDisplaysDataType (Apple Silicon unified memory via sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic path win32_VideoController get AdapterRAM with PowerShell fallback. Linux /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5, mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited model cards. Added _probe_ollama_model() that runs 'ollama list' as a last-resort fallback. run_command() now wraps PowerShell cmdlets on Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]). - vram-detect-cross-platform-tests: 11 new monkeypatched tests in tests/test_vram_detect.py covering Linux/Darwin/Windows branches for detect_ram and detect_gpu_vram, prefix-match for unknown model names, ollama probe, Windows PowerShell wrapper, and a LOCKED regression test for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are mocked; no live hardware probes. Suite total: 235 passed, 0 errors. Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory), target context correctly jumped 12k -> 42k. Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak verifier ran as the independent checker model (article #2/#9/#13 in 'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced inline branches where subtask-3 tests expected private _detect_ram_linux() helper; extracted helper to match the test contract without weakening tests. Parent + all 3 subtasks complete. Prior opencode-subagent implementation of subtask-2 preserved in git stash for reference.
124 lines
5.9 KiB
Markdown
124 lines
5.9 KiB
Markdown
# SPEC — vram-detect-cross-platform
|
|
|
|
Parent: `runnable-test-suite` (see PARENT_SPEC.md).
|
|
|
|
## Scope
|
|
|
|
Make `scripts/vram_detect.py` work on macOS, Windows, and Linux without behavior
|
|
change on Cachyos/Linux. The existing version was developed on Cachyos and only
|
|
fully works on Linux. Pure detection logic — no CLI/JSON-shape changes.
|
|
|
|
## Files this sub-task touches (and ONLY this)
|
|
|
|
- `scripts/vram_detect.py` — surgical edits to detection functions only.
|
|
|
|
## MUST NOT touch
|
|
|
|
- `tests/test_vram_detect.py` (subtask-3 owns all vram_detect tests)
|
|
- Any other test file
|
|
- Any documentation, prompt, or install script (subtask-1 owns those)
|
|
- CLI args, JSON output shape, main() flow — only detection internals
|
|
|
|
## Functions to refactor (by current line in `scripts/vram_detect.py`)
|
|
|
|
### `run_command()` (vram_detect.py:54-70)
|
|
- On Windows, route PowerShell cmdlets via `powershell -NoProfile -NoLogo -Command "..."` wrapper.
|
|
- Keep `shutil.which()` gating so missing tools return None on all OSes.
|
|
- Do not break Linux path.
|
|
|
|
### `detect_gpu_vram()` (vram_detect.py:73-108)
|
|
Branch on `platform.system()`:
|
|
- **Linux**: keep `nvidia-smi` → `lspci -vnn` path exactly as-is. Regression guard.
|
|
- **Darwin** (macOS): add `system_profiler SPDisplaysDataType` parser.
|
|
Parse `VRAM (Total):` line for Intel Macs, and for Apple Silicon unified memory,
|
|
detect `Chipset Model: Apple M*` and treat total RAM as shared VRAM (call
|
|
`sysctl -n hw.memsize` once and use that number, since Apple Silicon has no
|
|
dedicated VRAM). Print clearly which kind was detected.
|
|
Fallback (if `system_profiler` missing): `ioreg -c IOPlatformDevice -r -d 1`.
|
|
- **Windows**: add `wmic path win32_VideoController get AdapterRAM,Name /format:list`
|
|
(deprecated but ubiquitous; works on Win10/11). Sum `AdapterRAM=` values across
|
|
GPUs. PowerShell fallback:
|
|
`powershell -NoProfile -Command "Get-CimInstance Win32_VideoController | Select-Object AdapterRAM"`
|
|
|
|
### `detect_ram()` (vram_detect.py:137-162)
|
|
Branch on `platform.system()`:
|
|
- **Linux**: keep `/proc/meminfo` path. Parse `MemTotal` + `MemAvailable`. Regression guard.
|
|
- **Darwin**: keep `sysctl -n hw.memsize` path (currently the fallback; promote to
|
|
the macOS branch as primary). Returns (total_kb, total_kb) because macOS doesn't
|
|
expose "available RAM" via sysctl directly — leave available == total. Print clear
|
|
"available RAM detection not supported on macOS, reporting total" message once.
|
|
- **Windows**: add `wmic ComputerSystem get TotalPhysicalMemory /format:list`.
|
|
Returns bytes — divide by 1024 for KB. PowerShell fallback:
|
|
`powershell -NoProfile -Command "(Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory"`
|
|
|
|
### `MODEL_CONTEXT_WINDOWS` dict (vram_detect.py:24-47)
|
|
Additive only. Add these local-LLM entries with context sizes from public model cards:
|
|
- `llama-3.1-8b`: 128_000
|
|
- `llama-3.3-70b`: 128_000
|
|
- `qwen2.5-7b`: 128_000 (Qwen2.5 supports up to 128k per model card)
|
|
- `qwen2.5-72b`: 128_000
|
|
- `mistral-7b`: 32_000
|
|
- `mistral-large`: 128_000
|
|
- `deepseek-r1`: 64_000 (DeepSeek-R1)
|
|
- `deepseek-v3`: 64_000
|
|
- `glm-4`: 128_000
|
|
- `glm-4.5`: 128_000
|
|
- `gemma-2`: 8_000
|
|
- `gemma-2-27b`: 8_000
|
|
- `phi-3`: 128_000
|
|
- `phi-4`: 16_000
|
|
|
|
Add a docstring comment above each: `# Source: <model card URL or repo>` — never
|
|
fabricate. If a number is uncertain, use the smaller conservative value and
|
|
leave a comment noting the uncertainty.
|
|
|
|
### `detect_model_context()` (vram_detect.py:175-228)
|
|
Add an `ollama list` probe when no config file names a model AND the system has
|
|
`ollama` on PATH. Steps:
|
|
1. `ollama list` → parse first non-header row's NAME column (strip `:latest` tag).
|
|
2. Look up the cleaned name in `MODEL_CONTEXT_WINDOWS` via existing `_lookup_model_context`.
|
|
3. If matched, return that context. If not matched, fall through to fail-open `0`.
|
|
|
|
Do NOT modify any other code path in `detect_model_context`.
|
|
|
|
## MUST NOT regress
|
|
|
|
- The existing Linux output of `python3 vram_detect.py` must produce byte-identical
|
|
stdout (after the equivalent hardware probe) on the original Cachyos box. Subtask-3
|
|
will write a Linux-fixture test to lock this in.
|
|
- Do not remove or alter any existing OpenAI/Anthropic entry in `MODEL_CONTEXT_WINDOWS`.
|
|
|
|
## Acceptance criteria
|
|
|
|
1. `python3 scripts/vram_detect.py` on Darwin prints `gpu_vram_gb > 0` (currently prints 0).
|
|
2. `python3 scripts/vram_detect.py` JSON contains the same keys, same order, same types.
|
|
3. `python3 -m py_compile scripts/vram_detect.py` exits 0.
|
|
4. On a Linux fixture (simulated by subtask-3 tests with patched `platform.system`),
|
|
stdout matches the pre-refactor output line-by-line for GPU/RAM sections.
|
|
5. `python3 scripts/vram_detect.py` does not crash on Windows stub (subtask-3 sets
|
|
`monkeypatch.setattr(platform, "system", lambda: "Windows")` and mocks subprocess).
|
|
6. No new third-party imports.
|
|
|
|
## Anti-spin rails
|
|
|
|
- Unknown model name → fail open with `0`. NEVER guess a context window.
|
|
- If `platform.system()` returns an unexpected string (e.g. "AIX"), fall through to
|
|
Linux path or print "Unsupported OS: X" and return 0s. Do not crash.
|
|
- If `system_profiler` output format on the local M-series Mac is different from what
|
|
you parsed, STOP and report. Don't patch a half-working parser.
|
|
|
|
## Hardware context (this box)
|
|
|
|
- Darwin arm64, Python 3.9.6 (stock CommandLineTools).
|
|
- `system_profiler SPDisplaysDataType` is the canonical probe.
|
|
- You can iterate locally by running `python3 scripts/vram_detect.py` after each edit.
|
|
|
|
## Recommended approach
|
|
|
|
1. Add the local-LLM entries to `MODEL_CONTEXT_WINDOWS` first (mechanical).
|
|
2. Refactor `detect_ram()` with a `platform.system()` dispatch — easiest, lowest risk.
|
|
3. Refactor `detect_gpu_vram()` — hardest, leave for after RAM is green.
|
|
4. Add the `ollama list` probe.
|
|
5. Wrapping `run_command()` for Windows PowerShell — defer until last.
|
|
6. After each function, run `python3 scripts/vram_detect.py` and confirm no crash + correct output.
|
|
7. Do NOT touch `tests/test_vram_detect.py`. Subtask-3 will write tests against your function signatures. |