CI / build (push) Has been cancelled
runnable-test-suite (parent) — complete. Three sub-tasks all complete: - make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships python3), add idempotent .venv install block to scripts/install.sh, and add 'from __future__ import annotations' to 3 dashboard modules using PEP 604 union syntax at definition time so they import on Python 3.9+. The PEP 604 bug was caught by the streak verifier itself during implementation. - vram-detect-cross-platform: scripts/vram_detect.py now branches on platform.system() for Linux/Darwin/Windows. macOS path uses system_profiler SPDisplaysDataType (Apple Silicon unified memory via sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic path win32_VideoController get AdapterRAM with PowerShell fallback. Linux /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5, mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited model cards. Added _probe_ollama_model() that runs 'ollama list' as a last-resort fallback. run_command() now wraps PowerShell cmdlets on Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]). - vram-detect-cross-platform-tests: 11 new monkeypatched tests in tests/test_vram_detect.py covering Linux/Darwin/Windows branches for detect_ram and detect_gpu_vram, prefix-match for unknown model names, ollama probe, Windows PowerShell wrapper, and a LOCKED regression test for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are mocked; no live hardware probes. Suite total: 235 passed, 0 errors. Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory), target context correctly jumped 12k -> 42k. Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak verifier ran as the independent checker model (article #2/#9/#13 in 'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced inline branches where subtask-3 tests expected private _detect_ram_linux() helper; extracted helper to match the test contract without weakening tests. Parent + all 3 subtasks complete. Prior opencode-subagent implementation of subtask-2 preserved in git stash for reference.
126 lines
5.0 KiB
Markdown
126 lines
5.0 KiB
Markdown
# SPEC — vram-detect-cross-platform-tests
|
||
|
||
Parent: `runnable-test-suite` (see PARENT_SPEC.md).
|
||
|
||
**Dependency:** This sub-task runs AFTER `vram-detect-cross-platform` is complete.
|
||
Its function signatures and detection logic are the contract you test against.
|
||
|
||
## Scope
|
||
|
||
pytest tests proving cross-platform branches of `scripts/vram_detect.py` work
|
||
WITHOUT hitting real hardware. Every `subprocess.run` / `Path.exists` / `sysctl`
|
||
call is mocked.
|
||
|
||
## Files this sub-task touches (and ONLY this)
|
||
|
||
- `tests/test_vram_detect.py` — extend or rewrite. Preserve any existing
|
||
passing test (Linux regression tests especially).
|
||
|
||
## MUST NOT touch
|
||
|
||
- `scripts/vram_detect.py` (if it has a bug, escalate back to Orchestrator —
|
||
subtask-2 owns it)
|
||
- Any documentation, prompt, install script, or other source file
|
||
- Any other test file
|
||
|
||
## Test patterns (parametrize + monkeypatch)
|
||
|
||
### 1. `test_detect_ram_linux` (regression — must already be there, keep it)
|
||
- Patches `Path("/proc/meminfo")` text with synthetic `MemTotal: 16384000 kB\nMemAvailable: 8192000 kB`.
|
||
- Asserts `(16_384_000, 8_192_000)` returned.
|
||
|
||
### 2. `test_detect_ram_macos`
|
||
- `monkeypatch.setattr(platform, "system", lambda: "Darwin")`.
|
||
- Patches `subprocess.run` so `sysctl -n hw.memsize` returns `"34359738368"` (32GB).
|
||
- Asserts total_kb correct (33_554_432), available == total.
|
||
|
||
### 3. `test_detect_ram_windows`
|
||
- `monkeypatch.setattr(platform, "system", lambda: "Windows")`.
|
||
- Patches subprocess to return wmic output `TotalPhysicalMemory=34359738368`.
|
||
- Asserts total_kb == 33_554_432.
|
||
|
||
### 4. `test_detect_gpu_vram_nvidia_linux`
|
||
- `monkeypatch.setattr(platform, "system", lambda: "Linux")`.
|
||
- Patches `nvidia-smi` output `"24576\n"`.
|
||
- Asserts `(25_165_824, 25_165_824, 1)` (24GB × 1024 = KB).
|
||
|
||
### 5. `test_detect_gpu_vram_apple_silicon`
|
||
- `monkeypatch.setattr(platform, "system", lambda: "Darwin")`.
|
||
- Patches `system_profiler SPDisplaysDataType` with a real Apple M-series snippet:
|
||
```
|
||
Graphics/Displays:
|
||
Apple M2:
|
||
Chipset Model: Apple M2
|
||
Type: GPU
|
||
Bus: Built-In
|
||
Total Number of Cores: 10
|
||
Vendor: Apple (0x106b)
|
||
Metal: Supported, version 2
|
||
```
|
||
- Patches `sysctl -n hw.memsize` with `"17179869184"` (16GB).
|
||
- Asserts gpu_vram_kb > 0, num_gpus >= 1.
|
||
|
||
### 6. `test_detect_gpu_vram_windows_wmic`
|
||
- `monkeypatch.setattr(platform, "system", lambda: "Windows")`.
|
||
- Patches wmic output:
|
||
```
|
||
AdapterRAM=8589934592
|
||
Name=NVIDIA GeForce RTX 3060
|
||
```
|
||
- Asserts total_vram_kb == 8_388_608 (8GB).
|
||
|
||
### 7. `test_lookup_model_context_unknown_returns_zero`
|
||
- `_lookup_model_context("completely-unknown-model")` returns `0`.
|
||
|
||
### 8. `test_lookup_model_context_prefix_match`
|
||
- `_lookup_model_context("deepseek-r1:7b")` matches `deepseek-r1` entry and returns its context.
|
||
- `_lookup_model_context("llama-3.1-8b-instruct")` matches `llama-3.1-8b`.
|
||
|
||
### 9. `test_detect_model_context_ollama_probe`
|
||
- No config file specifies a model.
|
||
- `ollama list` is on PATH (patch `shutil.which`).
|
||
- Patches subprocess to return:
|
||
```
|
||
NAME ID SIZE MODIFIED
|
||
llama-3.1-8b abc 4.7GB 2 days ago
|
||
```
|
||
- Asserts the returned context matches the `llama-3.1-8b` entry.
|
||
|
||
### 10. `test_run_command_windows_powershell_wrapper`
|
||
- Show that on Windows, a PowerShell cmdlet invocation goes through
|
||
`powershell -NoProfile -NoLogo -Command "..."` (assert argv[0] is powershell
|
||
when cmd is a PS string).
|
||
|
||
### 11. `test_detect_ram_linux_regression` (LOCKED — do not modify)
|
||
- If this test already exists, keep its byte-for-byte assertions. If not, add
|
||
an exact match on `(total_kb, available_kb)` for a specific `/proc/meminfo`
|
||
fixture to prevent subtask-2 from regressing Linux output.
|
||
|
||
## Acceptance criteria
|
||
|
||
1. `python3 -m pytest tests/test_vram_detect.py -v` exits 0 with all new tests passing.
|
||
2. Total `python3 -m pytest tests/ -v` still exits 0.
|
||
3. **Streak:** 10 consecutive clean `python3 -m pytest tests/ -v` runs.
|
||
4. No live `subprocess` against real `system_profiler`/`nvidia-smi`/`wmic` — every
|
||
external call goes through `monkeypatch.setattr`.
|
||
5. Every new test runs < 500ms (mock only, no I/O).
|
||
|
||
## Anti-spin rails
|
||
|
||
- If a subtask-2 function signature doesn't support what a test needs, STOP and
|
||
report. Do not edit `vram_detect.py` yourself and do not weaken the test to fit.
|
||
Escalate via Orchestrator.
|
||
- No `pytest.mark.skip` unless the platform genuinely doesn't support the feature
|
||
(e.g. skip a Windows test on Linux only if it can't be mocked — but mocking is
|
||
the whole point, so this should never happen).
|
||
- Do not add `pytest-cov` or any new deps.
|
||
|
||
## Recommended approach
|
||
|
||
1. Read `scripts/vram_detect.py` and confirm subtask-2's signatures.
|
||
2. For each test, write the fixture data as a module-level constant (multiline string).
|
||
3. Use `monkeypatch.setattr` for `platform.system`, `subprocess.run`, `shutil.which`,
|
||
`Path.exists`, `Path.read_text`. Never call the real thing.
|
||
4. Use `capfd` for stdout assertions where the SPEC calls for print messages.
|
||
5. Run `python3 -m pytest tests/test_vram_detect.py -v` until green.
|
||
6. Run the full suite streak verifier. |