Files
automaton/tasks/complete/runnable-test-suite/subtasks/vram-detect-cross-platform/SPEC.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

5.9 KiB

SPEC — vram-detect-cross-platform

Parent: runnable-test-suite (see PARENT_SPEC.md).

Scope

Make scripts/vram_detect.py work on macOS, Windows, and Linux without behavior change on Cachyos/Linux. The existing version was developed on Cachyos and only fully works on Linux. Pure detection logic — no CLI/JSON-shape changes.

Files this sub-task touches (and ONLY this)

  • scripts/vram_detect.py — surgical edits to detection functions only.

MUST NOT touch

  • tests/test_vram_detect.py (subtask-3 owns all vram_detect tests)
  • Any other test file
  • Any documentation, prompt, or install script (subtask-1 owns those)
  • CLI args, JSON output shape, main() flow — only detection internals

Functions to refactor (by current line in scripts/vram_detect.py)

run_command() (vram_detect.py:54-70)

  • On Windows, route PowerShell cmdlets via powershell -NoProfile -NoLogo -Command "..." wrapper.
  • Keep shutil.which() gating so missing tools return None on all OSes.
  • Do not break Linux path.

detect_gpu_vram() (vram_detect.py:73-108)

Branch on platform.system():

  • Linux: keep nvidia-smi → lspci -vnn path exactly as-is. Regression guard.
  • Darwin (macOS): add system_profiler SPDisplaysDataType parser. Parse VRAM (Total): line for Intel Macs, and for Apple Silicon unified memory, detect Chipset Model: Apple M* and treat total RAM as shared VRAM (call sysctl -n hw.memsize once and use that number, since Apple Silicon has no dedicated VRAM). Print clearly which kind was detected. Fallback (if system_profiler missing): ioreg -c IOPlatformDevice -r -d 1.
  • Windows: add wmic path win32_VideoController get AdapterRAM,Name /format:list (deprecated but ubiquitous; works on Win10/11). Sum AdapterRAM= values across GPUs. PowerShell fallback: powershell -NoProfile -Command "Get-CimInstance Win32_VideoController | Select-Object AdapterRAM"

detect_ram() (vram_detect.py:137-162)

Branch on platform.system():

  • Linux: keep /proc/meminfo path. Parse MemTotal + MemAvailable. Regression guard.
  • Darwin: keep sysctl -n hw.memsize path (currently the fallback; promote to the macOS branch as primary). Returns (total_kb, total_kb) because macOS doesn't expose "available RAM" via sysctl directly — leave available == total. Print clear "available RAM detection not supported on macOS, reporting total" message once.
  • Windows: add wmic ComputerSystem get TotalPhysicalMemory /format:list. Returns bytes — divide by 1024 for KB. PowerShell fallback: powershell -NoProfile -Command "(Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory"

MODEL_CONTEXT_WINDOWS dict (vram_detect.py:24-47)

Additive only. Add these local-LLM entries with context sizes from public model cards:

  • llama-3.1-8b: 128_000
  • llama-3.3-70b: 128_000
  • qwen2.5-7b: 128_000 (Qwen2.5 supports up to 128k per model card)
  • qwen2.5-72b: 128_000
  • mistral-7b: 32_000
  • mistral-large: 128_000
  • deepseek-r1: 64_000 (DeepSeek-R1)
  • deepseek-v3: 64_000
  • glm-4: 128_000
  • glm-4.5: 128_000
  • gemma-2: 8_000
  • gemma-2-27b: 8_000
  • phi-3: 128_000
  • phi-4: 16_000

Add a docstring comment above each: # Source: <model card URL or repo> — never fabricate. If a number is uncertain, use the smaller conservative value and leave a comment noting the uncertainty.

detect_model_context() (vram_detect.py:175-228)

Add an ollama list probe when no config file names a model AND the system has ollama on PATH. Steps:

  1. ollama list → parse first non-header row's NAME column (strip :latest tag).
  2. Look up the cleaned name in MODEL_CONTEXT_WINDOWS via existing _lookup_model_context.
  3. If matched, return that context. If not matched, fall through to fail-open 0.

Do NOT modify any other code path in detect_model_context.

MUST NOT regress

  • The existing Linux output of python3 vram_detect.py must produce byte-identical stdout (after the equivalent hardware probe) on the original Cachyos box. Subtask-3 will write a Linux-fixture test to lock this in.
  • Do not remove or alter any existing OpenAI/Anthropic entry in MODEL_CONTEXT_WINDOWS.

Acceptance criteria

  1. python3 scripts/vram_detect.py on Darwin prints gpu_vram_gb > 0 (currently prints 0).
  2. python3 scripts/vram_detect.py JSON contains the same keys, same order, same types.
  3. python3 -m py_compile scripts/vram_detect.py exits 0.
  4. On a Linux fixture (simulated by subtask-3 tests with patched platform.system), stdout matches the pre-refactor output line-by-line for GPU/RAM sections.
  5. python3 scripts/vram_detect.py does not crash on Windows stub (subtask-3 sets monkeypatch.setattr(platform, "system", lambda: "Windows") and mocks subprocess).
  6. No new third-party imports.

Anti-spin rails

  • Unknown model name → fail open with 0. NEVER guess a context window.
  • If platform.system() returns an unexpected string (e.g. "AIX"), fall through to Linux path or print "Unsupported OS: X" and return 0s. Do not crash.
  • If system_profiler output format on the local M-series Mac is different from what you parsed, STOP and report. Don't patch a half-working parser.

Hardware context (this box)

  • Darwin arm64, Python 3.9.6 (stock CommandLineTools).
  • system_profiler SPDisplaysDataType is the canonical probe.
  • You can iterate locally by running python3 scripts/vram_detect.py after each edit.
  1. Add the local-LLM entries to MODEL_CONTEXT_WINDOWS first (mechanical).
  2. Refactor detect_ram() with a platform.system() dispatch — easiest, lowest risk.
  3. Refactor detect_gpu_vram() — hardest, leave for after RAM is green.
  4. Add the ollama list probe.
  5. Wrapping run_command() for Windows PowerShell — defer until last.
  6. After each function, run python3 scripts/vram_detect.py and confirm no crash + correct output.
  7. Do NOT touch tests/test_vram_detect.py. Subtask-3 will write tests against your function signatures.