Files
automaton/tasks/runnable-test-suite/DECOMPOSITION.md
T
Lap Tran f32f98575b
CI / build (push) Has been cancelled
Make test suite runnable from clean checkout + cross-platform vram_detect
runnable-test-suite (parent) — complete. Three sub-tasks all complete:

- make-tests-runnable: add requirements.txt pinning pytest==7.4.4, sweep all
  docs/prompts from bare 'python' to 'python3' (stock macOS/Windows ships
  python3), add idempotent .venv install block to scripts/install.sh, and add
  'from __future__ import annotations' to 3 dashboard modules using PEP 604
  union syntax at definition time so they import on Python 3.9+. The
  PEP 604 bug was caught by the streak verifier itself during implementation.

- vram-detect-cross-platform: scripts/vram_detect.py now branches on
  platform.system() for Linux/Darwin/Windows. macOS path uses
  system_profiler SPDisplaysDataType (Apple Silicon unified memory via
  sysctl hw.memsize; Intel Macs via 'VRAM (Total):'). Windows uses wmic
  path win32_VideoController get AdapterRAM with PowerShell fallback. Linux
  /proc/meminfo and nvidia-smi/lspci paths unchanged (regression test locks
  them). Added 14 local-LLM context-window entries (llama-3.1, qwen2.5,
  mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited
  model cards. Added _probe_ollama_model() that runs 'ollama list' as a
  last-resort fallback. run_command() now wraps PowerShell cmdlets on
  Windows (['powershell', '-NoProfile', '-NoLogo', '-Command', ...]).

- vram-detect-cross-platform-tests: 11 new monkeypatched tests in
  tests/test_vram_detect.py covering Linux/Darwin/Windows branches for
  detect_ram and detect_gpu_vram, prefix-match for unknown model names,
  ollama probe, Windows PowerShell wrapper, and a LOCKED regression test
  for _detect_ram_linux(). All external subprocess/sysctl/wmic calls are
  mocked; no live hardware probes. Suite total: 235 passed, 0 errors.

Verified on this box: gpu_vram_gb 0 -> 32 on Apple M5 (32GB unified memory),
target context correctly jumped 12k -> 42k.

Subtask-2 implementation was authored by local LLM (gemma-4-26B-A4B-it
via headroom proxy @ localhost:8787). The 10-consecutive-clean-pass streak
verifier ran as the independent checker model (article #2/#9/#13 in
'WTF Is a Loop? Part 2'). One anti-spin rail fired: local LLM produced
inline branches where subtask-3 tests expected private _detect_ram_linux()
helper; extracted helper to match the test contract without weakening tests.

Parent + all 3 subtasks complete. Prior opencode-subagent implementation
of subtask-2 preserved in git stash for reference.
2026-06-21 18:28:15 -04:00

6.0 KiB
Raw Blame History

DECOMPOSITION — runnable-test-suite

Method

Decompose by capability boundary, not by file. Each sub-task is independently verifiable and independently mergeable. Local-LLM context budget per sub-task: max 9k tokens peak on this box (32GB RAM, no GPU, Model: auto).

Sub-tasks (3)

subtask-1: make-tests-runnable

Scope: pytest install path + python → python3 doc sweep + streak verifier. Files touched: requirements.txt (new), AGENTS.md, README.md, automaton/dashboard/README.md, prompts/orchestrate.md, scripts/install.sh (append venv snippet, idempotent), CHANGELOG.md. Not touched: vram_detect.py, status.py, any test logic. Acceptance: pip3 install -r requirements.txt && python3 -m pytest tests/ -v exits 0 from a clean clone; 10 consecutive clean streak; rg "^python " returns zero matches in docs/prompts. Peak context estimate: ~4k tokens (mostly mechanical doc edits). Fits easily. Run order: first. Establishes the green-test baseline the other sub-tasks need.

subtask-2: vram-detect-cross-platform

Scope: Make scripts/vram_detect.py work on macOS, Windows, Linux without behavior change on Cachyos/Linux. Pure detection logic — no CLI/JSON-shape changes. Files touched: scripts/vram_detect.py only. Functions to refactor (by line in current file):

  • detect_gpu_vram() (vram_detect.py:73-108): branch on platform.system().
    • Linux: keep nvidia-smi → lspci -vnn path.
    • macOS: add system_profiler SPDisplaysDataType → parse VRAM (Total) and Chipset Vendor (Apple Unified Memory counts as VRAM). Probe ioreg -c IOPlatformDevice only if system_profiler is unavailable.
    • Windows: add wmic path win32_VideoController get AdapterRAM,Name (deprecated but ubiquitous); fallback to PowerShell Get-CimInstance Win32_VideoController -Property AdapterRAM. Sum across GPUs.
  • detect_ram() (vram_detect.py:137-162): branch on platform.system().
    • Linux: keep /proc/meminfo.
    • macOS: keep sysctl -n hw.memsize (already works as fallback).
    • Windows: add wmic ComputerSystem get TotalPhysicalMemory; PowerShell fallback (Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory.
  • MODEL_CONTEXT_WINDOWS (vram_detect.py:24-47): add local-LLM entries: llama-3.1-8b, llama-3.3-70b, qwen2.5-7b, qwen2.5-72b, mistral-7b, mistral-large, deepseek-r1, deepseek-v3, glm-4, glm-4.5, gemma-2, gemma-2-27b, phi-3, phi-4. Use community-published context sizes. No fabricating — every entry must cite the source model card in a comment.
  • detect_model_context() (vram_detect.py:175-228): add ollama list probe when no config file specifies a model. Pick the first running model name and look it up.
  • run_command() (vram_detect.py:54-70): on Windows, route PowerShell cmdlets via powershell -NoProfile -Command "..." wrapper. Keep shutil.which gating. Not touched: test files (those are subtask-3), JSON output shape, CLI args. Acceptance on this machine (Darwin): python3 vram_detect.py prints macOS VRAM (non-zero on Apple Silicon), RAM 32GB, recommends ≥8k target. JSON has gpu_vram_gb > 0. On Linux (CI), output unchanged from current. Anti-spin rail: if a new entry in MODEL_CONTEXT_WINDOWS is unknown, fail open with 0, do NOT guess. Per SPEC, wrong-context detection is worse than none. Peak context estimate: ~7k tokens (single 536-line file, surgical edits). Fits. Run order: second, parallel-ok with subtask-1 (independent files).

subtask-3: vram-detect-cross-platform-tests

Scope: pytest tests proving cross-platform branches without hitting real hardware. Files touched: tests/test_vram_detect.py only. Patterns:

  • Parametrize detect_ram across Linux/Darwin/Windows with monkeypatch on platform.system, Path.exists, subprocess.run, and Path.read_text; assert correct KB returned and correct print lines emitted (capfd).
  • Parametrize detect_gpu_vram with mocked system_profiler / wmic / nvidia-smi stdout fixtures (kept as multiline string constants).
  • Assert unknown model name returns 0 (fail-open contract from subtask-2).
  • Assert _lookup_model_context picks the longest matching prefix (so llama-3.1-8b-instruct matches llama-3.1-8b).
  • Assert existing Linux/Cachyos path still parses /proc/meminfo (regression).
  • No live subprocess against real system_profiler/nvidia-smi — every call goes through monkeypatch.setattr. Not touched: vram_detect.py itself, any other source file, any prompt. Acceptance: added tests pass; total suite still 10-streak clean. Peak context estimate: ~5k tokens. Fits. Run order: third, AFTER subtask-2 (depends on its function signatures).

Dependency graph

subtask-1 ─┐
            ├─> parent done
subtask-2 ─┤
            └─> subtask-3 ──> parent done

Parent runnable-test-suite is complete only when ALL three sub-tasks pass the streak verifier from the SPEC (10 consecutive clean python3 -m pytest tests/ -v).

Parent non-goals

  • No usage of psutil, wmi, pywin32, or other new third-party deps. Pure stdlib (platform, subprocess, shutil, re, sys). Per SPEC constraint #1.
  • No regression allowed on Cachyos/Linux output — the original author's box must produce identical JSON. Add a Linux-fixture test to lock this in.
  • No removal of the MODEL_CONTEXT_WINDOWS OpenAI/Anthropic entries — additive only.

Fallback

If any sub-task hits the captures-skipped-behavior it must report back to the Orchestrator rather than edit a passing test to make itself happy. That is the SPEC anti-spin rule (#9 in the article).

Verifier (the independent eyes inside the loop)

The streak verifier — python3 -m pytest tests/ -v × 10 — IS the separate checker model from the article (#2, Boris's verifier loop). The local LLM does not grade its own homework: a different invocation runs the suite after each implement pass and the count resets on any non-zero exit.