- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
6.0 KiB
DECOMPOSITION — runnable-test-suite
Method
Decompose by capability boundary, not by file. Each sub-task is independently
verifiable and independently mergeable. Local-LLM context budget per sub-task:
max 9k tokens peak on this box (32GB RAM, no GPU, Model: auto).
Sub-tasks (3)
subtask-1: make-tests-runnable
Scope: pytest install path + python → python3 doc sweep + streak verifier.
Files touched: requirements.txt (new), AGENTS.md, README.md,
automaton/dashboard/README.md, prompts/orchestrate.md, scripts/install.sh
(append venv snippet, idempotent), CHANGELOG.md.
Not touched: vram_detect.py, status.py, any test logic.
Acceptance: pip3 install -r requirements.txt && python3 -m pytest tests/ -v
exits 0 from a clean clone; 10 consecutive clean streak; rg "^python "
returns zero matches in docs/prompts.
Peak context estimate: ~4k tokens (mostly mechanical doc edits). Fits easily.
Run order: first. Establishes the green-test baseline the other sub-tasks need.
subtask-2: vram-detect-cross-platform
Scope: Make scripts/vram_detect.py work on macOS, Windows, Linux without
behavior change on Cachyos/Linux. Pure detection logic — no CLI/JSON-shape changes.
Files touched: scripts/vram_detect.py only.
Functions to refactor (by line in current file):
detect_gpu_vram()(vram_detect.py:73-108): branch onplatform.system().- Linux: keep
nvidia-smi→lspci -vnnpath. - macOS: add
system_profiler SPDisplaysDataType→ parseVRAM (Total)andChipset Vendor(Apple Unified Memory counts as VRAM). Probeioreg -c IOPlatformDeviceonly ifsystem_profileris unavailable. - Windows: add
wmic path win32_VideoController get AdapterRAM,Name(deprecated but ubiquitous); fallback to PowerShellGet-CimInstance Win32_VideoController -Property AdapterRAM. Sum across GPUs.
- Linux: keep
detect_ram()(vram_detect.py:137-162): branch onplatform.system().- Linux: keep
/proc/meminfo. - macOS: keep
sysctl -n hw.memsize(already works as fallback). - Windows: add
wmic ComputerSystem get TotalPhysicalMemory; PowerShell fallback(Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory.
- Linux: keep
MODEL_CONTEXT_WINDOWS(vram_detect.py:24-47): add local-LLM entries:llama-3.1-8b,llama-3.3-70b,qwen2.5-7b,qwen2.5-72b,mistral-7b,mistral-large,deepseek-r1,deepseek-v3,glm-4,glm-4.5,gemma-2,gemma-2-27b,phi-3,phi-4. Use community-published context sizes. No fabricating — every entry must cite the source model card in a comment.detect_model_context()(vram_detect.py:175-228): addollama listprobe when no config file specifies a model. Pick the first running model name and look it up.run_command()(vram_detect.py:54-70): on Windows, route PowerShell cmdlets viapowershell -NoProfile -Command "..."wrapper. Keepshutil.whichgating. Not touched: test files (those are subtask-3), JSON output shape, CLI args. Acceptance on this machine (Darwin):python3 vram_detect.pyprints macOS VRAM (non-zero on Apple Silicon), RAM 32GB, recommends ≥8k target. JSON hasgpu_vram_gb > 0. On Linux (CI), output unchanged from current. Anti-spin rail: if a new entry inMODEL_CONTEXT_WINDOWSis unknown, fail open with0, do NOT guess. Per SPEC, wrong-context detection is worse than none. Peak context estimate: ~7k tokens (single 536-line file, surgical edits). Fits. Run order: second, parallel-ok with subtask-1 (independent files).
subtask-3: vram-detect-cross-platform-tests
Scope: pytest tests proving cross-platform branches without hitting real hardware.
Files touched: tests/test_vram_detect.py only.
Patterns:
- Parametrize
detect_ramacrossLinux/Darwin/Windowswithmonkeypatchonplatform.system,Path.exists,subprocess.run, andPath.read_text; assert correct KB returned and correct print lines emitted (capfd). - Parametrize
detect_gpu_vramwith mockedsystem_profiler/wmic/nvidia-smistdout fixtures (kept as multiline string constants). - Assert unknown model name returns 0 (fail-open contract from subtask-2).
- Assert
_lookup_model_contextpicks the longest matching prefix (sollama-3.1-8b-instructmatchesllama-3.1-8b). - Assert existing Linux/Cachyos path still parses
/proc/meminfo(regression). - No live
subprocessagainst realsystem_profiler/nvidia-smi— every call goes throughmonkeypatch.setattr. Not touched:vram_detect.pyitself, any other source file, any prompt. Acceptance: added tests pass; total suite still 10-streak clean. Peak context estimate: ~5k tokens. Fits. Run order: third, AFTER subtask-2 (depends on its function signatures).
Dependency graph
subtask-1 ─┐
├─> parent done
subtask-2 ─┤
└─> subtask-3 ──> parent done
Parent runnable-test-suite is complete only when ALL three sub-tasks pass the
streak verifier from the SPEC (10 consecutive clean python3 -m pytest tests/ -v).
Parent non-goals
- No usage of
psutil,wmi,pywin32, or other new third-party deps. Pure stdlib (platform,subprocess,shutil,re,sys). Per SPEC constraint #1. - No regression allowed on Cachyos/Linux output — the original author's box must produce identical JSON. Add a Linux-fixture test to lock this in.
- No removal of the
MODEL_CONTEXT_WINDOWSOpenAI/Anthropic entries — additive only.
Fallback
If any sub-task hits the captures-skipped-behavior it must report back to the Orchestrator rather than edit a passing test to make itself happy. That is the SPEC anti-spin rule (#9 in the article).
Verifier (the independent eyes inside the loop)
The streak verifier — python3 -m pytest tests/ -v × 10 — IS the separate
checker model from the article (#2, Boris's verifier loop). The local LLM does
not grade its own homework: a different invocation runs the suite after each
implement pass and the count resets on any non-zero exit.