Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

6.0 KiB
Raw Permalink Blame History

DECOMPOSITION — runnable-test-suite

Method

Decompose by capability boundary, not by file. Each sub-task is independently verifiable and independently mergeable. Local-LLM context budget per sub-task: max 9k tokens peak on this box (32GB RAM, no GPU, Model: auto).

Sub-tasks (3)

subtask-1: make-tests-runnable

Scope: pytest install path + python → python3 doc sweep + streak verifier. Files touched: requirements.txt (new), AGENTS.md, README.md, automaton/dashboard/README.md, prompts/orchestrate.md, scripts/install.sh (append venv snippet, idempotent), CHANGELOG.md. Not touched: vram_detect.py, status.py, any test logic. Acceptance: pip3 install -r requirements.txt && python3 -m pytest tests/ -v exits 0 from a clean clone; 10 consecutive clean streak; rg "^python " returns zero matches in docs/prompts. Peak context estimate: ~4k tokens (mostly mechanical doc edits). Fits easily. Run order: first. Establishes the green-test baseline the other sub-tasks need.

subtask-2: vram-detect-cross-platform

Scope: Make scripts/vram_detect.py work on macOS, Windows, Linux without behavior change on Cachyos/Linux. Pure detection logic — no CLI/JSON-shape changes. Files touched: scripts/vram_detect.py only. Functions to refactor (by line in current file):

  • detect_gpu_vram() (vram_detect.py:73-108): branch on platform.system().
    • Linux: keep nvidia-smi → lspci -vnn path.
    • macOS: add system_profiler SPDisplaysDataType → parse VRAM (Total) and Chipset Vendor (Apple Unified Memory counts as VRAM). Probe ioreg -c IOPlatformDevice only if system_profiler is unavailable.
    • Windows: add wmic path win32_VideoController get AdapterRAM,Name (deprecated but ubiquitous); fallback to PowerShell Get-CimInstance Win32_VideoController -Property AdapterRAM. Sum across GPUs.
  • detect_ram() (vram_detect.py:137-162): branch on platform.system().
    • Linux: keep /proc/meminfo.
    • macOS: keep sysctl -n hw.memsize (already works as fallback).
    • Windows: add wmic ComputerSystem get TotalPhysicalMemory; PowerShell fallback (Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory.
  • MODEL_CONTEXT_WINDOWS (vram_detect.py:24-47): add local-LLM entries: llama-3.1-8b, llama-3.3-70b, qwen2.5-7b, qwen2.5-72b, mistral-7b, mistral-large, deepseek-r1, deepseek-v3, glm-4, glm-4.5, gemma-2, gemma-2-27b, phi-3, phi-4. Use community-published context sizes. No fabricating — every entry must cite the source model card in a comment.
  • detect_model_context() (vram_detect.py:175-228): add ollama list probe when no config file specifies a model. Pick the first running model name and look it up.
  • run_command() (vram_detect.py:54-70): on Windows, route PowerShell cmdlets via powershell -NoProfile -Command "..." wrapper. Keep shutil.which gating. Not touched: test files (those are subtask-3), JSON output shape, CLI args. Acceptance on this machine (Darwin): python3 vram_detect.py prints macOS VRAM (non-zero on Apple Silicon), RAM 32GB, recommends ≥8k target. JSON has gpu_vram_gb > 0. On Linux (CI), output unchanged from current. Anti-spin rail: if a new entry in MODEL_CONTEXT_WINDOWS is unknown, fail open with 0, do NOT guess. Per SPEC, wrong-context detection is worse than none. Peak context estimate: ~7k tokens (single 536-line file, surgical edits). Fits. Run order: second, parallel-ok with subtask-1 (independent files).

subtask-3: vram-detect-cross-platform-tests

Scope: pytest tests proving cross-platform branches without hitting real hardware. Files touched: tests/test_vram_detect.py only. Patterns:

  • Parametrize detect_ram across Linux/Darwin/Windows with monkeypatch on platform.system, Path.exists, subprocess.run, and Path.read_text; assert correct KB returned and correct print lines emitted (capfd).
  • Parametrize detect_gpu_vram with mocked system_profiler / wmic / nvidia-smi stdout fixtures (kept as multiline string constants).
  • Assert unknown model name returns 0 (fail-open contract from subtask-2).
  • Assert _lookup_model_context picks the longest matching prefix (so llama-3.1-8b-instruct matches llama-3.1-8b).
  • Assert existing Linux/Cachyos path still parses /proc/meminfo (regression).
  • No live subprocess against real system_profiler/nvidia-smi — every call goes through monkeypatch.setattr. Not touched: vram_detect.py itself, any other source file, any prompt. Acceptance: added tests pass; total suite still 10-streak clean. Peak context estimate: ~5k tokens. Fits. Run order: third, AFTER subtask-2 (depends on its function signatures).

Dependency graph

subtask-1 ─┐
            ├─> parent done
subtask-2 ─┤
            └─> subtask-3 ──> parent done

Parent runnable-test-suite is complete only when ALL three sub-tasks pass the streak verifier from the SPEC (10 consecutive clean python3 -m pytest tests/ -v).

Parent non-goals

  • No usage of psutil, wmi, pywin32, or other new third-party deps. Pure stdlib (platform, subprocess, shutil, re, sys). Per SPEC constraint #1.
  • No regression allowed on Cachyos/Linux output — the original author's box must produce identical JSON. Add a Linux-fixture test to lock this in.
  • No removal of the MODEL_CONTEXT_WINDOWS OpenAI/Anthropic entries — additive only.

Fallback

If any sub-task hits the captures-skipped-behavior it must report back to the Orchestrator rather than edit a passing test to make itself happy. That is the SPEC anti-spin rule (#9 in the article).

Verifier (the independent eyes inside the loop)

The streak verifier — python3 -m pytest tests/ -v × 10 — IS the separate checker model from the article (#2, Boris's verifier loop). The local LLM does not grade its own homework: a different invocation runs the suite after each implement pass and the count resets on any non-zero exit.