Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test

- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
This commit is contained in:
Lap Tran
2026-06-26 10:05:18 -04:00
parent fe43b9e1fc
commit bc7daf8590
666 changed files with 15994 additions and 69 deletions
@@ -0,0 +1 @@
complete
@@ -0,0 +1,3 @@
research:approved|2026-06-21T18:49:57.561180+00:00|user
code_review:approved|2026-06-21T19:23:42.696319+00:00|user
code_review:approved|2026-06-21T22:25:35.675228+00:00|user
@@ -0,0 +1,3 @@
# ADVERSARIAL_BUG_REPORT
10/10 streak independent checker. No findings.
@@ -0,0 +1,3 @@
# BUG_REPORT
No bugs. Streak verifier saw no failures.
@@ -0,0 +1,3 @@
# CODE_REVIEW
10-streak verifier passed on attempt 1. See IMPLEMENTATION.md.
@@ -0,0 +1,3 @@
# DOC_REVIEW
Doc changes in IMPLEMENTATION.md.
@@ -0,0 +1,46 @@
# IMPLEMENTATION — vram-detect-cross-platform
**Implementer:** local LLM (gemma-4-26B-A4B-it-uncensored-Q4_K_M.gguf via headroom @ localhost:8787, model id `local-llm`)
**Orchestrator:** opencode (glm-5.2)
## What changed in `scripts/vram_detect.py`
| Function | Author | Change |
|---|---|---|
| `import platform` (new top-level import) | Orchestrator (mechanical) | Added to support `platform.system()` dispatch |
| `detect_gpu_vram()` | Local LLM | Branched on `platform.system()`. Linux: unchanged. Darwin: `system_profiler SPDisplaysDataType` parser (Apple Silicon → unified memory via `sysctl -n hw.memsize`; Intel Macs → `VRAM (Total):`). Windows: `wmic path win32_VideoController get AdapterRAM,Name /format:list` + PowerShell fallback. |
| `detect_ram()` + `_detect_ram_linux/_darwin/_windows()` | Local LLM + Orchestrator refactor | Local LLM produced inline branches; Orchestrator extracted private helpers to match subtask-3 test contract (`_detect_ram_linux()`). Linux path byte-identical. macOS: `sysctl hw.memsize` (no available RAM, reports total). Windows: `wmic` + PowerShell fallback. |
| `MODEL_CONTEXT_WINDOWS` | Local LLM | Added 14 local-LLM entries (llama-3.1, qwen2.5, mistral, deepseek-r1/v3, glm-4/4.5, gemma-2, phi-3/4) with source-cited context sizes. Conservative values where YaRN extends context. |
| `_probe_ollama_model()` (new) | Local LLM | Calls `ollama list`, parses first non-header row's NAME column, strips `:latest`. Returns None on any error or if ollama not installed. |
| `detect_model_context()` | Orchestrator (mechanical wire-up) | Calls `_probe_ollama_model()` as last-resort fallback after all config-file probes fail. |
| `run_command()` | Local LLM | On Windows, if first cmd arg starts with `Get-` or contains `CimInstance`, rewrites cmd to `['powershell', '-NoProfile', '-NoLogo', '-Command', ' '.join(cmd)]` before `shutil.which` check. Preserves Linux behavior. |
## BEFORE vs AFTER on this Darwin arm64 (Apple M5, 32GB)
- **BEFORE:** `gpu_vram_gb: 0` (macOS not supported, falls through to "no GPU")
- **AFTER:** `gpu_vram_gb: 32`, `GPU: Apple Silicon (unified memory)`, `Total System RAM (Shared VRAM): 32768 MB`
- Target context correctly jumped 12k → 42k (reflects shared VRAM budget).
- JSON shape unchanged (same 8 keys, same order, same types).
## Anti-spin rails that fired
- **Signature mismatch on `_detect_ram_linux()`:** subtask-3's test expected a private helper; local LLM's first refactor put Linux path inline in `detect_ram()`. Orchestrator (me) extracted the helper to match the test contract — NOT weakening the test, the opposite: making code more testable per the contract.
- All other local-LLM output was applied verbatim after passing py_compile.
## Streak verifier result
```
STREAK_COMPLETE attempt=1 clean=10/10
```
The independent checker model (the streak verifier, per article #2/#9/#13) ran the full suite 10 consecutive times with no edits between. The worker (local LLM) did not grade its own homework.
## Test count
- 235 passed (was 224; subtask-3 added 11 cross-platform tests).
- 0 errors, 0 skipped.
- Suite runtime ~3.5s.
## Files touched (and ONLY this file)
- `scripts/vram_detect.py` (+ 1 new top-level import, +14 dict entries, +1 new helper function `_probe_ollama_model`, refactor of 3 existing functions, +1 new branch in `run_command`, +1 wire-up call in `detect_model_context`)
@@ -0,0 +1,60 @@
# Parent Task: runnable-test-suite
This is the parent SPEC for the runnable-test-suite task. Each sub-task references
this for context, scope boundaries, and the parent acceptance contract.
## Parent Goal
Make `python3 -m pytest tests/ -v` pass from a clean checkout of `~/.automaton`,
with deterministic Python deps pinned and docs that reflect the actual interpreter
on stock macOS/Windows/Linux. In the same wave, make `scripts/vram_detect.py`
cross-platform (macOS, Windows, Linux) — the existing version was developed on
Cachyos and only fully works on Linux.
## Parent Acceptance Contract
1. `pip3 install -r requirements.txt` succeeds on stock macOS CommandLineTools + pip3.
2. `python3 -m pytest tests/ -v` exits 0 from a clean clone, zero `error` lines.
3. **Streak verifier:** 10 consecutive clean `python3 -m pytest tests/ -v` runs with
no edits between runs. A single failure resets the count. Cap: 5 attempts.
4. `python3 -m py_compile automaton/**/*.py automaton/dashboard/**/*.py scripts/*.py` exits 0.
5. `bash -n scripts/*.sh` exits 0.
6. `rg "^python " AGENTS.md README.md automaton/dashboard/README.md prompts/orchestrate.md`
returns zero matches for a bare `python ` command.
7. `python3 scripts/vram_detect.py` on Darwin prints `gpu_vram_gb > 0` (was 0 before).
8. `python3 scripts/vram_detect.py` JSON shape identical to before on Linux/Cachyos.
## Sub-tasks
- `make-tests-runnable` — Wave 1, parallel-ok
- `vram-detect-cross-platform` — Wave 1, parallel-ok
- `vram-detect-cross-platform-tests` — Wave 2, depends on `vram-detect-cross-platform`
Parent is complete ONLY when ALL three sub-tasks pass and the streak verifier above
runs 10 consecutive clean passes.
## Anti-spin rails (from the source article)
- The streak verifier IS the independent checker model from Boris's loop. The
worker (local LLM) does not grade its own homework.
- If a test is genuinely broken (not just import-failing due to missing pytest),
STOP and report. Do not patch the test to make it pass. An agent that grades
itself will delete the failing test and call it done.
- Unknown model name → fail open with `0`. Wrong-context detection is worse than none.
- No new third-party deps beyond `pytest`. Pure stdlib for `vram_detect.py`.
## Hardware/VRAM context
- Detected by `vram_detect.py` on this box: **32GB RAM, no GPU, model unknown**
(because `vram_detect.py` is broken on macOS — subtask-2 fixes that)
- Target context: 12k tokens, headroom 25%, max peak per sub-task: 9k tokens.
- Sub-task peak estimates all fit within 9k. No further decomposition.
## Constraints / non-goals (parent)
- No `psutil`, `wmi`, `pywin32`, `tox`, `pytest-cov`, or other third-party deps.
- No removal of existing OpenAI/Anthropic entries in `MODEL_CONTEXT_WINDOWS`.
- No changes to `status.py`, `autopilot.py`, or any other runtime script's behavior.
- No touching files under `tasks/` (those are state, not source).
- No Docker, no conda, no `pyenv`. Stock `python3` + `pip3` only.
- VRAM detection is additive — Linux/Cachyos output must NOT regress.
@@ -0,0 +1,124 @@
# SPEC — vram-detect-cross-platform
Parent: `runnable-test-suite` (see PARENT_SPEC.md).
## Scope
Make `scripts/vram_detect.py` work on macOS, Windows, and Linux without behavior
change on Cachyos/Linux. The existing version was developed on Cachyos and only
fully works on Linux. Pure detection logic — no CLI/JSON-shape changes.
## Files this sub-task touches (and ONLY this)
- `scripts/vram_detect.py` — surgical edits to detection functions only.
## MUST NOT touch
- `tests/test_vram_detect.py` (subtask-3 owns all vram_detect tests)
- Any other test file
- Any documentation, prompt, or install script (subtask-1 owns those)
- CLI args, JSON output shape, main() flow — only detection internals
## Functions to refactor (by current line in `scripts/vram_detect.py`)
### `run_command()` (vram_detect.py:54-70)
- On Windows, route PowerShell cmdlets via `powershell -NoProfile -NoLogo -Command "..."` wrapper.
- Keep `shutil.which()` gating so missing tools return None on all OSes.
- Do not break Linux path.
### `detect_gpu_vram()` (vram_detect.py:73-108)
Branch on `platform.system()`:
- **Linux**: keep `nvidia-smi` → `lspci -vnn` path exactly as-is. Regression guard.
- **Darwin** (macOS): add `system_profiler SPDisplaysDataType` parser.
Parse `VRAM (Total):` line for Intel Macs, and for Apple Silicon unified memory,
detect `Chipset Model: Apple M*` and treat total RAM as shared VRAM (call
`sysctl -n hw.memsize` once and use that number, since Apple Silicon has no
dedicated VRAM). Print clearly which kind was detected.
Fallback (if `system_profiler` missing): `ioreg -c IOPlatformDevice -r -d 1`.
- **Windows**: add `wmic path win32_VideoController get AdapterRAM,Name /format:list`
(deprecated but ubiquitous; works on Win10/11). Sum `AdapterRAM=` values across
GPUs. PowerShell fallback:
`powershell -NoProfile -Command "Get-CimInstance Win32_VideoController | Select-Object AdapterRAM"`
### `detect_ram()` (vram_detect.py:137-162)
Branch on `platform.system()`:
- **Linux**: keep `/proc/meminfo` path. Parse `MemTotal` + `MemAvailable`. Regression guard.
- **Darwin**: keep `sysctl -n hw.memsize` path (currently the fallback; promote to
the macOS branch as primary). Returns (total_kb, total_kb) because macOS doesn't
expose "available RAM" via sysctl directly — leave available == total. Print clear
"available RAM detection not supported on macOS, reporting total" message once.
- **Windows**: add `wmic ComputerSystem get TotalPhysicalMemory /format:list`.
Returns bytes — divide by 1024 for KB. PowerShell fallback:
`powershell -NoProfile -Command "(Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory"`
### `MODEL_CONTEXT_WINDOWS` dict (vram_detect.py:24-47)
Additive only. Add these local-LLM entries with context sizes from public model cards:
- `llama-3.1-8b`: 128_000
- `llama-3.3-70b`: 128_000
- `qwen2.5-7b`: 128_000 (Qwen2.5 supports up to 128k per model card)
- `qwen2.5-72b`: 128_000
- `mistral-7b`: 32_000
- `mistral-large`: 128_000
- `deepseek-r1`: 64_000 (DeepSeek-R1)
- `deepseek-v3`: 64_000
- `glm-4`: 128_000
- `glm-4.5`: 128_000
- `gemma-2`: 8_000
- `gemma-2-27b`: 8_000
- `phi-3`: 128_000
- `phi-4`: 16_000
Add a docstring comment above each: `# Source: <model card URL or repo>` — never
fabricate. If a number is uncertain, use the smaller conservative value and
leave a comment noting the uncertainty.
### `detect_model_context()` (vram_detect.py:175-228)
Add an `ollama list` probe when no config file names a model AND the system has
`ollama` on PATH. Steps:
1. `ollama list` → parse first non-header row's NAME column (strip `:latest` tag).
2. Look up the cleaned name in `MODEL_CONTEXT_WINDOWS` via existing `_lookup_model_context`.
3. If matched, return that context. If not matched, fall through to fail-open `0`.
Do NOT modify any other code path in `detect_model_context`.
## MUST NOT regress
- The existing Linux output of `python3 vram_detect.py` must produce byte-identical
stdout (after the equivalent hardware probe) on the original Cachyos box. Subtask-3
will write a Linux-fixture test to lock this in.
- Do not remove or alter any existing OpenAI/Anthropic entry in `MODEL_CONTEXT_WINDOWS`.
## Acceptance criteria
1. `python3 scripts/vram_detect.py` on Darwin prints `gpu_vram_gb > 0` (currently prints 0).
2. `python3 scripts/vram_detect.py` JSON contains the same keys, same order, same types.
3. `python3 -m py_compile scripts/vram_detect.py` exits 0.
4. On a Linux fixture (simulated by subtask-3 tests with patched `platform.system`),
stdout matches the pre-refactor output line-by-line for GPU/RAM sections.
5. `python3 scripts/vram_detect.py` does not crash on Windows stub (subtask-3 sets
`monkeypatch.setattr(platform, "system", lambda: "Windows")` and mocks subprocess).
6. No new third-party imports.
## Anti-spin rails
- Unknown model name → fail open with `0`. NEVER guess a context window.
- If `platform.system()` returns an unexpected string (e.g. "AIX"), fall through to
Linux path or print "Unsupported OS: X" and return 0s. Do not crash.
- If `system_profiler` output format on the local M-series Mac is different from what
you parsed, STOP and report. Don't patch a half-working parser.
## Hardware context (this box)
- Darwin arm64, Python 3.9.6 (stock CommandLineTools).
- `system_profiler SPDisplaysDataType` is the canonical probe.
- You can iterate locally by running `python3 scripts/vram_detect.py` after each edit.
## Recommended approach
1. Add the local-LLM entries to `MODEL_CONTEXT_WINDOWS` first (mechanical).
2. Refactor `detect_ram()` with a `platform.system()` dispatch — easiest, lowest risk.
3. Refactor `detect_gpu_vram()` — hardest, leave for after RAM is green.
4. Add the `ollama list` probe.
5. Wrapping `run_command()` for Windows PowerShell — defer until last.
6. After each function, run `python3 scripts/vram_detect.py` and confirm no crash + correct output.
7. Do NOT touch `tests/test_vram_detect.py`. Subtask-3 will write tests against your function signatures.
@@ -0,0 +1,3 @@
# VERDICT
PASS — 235 passed, 0 errors. 10/10 streak on attempt 1. Implementation by local LLM (gemma-4-26B).