Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
This commit is contained in:
@@ -0,0 +1 @@
|
||||
complete
|
||||
@@ -0,0 +1,2 @@
|
||||
research:approved|2026-06-21T19:11:28.211174+00:00|user
|
||||
code_review:approved|2026-06-21T19:23:42.983607+00:00|user
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
# ADVERSARIAL_BUG_REPORT
|
||||
|
||||
10/10 streak = independent checker (article #2). No findings.
|
||||
@@ -0,0 +1,3 @@
|
||||
# BUG_REPORT
|
||||
|
||||
No bugs found. Streak verifier saw no failures.
|
||||
@@ -0,0 +1,3 @@
|
||||
# CODE_REVIEW
|
||||
|
||||
10-streak verifier (Boris loop, article #2). See IMPLEMENTATION.md.
|
||||
@@ -0,0 +1,3 @@
|
||||
# DOC_REVIEW
|
||||
|
||||
Doc changes in IMPLEMENTATION.md. No further work.
|
||||
@@ -0,0 +1,90 @@
|
||||
# IMPLEMENTATION — vram-detect-cross-platform-tests
|
||||
|
||||
Parent: `runnable-test-suite` (see PARENT_SPEC.md).
|
||||
SPEC: `tasks/runnable-test-suite/subtasks/vram-detect-cross-platform-tests/SPEC.md`.
|
||||
|
||||
## File touched
|
||||
|
||||
- `tests/test_vram_detect.py` — extended with 11 new test functions + helpers.
|
||||
No other file was modified.
|
||||
|
||||
## Test functions added (11)
|
||||
|
||||
| # | Function | What it verifies |
|
||||
|---|----------|------------------|
|
||||
| 1 | `test_detect_ram_linux` | `/proc/meminfo` parse → `(16_384_000, 8_192_000)` via `detect_ram()` with `platform.system()` patched to `Linux`. |
|
||||
| 2 | `test_detect_ram_macos` | `sysctl -n hw.memsize` → `"34359738368"` → total_kb `33_554_432`, available == total. |
|
||||
| 3 | `test_detect_ram_windows` | `wmic ComputerSystem` → `TotalPhysicalMemory=34359738368` → total_kb `33_554_432`. |
|
||||
| 4 | `test_detect_gpu_vram_nvidia_linux` | `nvidia-smi` → `"24576\n"` → `(25_165_824, 25_165_824, 1)`. |
|
||||
| 5 | `test_detect_gpu_vram_apple_silicon` | `system_profiler` Apple M2 snippet + `sysctl hw.memsize=17179869184` → total > 0, num >= 1. |
|
||||
| 6 | `test_detect_gpu_vram_windows_wmic` | `wmic win32_VideoController` → `AdapterRAM=8589934592` → total_vram_kb `8_388_608`. |
|
||||
| 7 | `test_lookup_model_context_unknown_returns_zero` | `_lookup_model_context("completely-unknown-model")` → `0`. |
|
||||
| 8 | `test_lookup_model_context_prefix_match` | `deepseek-r1:7b` → `64_000`; `llama-3.1-8b-instruct` → `128_000`. |
|
||||
| 9 | `test_detect_model_context_ollama_probe` | No config model; `ollama list` returns `llama-3.1-8b` → context `128_000`. |
|
||||
| 10 | `test_run_command_windows_powershell_wrapper` | On Windows, `Get-CimInstance ...` is routed through `powershell -NoProfile -NoLogo -Command "..."`. |
|
||||
| 11 | `test_detect_ram_linux_regression` (LOCKED) | Exact match `(32_768_000, 16_384_000)` for a 32GB `/proc/meminfo` fixture via `_detect_ram_linux()`. |
|
||||
|
||||
Counting parametrize cases: **11 new tests** (no parametrization used; each function
|
||||
is a single case).
|
||||
|
||||
## Existing tests modified
|
||||
|
||||
None. All 9 pre-existing test functions (`test_lookup_model_context`,
|
||||
`test_parse_token_value`, `test_extract_value`, `test_parse_config_model`,
|
||||
`test_parse_config_model_skips_code_blocks`, `test_parse_vram_config_manual`,
|
||||
`test_recommend_context_api_model`, `test_recommend_context_manual_mode`,
|
||||
`test_extract_model_from_file_respects_10kb_limit`) were preserved byte-for-byte.
|
||||
No diffs to existing code.
|
||||
|
||||
## Mocking strategy
|
||||
|
||||
Every external call is mocked via `monkeypatch.setattr` — no live subprocess,
|
||||
`system_profiler`, `nvidia-smi`, or `wmic` invocation:
|
||||
|
||||
- `vram.platform.system` → lambda returning the target OS string.
|
||||
- `vram.subprocess.run` → `_make_fake_run(responses)` mapping `cmd[0]` (or the
|
||||
joined PowerShell string) to a canned `_FakeResult(stdout, returncode=0)`.
|
||||
- `vram.shutil.which` → lambda returning a truthy name (or `None` for non-ollama
|
||||
in the ollama probe test).
|
||||
- `Path.exists` / `Path.read_text` → `_patch_meminfo` serves the fixture only
|
||||
for `/proc/meminfo` and falls through to the original for any other path
|
||||
(keeps `tmp_path` and pytest internals working during the test).
|
||||
- `vram.Path.home` → `tmp_path` in the ollama probe test so the real
|
||||
`~/.automaton/config.md` is never consulted.
|
||||
|
||||
Module-level multiline string fixtures: `MEMINFO_LINUX_16GB`,
|
||||
`MEMINFO_LINUX_32GB`, `APPLE_M2_PROFILER`, `WMIC_VIDEOCONTROLLER`,
|
||||
`WMIC_COMPUTERSYSTEM`, `OLLAMA_LIST`.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
| Criterion | Result |
|
||||
|-----------|--------|
|
||||
| `pytest tests/test_vram_detect.py -v` exits 0, all new tests pass | PASS — 20/20 (9 existing + 11 new) |
|
||||
| `pytest tests/ -v` exits 0 | PASS — 235 passed, 0 errors |
|
||||
| Streak: 10 consecutive clean `pytest tests/ -v` runs | PASS — `STREAK_COMPLETE attempt=1 clean=10/10` |
|
||||
| No live subprocess against real hardware | PASS — every `subprocess.run` / `shutil.which` / `Path` I/O patched |
|
||||
| Every new test < 500ms | PASS — entire file 0.01s; slowest 60 durations < 0.005s |
|
||||
| No `pytest.mark.skip` | PASS — none used |
|
||||
| No new deps | PASS — stdlib + pytest only |
|
||||
| Did NOT touch `scripts/vram_detect.py` | CONFIRMED — only `tests/test_vram_detect.py` edited |
|
||||
|
||||
## Final pass counts
|
||||
|
||||
- Baseline (`tests/test_vram_detect.py`): 9 passed.
|
||||
- After implementation (`tests/test_vram_detect.py`): 20 passed (+11).
|
||||
- Full suite baseline: 224 passed.
|
||||
- Full suite after: **235 passed** (+11), 0 errors, 0 skipped.
|
||||
|
||||
## Streak result
|
||||
|
||||
```
|
||||
STREAK_COMPLETE attempt=1 clean=10/10
|
||||
```
|
||||
|
||||
## STOP-and-report triggers hit
|
||||
|
||||
None. All 11 SPEC-required tests passed against subtask-2's `vram_detect.py`
|
||||
without any signature mismatch. No edits to `scripts/vram_detect.py` were
|
||||
required or made. The `ollama list` probe (SPEC test 9) is present in
|
||||
`vram_detect.py:525-539` and works as specified.
|
||||
@@ -0,0 +1,60 @@
|
||||
# Parent Task: runnable-test-suite
|
||||
|
||||
This is the parent SPEC for the runnable-test-suite task. Each sub-task references
|
||||
this for context, scope boundaries, and the parent acceptance contract.
|
||||
|
||||
## Parent Goal
|
||||
|
||||
Make `python3 -m pytest tests/ -v` pass from a clean checkout of `~/.automaton`,
|
||||
with deterministic Python deps pinned and docs that reflect the actual interpreter
|
||||
on stock macOS/Windows/Linux. In the same wave, make `scripts/vram_detect.py`
|
||||
cross-platform (macOS, Windows, Linux) — the existing version was developed on
|
||||
Cachyos and only fully works on Linux.
|
||||
|
||||
## Parent Acceptance Contract
|
||||
|
||||
1. `pip3 install -r requirements.txt` succeeds on stock macOS CommandLineTools + pip3.
|
||||
2. `python3 -m pytest tests/ -v` exits 0 from a clean clone, zero `error` lines.
|
||||
3. **Streak verifier:** 10 consecutive clean `python3 -m pytest tests/ -v` runs with
|
||||
no edits between runs. A single failure resets the count. Cap: 5 attempts.
|
||||
4. `python3 -m py_compile automaton/**/*.py automaton/dashboard/**/*.py scripts/*.py` exits 0.
|
||||
5. `bash -n scripts/*.sh` exits 0.
|
||||
6. `rg "^python " AGENTS.md README.md automaton/dashboard/README.md prompts/orchestrate.md`
|
||||
returns zero matches for a bare `python ` command.
|
||||
7. `python3 scripts/vram_detect.py` on Darwin prints `gpu_vram_gb > 0` (was 0 before).
|
||||
8. `python3 scripts/vram_detect.py` JSON shape identical to before on Linux/Cachyos.
|
||||
|
||||
## Sub-tasks
|
||||
|
||||
- `make-tests-runnable` — Wave 1, parallel-ok
|
||||
- `vram-detect-cross-platform` — Wave 1, parallel-ok
|
||||
- `vram-detect-cross-platform-tests` — Wave 2, depends on `vram-detect-cross-platform`
|
||||
|
||||
Parent is complete ONLY when ALL three sub-tasks pass and the streak verifier above
|
||||
runs 10 consecutive clean passes.
|
||||
|
||||
## Anti-spin rails (from the source article)
|
||||
|
||||
- The streak verifier IS the independent checker model from Boris's loop. The
|
||||
worker (local LLM) does not grade its own homework.
|
||||
- If a test is genuinely broken (not just import-failing due to missing pytest),
|
||||
STOP and report. Do not patch the test to make it pass. An agent that grades
|
||||
itself will delete the failing test and call it done.
|
||||
- Unknown model name → fail open with `0`. Wrong-context detection is worse than none.
|
||||
- No new third-party deps beyond `pytest`. Pure stdlib for `vram_detect.py`.
|
||||
|
||||
## Hardware/VRAM context
|
||||
|
||||
- Detected by `vram_detect.py` on this box: **32GB RAM, no GPU, model unknown**
|
||||
(because `vram_detect.py` is broken on macOS — subtask-2 fixes that)
|
||||
- Target context: 12k tokens, headroom 25%, max peak per sub-task: 9k tokens.
|
||||
- Sub-task peak estimates all fit within 9k. No further decomposition.
|
||||
|
||||
## Constraints / non-goals (parent)
|
||||
|
||||
- No `psutil`, `wmi`, `pywin32`, `tox`, `pytest-cov`, or other third-party deps.
|
||||
- No removal of existing OpenAI/Anthropic entries in `MODEL_CONTEXT_WINDOWS`.
|
||||
- No changes to `status.py`, `autopilot.py`, or any other runtime script's behavior.
|
||||
- No touching files under `tasks/` (those are state, not source).
|
||||
- No Docker, no conda, no `pyenv`. Stock `python3` + `pip3` only.
|
||||
- VRAM detection is additive — Linux/Cachyos output must NOT regress.
|
||||
@@ -0,0 +1,126 @@
|
||||
# SPEC — vram-detect-cross-platform-tests
|
||||
|
||||
Parent: `runnable-test-suite` (see PARENT_SPEC.md).
|
||||
|
||||
**Dependency:** This sub-task runs AFTER `vram-detect-cross-platform` is complete.
|
||||
Its function signatures and detection logic are the contract you test against.
|
||||
|
||||
## Scope
|
||||
|
||||
pytest tests proving cross-platform branches of `scripts/vram_detect.py` work
|
||||
WITHOUT hitting real hardware. Every `subprocess.run` / `Path.exists` / `sysctl`
|
||||
call is mocked.
|
||||
|
||||
## Files this sub-task touches (and ONLY this)
|
||||
|
||||
- `tests/test_vram_detect.py` — extend or rewrite. Preserve any existing
|
||||
passing test (Linux regression tests especially).
|
||||
|
||||
## MUST NOT touch
|
||||
|
||||
- `scripts/vram_detect.py` (if it has a bug, escalate back to Orchestrator —
|
||||
subtask-2 owns it)
|
||||
- Any documentation, prompt, install script, or other source file
|
||||
- Any other test file
|
||||
|
||||
## Test patterns (parametrize + monkeypatch)
|
||||
|
||||
### 1. `test_detect_ram_linux` (regression — must already be there, keep it)
|
||||
- Patches `Path("/proc/meminfo")` text with synthetic `MemTotal: 16384000 kB\nMemAvailable: 8192000 kB`.
|
||||
- Asserts `(16_384_000, 8_192_000)` returned.
|
||||
|
||||
### 2. `test_detect_ram_macos`
|
||||
- `monkeypatch.setattr(platform, "system", lambda: "Darwin")`.
|
||||
- Patches `subprocess.run` so `sysctl -n hw.memsize` returns `"34359738368"` (32GB).
|
||||
- Asserts total_kb correct (33_554_432), available == total.
|
||||
|
||||
### 3. `test_detect_ram_windows`
|
||||
- `monkeypatch.setattr(platform, "system", lambda: "Windows")`.
|
||||
- Patches subprocess to return wmic output `TotalPhysicalMemory=34359738368`.
|
||||
- Asserts total_kb == 33_554_432.
|
||||
|
||||
### 4. `test_detect_gpu_vram_nvidia_linux`
|
||||
- `monkeypatch.setattr(platform, "system", lambda: "Linux")`.
|
||||
- Patches `nvidia-smi` output `"24576\n"`.
|
||||
- Asserts `(25_165_824, 25_165_824, 1)` (24GB × 1024 = KB).
|
||||
|
||||
### 5. `test_detect_gpu_vram_apple_silicon`
|
||||
- `monkeypatch.setattr(platform, "system", lambda: "Darwin")`.
|
||||
- Patches `system_profiler SPDisplaysDataType` with a real Apple M-series snippet:
|
||||
```
|
||||
Graphics/Displays:
|
||||
Apple M2:
|
||||
Chipset Model: Apple M2
|
||||
Type: GPU
|
||||
Bus: Built-In
|
||||
Total Number of Cores: 10
|
||||
Vendor: Apple (0x106b)
|
||||
Metal: Supported, version 2
|
||||
```
|
||||
- Patches `sysctl -n hw.memsize` with `"17179869184"` (16GB).
|
||||
- Asserts gpu_vram_kb > 0, num_gpus >= 1.
|
||||
|
||||
### 6. `test_detect_gpu_vram_windows_wmic`
|
||||
- `monkeypatch.setattr(platform, "system", lambda: "Windows")`.
|
||||
- Patches wmic output:
|
||||
```
|
||||
AdapterRAM=8589934592
|
||||
Name=NVIDIA GeForce RTX 3060
|
||||
```
|
||||
- Asserts total_vram_kb == 8_388_608 (8GB).
|
||||
|
||||
### 7. `test_lookup_model_context_unknown_returns_zero`
|
||||
- `_lookup_model_context("completely-unknown-model")` returns `0`.
|
||||
|
||||
### 8. `test_lookup_model_context_prefix_match`
|
||||
- `_lookup_model_context("deepseek-r1:7b")` matches `deepseek-r1` entry and returns its context.
|
||||
- `_lookup_model_context("llama-3.1-8b-instruct")` matches `llama-3.1-8b`.
|
||||
|
||||
### 9. `test_detect_model_context_ollama_probe`
|
||||
- No config file specifies a model.
|
||||
- `ollama list` is on PATH (patch `shutil.which`).
|
||||
- Patches subprocess to return:
|
||||
```
|
||||
NAME ID SIZE MODIFIED
|
||||
llama-3.1-8b abc 4.7GB 2 days ago
|
||||
```
|
||||
- Asserts the returned context matches the `llama-3.1-8b` entry.
|
||||
|
||||
### 10. `test_run_command_windows_powershell_wrapper`
|
||||
- Show that on Windows, a PowerShell cmdlet invocation goes through
|
||||
`powershell -NoProfile -NoLogo -Command "..."` (assert argv[0] is powershell
|
||||
when cmd is a PS string).
|
||||
|
||||
### 11. `test_detect_ram_linux_regression` (LOCKED — do not modify)
|
||||
- If this test already exists, keep its byte-for-byte assertions. If not, add
|
||||
an exact match on `(total_kb, available_kb)` for a specific `/proc/meminfo`
|
||||
fixture to prevent subtask-2 from regressing Linux output.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
1. `python3 -m pytest tests/test_vram_detect.py -v` exits 0 with all new tests passing.
|
||||
2. Total `python3 -m pytest tests/ -v` still exits 0.
|
||||
3. **Streak:** 10 consecutive clean `python3 -m pytest tests/ -v` runs.
|
||||
4. No live `subprocess` against real `system_profiler`/`nvidia-smi`/`wmic` — every
|
||||
external call goes through `monkeypatch.setattr`.
|
||||
5. Every new test runs < 500ms (mock only, no I/O).
|
||||
|
||||
## Anti-spin rails
|
||||
|
||||
- If a subtask-2 function signature doesn't support what a test needs, STOP and
|
||||
report. Do not edit `vram_detect.py` yourself and do not weaken the test to fit.
|
||||
Escalate via Orchestrator.
|
||||
- No `pytest.mark.skip` unless the platform genuinely doesn't support the feature
|
||||
(e.g. skip a Windows test on Linux only if it can't be mocked — but mocking is
|
||||
the whole point, so this should never happen).
|
||||
- Do not add `pytest-cov` or any new deps.
|
||||
|
||||
## Recommended approach
|
||||
|
||||
1. Read `scripts/vram_detect.py` and confirm subtask-2's signatures.
|
||||
2. For each test, write the fixture data as a module-level constant (multiline string).
|
||||
3. Use `monkeypatch.setattr` for `platform.system`, `subprocess.run`, `shutil.which`,
|
||||
`Path.exists`, `Path.read_text`. Never call the real thing.
|
||||
4. Use `capfd` for stdout assertions where the SPEC calls for print messages.
|
||||
5. Run `python3 -m pytest tests/test_vram_detect.py -v` until green.
|
||||
6. Run the full suite streak verifier.
|
||||
@@ -0,0 +1,3 @@
|
||||
# VERDICT
|
||||
|
||||
PASS — All acceptance criteria met. 10/10 streak on attempt 1.
|
||||
Reference in New Issue
Block a user