- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
2.8 KiB
2.8 KiB
Code Review: fix-context-sizing
Status: PASS
Author
AI assistant (per user directive "i approve all tasks. drive them to completion")
Files Reviewed
scripts/vram_detect.py(350-line region) —recommend_contextrewrite,main()extension with--loop-mode, JSON output additionsconfig.md— new## Loop Role Modelssectionprompts/decompose.md— 4k tier added, 16k floorREFUSEline addedtests/test_context_sizing.py— 15 new teststests/test_vram_detect.py:76— updatedtest_recommend_context_api_modelto assert the corrected single-headroom contract
Spec Conformance (R1–R6)
- R1 (single headroom): PASS —
recommend_contextreturnsrecommended_kbpre-headroom,max_peak_kb = net_kb * (100 - headroom_pct) // 100(headroom exactly once). Confirmed bytest_recommend_context_single_headroom. - R2 (no fake defaults): PASS —
else 8/else 6fallbacks removed;max(0, ...)clamp removed. Warning emitted in non-loop mode when budget ≤ 0 or model unknown. Confirmed bytest_no_fake_defaults_when_budget_zero,test_no_max_zero_clamp_in_output. - R3 (
--loop-moderefuse): PASS — unknown model exits 2 ("model context window is unknown"); sub-floor budget check gated onLOOP_MODE_CONTEXT_FLOOR_KB == 16_000. User override is authoritative per existing_parse_config_modelflow. Confirmed bytest_loop_mode_refuses_unknown_model,test_loop_mode_passes_for_known_model. - R4 (JSON fields): PASS —
available_context_kb,loop_mode_eligible,loop_modepresent;available_context_kb == max_peak_context_kb. Confirmed bytest_json_includes_available_context_kb_and_eligible. - R5 (
config.md## Loop Role Models): PASS — section added verbatim with role definitions, D12/D13 references. Confirmed bytest_config_md_includes_loop_role_models_section. - R6 (decompose.md 4k tier + 16k floor): PASS —
**4k VRAM**in guidelines,(4k VRAM)in size targets,≤ 16k ... REFUSEfloor line present. Confirmed bytest_decompose_md_includes_4k_tier,test_decompose_md_includes_16k_floor_refuse.
Test Results
- New tests: 15 passed
- Full suite: 264 passed (one pre-existing test updated to match fixed contract; no regression)
- Pre-existing self-consistency suite: 74 passed (no regression to prompt path enforcement or framework invariants)
Risks and Observations
- One pre-existing test (
test_recommend_context_api_model) was asserting the bug. Updated to assert the fixed contract. Documented inline as the reason. --loop-modeis opt-in via CLI flag. Non-loop callers preserve prior behavior + gain a human-readable warning. Backward-compatible.- The override-is-authoritative path (D13) was already honored by
_parse_config_model; no new parsing code needed.
Verdict
Ship. No defects blocking transition.