- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
32 lines
2.3 KiB
Markdown
32 lines
2.3 KiB
Markdown
# Bug Report: fix-context-sizing
|
|
|
|
## Status: NO_BUGS_FOUND
|
|
|
|
## Method
|
|
Static re-read of `vram_detect.py` changes against the SPEC and against pre-existing behaviors. Focused on:
|
|
1. **Backward compatibility**: non-loop callers must keep prior behavior.
|
|
2. **Override authority**: `_parse_config_model`'s override flow must still be authoritative in loop mode.
|
|
3. **Negative-budget propagation**: `max_peak_kb` can now be negative when overhead exceeds budget — does any consumer read it without wrapping in `max(0, ...)`?
|
|
4. **Exit codes**: `--loop-mode` refuse paths must exit `2`, not `0`.
|
|
|
|
## Findings
|
|
|
|
### F1. Non-loop behavior preserved — CONFIRMED SAFE
|
|
Pre-existing CLI callers (`vram_detect.py` without `--loop-mode`) and dashboard invocations are unaffected. Warning lines are emitted on degraded conditions but exit code stays `0`. Verified by `test_non_loop_mode_does_not_refuse_unknown_model`.
|
|
|
|
### F2. Override authority — CONFIRMED SAFE
|
|
In `detect_model_context`, the `if override_context and override_context != "auto"` branch (line 343-345) returns the override *before* reaching the fallback-return-zero path. So when a user has set `Override context window` in `config.md`, `detect_model_context` returns a positive value and the `--loop-mode` unknown-model refuse never fires. Matches D13 spec.
|
|
|
|
### F3. Negative `max_peak_kb` propagation — LOW RISK, OUT OF SCOPE
|
|
`recommend_context` now returns a negative `max_peak_kb` when `overhead_tokens > recommended_kb`. No current consumer reads it without arithmetic. The dashboard (`automaton/dashboard/__main__.py`) computes its own derived numbers and does not display this value directly. The loop runner (task `add-loop-runner`) must `max(0, available_context_kb)` before scheduling; that's the runner's obligation, not this task's. Acceptable.
|
|
|
|
### F4. Exit code correctness — CONFIRMED
|
|
`--loop-mode` refuse paths call `return 2`. Verified directly: `python3 vram_detect.py --model zzz --loop-mode; echo $?` returns `2`. Support: `test_loop_mode_refuses_unknown_model` writes a regression guard.
|
|
|
|
## Bugs Found
|
|
None.
|
|
|
|
## Out of Scope (for downstream tasks)
|
|
- Loop runner's `max(0, available_context_kb)` wrap (task 3 `add-loop-runner`)
|
|
- Dashboard showing the new `loop_mode_eligible` field (v1.1 dashboard panel)
|
|
- Tier 2 cleanup: `last-read-sha` drift detection, etc. |