Files
automaton/tasks/complete/fix-context-sizing/CODE_REVIEW.md
T
Lap Tran 4a2301b077
CI / build (push) Has been cancelled
Archive completed tasks, add cleanup commands, self-documenting dashboard UI
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2026-06-24 22:43:33 -04:00

37 lines
2.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Code Review: fix-context-sizing
## Status: PASS
## Author
AI assistant (per user directive "i approve all tasks. drive them to completion")
## Files Reviewed
- `scripts/vram_detect.py` (350-line region) — `recommend_context` rewrite, `main()` extension with `--loop-mode`, JSON output additions
- `config.md` — new `## Loop Role Models` section
- `prompts/decompose.md` — 4k tier added, 16k floor `REFUSE` line added
- `tests/test_context_sizing.py` — 15 new tests
- `tests/test_vram_detect.py:76` — updated `test_recommend_context_api_model` to assert the corrected single-headroom contract
## Spec Conformance (R1–R6)
- R1 (single headroom): PASS — `recommend_context` returns `recommended_kb` pre-headroom, `max_peak_kb = net_kb * (100 - headroom_pct) // 100` (headroom exactly once). Confirmed by `test_recommend_context_single_headroom`.
- R2 (no fake defaults): PASS — `else 8` / `else 6` fallbacks removed; `max(0, ...)` clamp removed. Warning emitted in non-loop mode when budget ≤ 0 or model unknown. Confirmed by `test_no_fake_defaults_when_budget_zero`, `test_no_max_zero_clamp_in_output`.
- R3 (`--loop-mode` refuse): PASS — unknown model exits 2 ("model context window is unknown"); sub-floor budget check gated on `LOOP_MODE_CONTEXT_FLOOR_KB == 16_000`. User override is authoritative per existing `_parse_config_model` flow. Confirmed by `test_loop_mode_refuses_unknown_model`, `test_loop_mode_passes_for_known_model`.
- R4 (JSON fields): PASS — `available_context_kb`, `loop_mode_eligible`, `loop_mode` present; `available_context_kb == max_peak_context_kb`. Confirmed by `test_json_includes_available_context_kb_and_eligible`.
- R5 (`config.md` `## Loop Role Models`): PASS — section added verbatim with role definitions, D12/D13 references. Confirmed by `test_config_md_includes_loop_role_models_section`.
- R6 (decompose.md 4k tier + 16k floor): PASS — `**4k VRAM**` in guidelines, `(4k VRAM)` in size targets, `≤ 16k ... REFUSE` floor line present. Confirmed by `test_decompose_md_includes_4k_tier`, `test_decompose_md_includes_16k_floor_refuse`.
## Test Results
- New tests: 15 passed
- Full suite: 264 passed (one pre-existing test updated to match fixed contract; no regression)
- Pre-existing self-consistency suite: 74 passed (no regression to prompt path enforcement or framework invariants)
## Risks and Observations
- One pre-existing test (`test_recommend_context_api_model`) was *asserting the bug*. Updated to assert the fixed contract. Documented inline as the reason.
- `--loop-mode` is opt-in via CLI flag. Non-loop callers preserve prior behavior + gain a human-readable warning. Backward-compatible.
- The override-is-authoritative path (D13) was already honored by `_parse_config_model`; no new parsing code needed.
## Verdict
Ship. No defects blocking transition.