CI / build (push) Has been cancelled
- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/ - status.py: add --cleanup-done and --install-cleanup-schedule commands - Add scripts/automaton-cleanup.sh for periodic task archiving - Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders - .rules.md: add Self-Documenting UI Names rule - New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
2.1 KiB
2.1 KiB
Verdict: fix-context-sizing
Status: PASS
Completion Date: 2026-06-22
Summary
Tier 1 context-sizing layer landed: vram_detect.py's recommend_context now applies headroom exactly once (was 3×); fake 8k/6k fallbacks and max(0,...) lying clamp removed; new --loop-mode CLI flag refuses unknown models and sub-16k available context per D13; JSON output exposes available_context_kb + loop_mode_eligible for the upcoming loop runner. config.md documents ## Loop Role Models; decompose.md gains a 4k tier and an explicit ≤ 16k: REFUSE floor.
Phase Outcomes
- Research: SPEC.md produced and approved by user before implementation.
- Implement: code edits + IMPLEMENTATION.md produced; 15 new tests in
tests/test_context_sizing.py; one pre-existing test (test_recommend_context_api_model) updated to assert the corrected single-headroom contract. - Code Review: PASS — full spec conformance R1–R6 verified; 264/264 tests green.
- Bug Find: NO_BUGS_FOUND — 4 adversarial vectors probed, no defects.
- Adversarial Bug Find: NO_NEW_DEFECTS — 5 attacks from design's §6 model all produce correct refuse/warn behavior.
- Doc Review: PASS —
config.md+decompose.mdupdated; doc cross-references verified. - Referee: PASS — all phase artifacts present and consistent.
Test Results
- New: 15 passed / 15
- Pre-existing updated: 1 (
test_recommend_context_api_model— was asserting the bug; now asserts the fix) - Full suite: 264 passed
- Self-consistency suite: 74 passed (no prompt path regressions)
Findings
- All six SPEC requirements (R1–R6) implemented and guard-tested.
- Pre-existing test that encoded the old buggy behavior was properly updated; reason documented inline in commit message.
Tasks for Review / Tie-Breaks
- None
Remaining Issues (out of scope, tracked)
- Loop runner must
max(0, available_context_kb)before scheduling — task 3 (add-loop-runner). README.mdhuman-facing CLI doc for--loop-mode— Tier 2 cleanup.- v1.1 dashboard "Loops" panel will surface
loop_mode_eligible— v1.1.
Score
+10