16 lines
919 B
Markdown
16 lines
919 B
Markdown
# Spec: fix-vram-model-prefix-match
|
|||
|
|
|
||
|
|
## Problem
|
||
|
|
`scripts/vram_detect.py:396` uses `model_name.lower().startswith(key.lower())` to match model names. This prefix matching causes false matches: `phi-4-mini` matches `phi-4` (16000), and unknown models starting with known prefixes get incorrect context windows instead of the fallback.
|
||
|
|
|
||
|
|
## Fix
|
||
|
|
Try exact match first, then longest-prefix match (sort keys by length descending). Only match if the model name equals the key or starts with `key + "-"` (to avoid `phi-4` matching `phi-40`).
|
||
|
|
|
||
|
|
## Acceptance Criteria
|
||
|
|
- `phi-4-mini-instruct` does NOT match `phi-4` — returns fallback (128000)
|
||
|
|
- `gpt-4o` still matches `gpt-4o` (exact) — returns 128000
|
||
|
|
- `gpt-4o-mini` matches `gpt-4o-mini` (exact) — returns 128000
|
||
|
|
- `claude-3-5-sonnet-20241022` matches exact entry — returns 200000
|
||
|
|
- Existing tests in `test_vram_detect.py` still pass
|
||
|
|
- Add test for the prefix edge case
|