25 lines
1.4 KiB
Markdown
25 lines
1.4 KiB
Markdown
# Implementation: fix-vram-model-prefix-match
|
|||
|
|
|
||
|
|
## Bug
|
||
|
|
`_lookup_model_context()` in `vram_detect.py` used raw `startswith()` for model name matching, causing false positives like `phi-4` matching `phi-40` or `phi-4-mini-instruct` (a different model with different context window).
|
||
|
|
|
||
|
|
## Fix
|
||
|
|
Replaced the raw `startswith()` with a three-tier matching strategy:
|
||
|
|
1. **Exact match** — `name_lower == key_lower`
|
||
|
|
2. **Ollama parameter tag** — `name_lower.startswith(key_lower + ":")` (e.g. `deepseek-r1:7b` matches `deepseek-r1`)
|
||
|
|
3. **Known instruction-tuning suffix** — `name_lower.startswith(key_lower + "-")` only if the next segment is in `_KNOWN_MODEL_SUFFIXES = {"instruct", "chat", "it", "fp16", "f16", "bf16"}` (e.g. `llama-3.1-8b-instruct` matches `llama-3.1-8b`)
|
||
|
|
|
||
|
|
Keys are sorted by length descending so the most specific match wins first.
|
||
|
|
|
||
|
|
This prevents false matches:
|
||
|
|
- `phi-4-mini-instruct` → `mini` not in known suffixes → no match ✓
|
||
|
|
- `gpt-4o-foo-unknown` → `foo` not in known suffixes → no match ✓
|
||
|
|
- `phi-40` → no `:` or known-suffix separator → no match ✓
|
||
|
|
|
||
|
|
## Files Changed
|
||
|
|
- `scripts/vram_detect.py`: Added `_KNOWN_MODEL_SUFFIXES` set, rewrote `_lookup_model_context()` with three-tier matching
|
||
|
|
|
||
|
|
## Tests
|
||
|
|
- `test_lookup_model_context_no_false_prefix_match`: Asserts `phi-4-mini-instruct` and `gpt-4o-foo-unknown` return 0
|
||
|
|
- Existing `test_lookup_model_context_prefix_match` still passes (deepseek-r1:7b and llama-3.1-8b-instruct)
|