# vram_detect.py model prefix matching `_lookup_model_context()` must avoid false prefix matches. The correct approach is three-tier: 1. **Exact match** — `name_lower == key_lower` 2. **Ollama parameter tag** — `name_lower.startswith(key_lower + ":")` (e.g. `deepseek-r1:7b` matches `deepseek-r1`) 3. **Known instruction-tuning suffix** — `name_lower.startswith(key_lower + "-")` only if next segment is in `_KNOWN_MODEL_SUFFIXES = {"instruct", "chat", "it", "fp16", "f16", "bf16"}` Keys sorted by length descending so most specific match wins first. This prevents `phi-4` matching `phi-4-mini-instruct` or `gpt-4o` matching `gpt-4o-foo-unknown`.