CI / build (push) Has been cancelled
7 memory files covering: - audit-bug-patterns: recurring status.py bug patterns - vram-model-matching: three-tier model prefix matching - dashboard-security: .state priority, CORS removal, register-guards - state-machine-workflow: legal transitions and approval gates - testing-conventions: pytest patterns and helpers - audit-process: report conventions and batch processing - framework-architecture: enforcement layers and key files
10 lines
656 B
Markdown
10 lines
656 B
Markdown
# vram_detect.py model prefix matching
|
|
|
|
`_lookup_model_context()` must avoid false prefix matches. The correct approach is three-tier:
|
|
|
|
1. **Exact match** — `name_lower == key_lower`
|
|
2. **Ollama parameter tag** — `name_lower.startswith(key_lower + ":")` (e.g. `deepseek-r1:7b` matches `deepseek-r1`)
|
|
3. **Known instruction-tuning suffix** — `name_lower.startswith(key_lower + "-")` only if next segment is in `_KNOWN_MODEL_SUFFIXES = {"instruct", "chat", "it", "fp16", "f16", "bf16"}`
|
|
|
|
Keys sorted by length descending so most specific match wins first. This prevents `phi-4` matching `phi-4-mini-instruct` or `gpt-4o` matching `gpt-4o-foo-unknown`.
|