7 memory files covering: - audit-bug-patterns: recurring status.py bug patterns - vram-model-matching: three-tier model prefix matching - dashboard-security: .state priority, CORS removal, register-guards - state-machine-workflow: legal transitions and approval gates - testing-conventions: pytest patterns and helpers - audit-process: report conventions and batch processing - framework-architecture: enforcement layers and key files
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
# vram_detect.py model prefix matching
|
||||
|
||||
`_lookup_model_context()` must avoid false prefix matches. The correct approach is three-tier:
|
||||
|
||||
1. **Exact match** — `name_lower == key_lower`
|
||||
2. **Ollama parameter tag** — `name_lower.startswith(key_lower + ":")` (e.g. `deepseek-r1:7b` matches `deepseek-r1`)
|
||||
3. **Known instruction-tuning suffix** — `name_lower.startswith(key_lower + "-")` only if next segment is in `_KNOWN_MODEL_SUFFIXES = {"instruct", "chat", "it", "fp16", "f16", "bf16"}`
|
||||
|
||||
Keys sorted by length descending so most specific match wins first. This prevents `phi-4` matching `phi-4-mini-instruct` or `gpt-4o` matching `gpt-4o-foo-unknown`.
|
||||
Reference in New Issue
Block a user