v0.31.1: mlx: tighten up gemma4 moe loading code (#16964)
🔒
https://github.com
«This change allows .experts.gate_proj / .up_proj / .down_proj tensor names to each
be used for both quantized (i.e. nvfp4 and mxfp8) and non-quantized (bf16) models.
Previous to this only non-quantized models used that t...»
Automatische Weiterleitung...
1.5s