What's Changed
Models run on MLX on Apple Silicon by default
In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.
ollama pull qwen3.8
ollama run qwen3.8
Additional models include gemma4, qwen3.6 and qwen3.5
Decision models are now available on MLX as well: Nimble tev1 clef clef-flash
We will continue testing and enabling additional models.
Full Changelog: v0.34.4...v0.40.0-rc3