🔥 What's New: v0.17.0
- Optimized Local Attention: Reduced memory overhead and enabled support for even longer contexts.
- Apple Silicon Acceleration: Added Metal residency support for improved inference performance on Apple devices.
- Extended Gemma 4 (12B): Introduced multimodal capabilities, multi-token prediction (MTP) acceleration, and extended context window support.
- Stability & Internals: General bug fixes and framework-level performance improvements.