🔥 What's New: v0.11.0
Gemma 4 Multi-token Prediction (MTP) Support: Supercharge Gemma 4 on-device inference with Single Position Multi Token Prediction (MTP), delivering >2x faster decode speeds on mobile GPUs with zero quality degradation (blog, documentation).
Windows Native Support: The LiteRT-LM CLI now runs natively on Windows with both CPU and GPU backend support.