🎥 Video | YoutubeFlutter: Primary constructors and private named parameters(17.09.2026 um 18:00 Uhr)
🎥 Video | YoutubeGoogle Workspace: Generate and edit graphics with Google Pics(17.09.2026 um 18:17 Uhr)
🎥 Video | YoutubeAndroid Developers: How Android Bench 2.0 pushes AI evaluations(17.09.2026 um 18:36 Uhr)
🎥 IT Security VideoBuild apps that talk with Firebase AI Logic text-to-speech(17.09.2026 um 18:20 Uhr)
🍏 iOS / Mac OSOnline-Audio erreicht drei Viertel der Bevölkerung(17.09.2026 um 18:28 Uhr)
🍏 iOS / Mac OSApple co-founder Woz launches new merch store(17.09.2026 um 18:30 Uhr)
🎥 Video | YoutubeFlutter: Primary constructors and private named parameters(17.09.2026 um 18:00 Uhr)
🎥 Video | YoutubeGoogle Workspace: Generate and edit graphics with Google Pics(17.09.2026 um 18:17 Uhr)
🎥 Video | YoutubeAndroid Developers: How Android Bench 2.0 pushes AI evaluations(17.09.2026 um 18:36 Uhr)
🎥 IT Security VideoBuild apps that talk with Firebase AI Logic text-to-speech(17.09.2026 um 18:20 Uhr)
🍏 iOS / Mac OSOnline-Audio erreicht drei Viertel der Bevölkerung(17.09.2026 um 18:28 Uhr)
🍏 iOS / Mac OSApple co-founder Woz launches new merch store(17.09.2026 um 18:30 Uhr)
💾 Downloads 🕛 vor 21 Min. 1 Min Lesezeit
0

GitHub Release: ollama/ollama v0.34.2-rc2 (17.09.2026)

↗ Quelle (GitHub · ollama/ollama)
🗣️ Stimme:
GitHub Release: ollama/ollama v0.34.2-rc2 (17.09.2026)
Avatar
$ git clone https://github.com/ollama/ollama.git

The decode loop releases MLX's pool of freed buffers every 256 generated

tokens, which is also how often the KV cache grows and drops its previous,

smaller buffers. The check fires only when the token count lands exactly on

a multiple of 256. Speculative decoding emits several tokens per round, so

most rounds step over the boundary and the pool is never released. Each

growth at a long context leaves several GB of buffers that no later

allocation can reuse, so the runner's footprint keeps climbing over a long

generation until the system runs out of memory.


We now release the pool whenever a round crosses a multiple of 256 tokens,

which is what a single-token round already did. With qwen3.8:27b-mlx at a

98k-token context on a 128 GB machine, a long speculative generation

previously grew the runner past 90 GB and panicked the kernel; it now stays

flat at 30 GB.

Vollständiger Original-Artikel
Den kompletten Beitrag mit allen Details direkt auf github.com lesen.
↗ Original-Artikel auf github.com lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Avision AD7100 & AD7100N - Dreifach kontrolliert gegen Doppelblätter und Papierstau
1 Quelle
Windows Server 2022: Mainstream-Support endet am 13. Oktober - ad-hoc-news.de
1 Quelle
Lenovo ThinkAgile VX850 V4: Neue Infrastruktur für KI und Virtualisierung - ad-hoc-news.de