🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🪟 Windows TippsIntegrated GPU is showing as Removable on Windows 11(14.09.2026 um 06:48 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenVorsicht: Android-Malware verschlüsselt Ihre Handys und nimmt heimlich Fotos auf(13.09.2026 um 09:30 Uhr)
🪟 Windows TippsIkea bringt neuen Bluetooth-Lautsprecher auf den Markt: Badkruka(14.09.2026 um 08:00 Uhr)
🪟 Windows TippsLinuxWelt Extra 3/2026 am Kiosk: Linux Grundlagen erklärt(14.09.2026 um 09:45 Uhr)
🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🪟 Windows TippsIntegrated GPU is showing as Removable on Windows 11(14.09.2026 um 06:48 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenVorsicht: Android-Malware verschlüsselt Ihre Handys und nimmt heimlich Fotos auf(13.09.2026 um 09:30 Uhr)
🪟 Windows TippsIkea bringt neuen Bluetooth-Lautsprecher auf den Markt: Badkruka(14.09.2026 um 08:00 Uhr)
🪟 Windows TippsLinuxWelt Extra 3/2026 am Kiosk: Linux Grundlagen erklärt(14.09.2026 um 09:45 Uhr)

🔧 Programmierung 🕛 vor 2 Monaten 3 Min Lesezeit
0

Title

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht




Title



vLLM PagedAttention KV Cache Corruption: Woke Up to This Nightmare

Image generated via Midjourney by author



But tbh I dont even know where to start. So like I was on call and at 8am my phone starts blowing up. Incident alert. Peak RPS was at 14720. And ngl I freaked out a bit. Because what even is that.



And so I jump into logs and see this crazy error message. Like wtf is going on here? But ok so lets get into it.




CODE
graph LR
A[vLLM] ,>|requests|> B[PagedAttention]
B ,>|cache query|> C[KV Store]
C ,>|corrupted response|> B
B ,>|error|> A






Because we use vLLM with paged attention for our model serving. And lol its been working great till now.



Edit: Wait, I was wrong about the batch size above. It was 32, not 16.



But today because of this cache corruption issue were seeing crazy errors like this:

$$tensor_shape = [B, S, H]$$

where $B$ is batch size $S$ is sequence length and $H$ is hidden size.



And when we try to access the KV store we get:




CODE
try:
kv_store.get(key)
except Exception as e:
print(f"Error accessing KV store: {e}")






And idk whats going on but the error log says:






Debugging Log






CODE
Traceback (most recent call last):
File "model_serving.py", line 123, in serve_model
response = paged_attention.query(cache_key)
File "paged_attention.py", line 45, in query
value = kv_store.get(cache_key)
File "kv_store.py", line 23, in get
raise Exception("Cache corruption detected")
Exception: Cache corruption detected






Because of this stupid cache corruption issue were down and idk how long its gonna take to fix.



But so first thing I did was jump into the codebase and start debugging. And lol its always something simple right?

So after hours of debugging we finally found the issue. It was a subtle bug in our cache eviction policy.



And now were pushing a fix and hoping itll resolve the issue.





What Didn't Work First



Before I found the real issue, I tried 3 other fixes that failed:





  1. Bumping timeouts - Changed NCCL_TIMEOUT=1800 in the env. Did nothing. Still failed at 8am.


  2. Restarting pods - kubectl rollout restart deployment/vllm. Came back up, same error. Wasted 10 mins.


  3. Checking GPU health - nvidia-smi showed all GPUs fine. I was convinced it was hardware tbh.



Spent 45 mins going down wrong paths. The fix was 1 line in Dockerfile. Im an idiot.





Monitoring We Added After



Because this sucked, we added 3 grafana alerts so Marcus never gets paged for this again:




CODE
# Alert if NCCL comms thread fails
rate(nccl_errors_total) > 0









CODE
# Alert if all_reduce latency > 50ms
histogram_quantile(0.99, nccl_allreduce_duration_seconds_bucket) > 0.05






Now if this breaks, pagerduty wakes us up before users notice.






FAQ Nobody Asked



Q: Why not use Gloo backend?

A: Gloo is slower. NCCL is 3x faster for all_reduce. Unless your network is trash.



Q: Could this happen on single-node?

A: No. This error only triggers multi-node. If you see this on 1 GPU, you have bigger problems.



Q: Do I need to update CUDA too?

A: Maybe. We were on 12.1. If you're on 11.8, upgrade everything or suffer.






Reproduce This



Full code: https://github.com/yourorg/voygr-vllm-pagedattention-kv-cache-corruption-debug






Distribution Taxonomy



Artificial Intelligence, Machine Learning, Data Science, Deep Learning, Programming, Software Engineering

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
2 Quellen
Why Companies Build Custom Photoshop Plugins (And When You Need One Too)
2 Quellen
Mastering Claude and ChatGPT: Developers Guide to Advanced Prompting
1 Quelle
SchemaCrawler LLM Context: Extract and Prune Relational DB Schemas
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Title

Thematisch verwandte Begriffe: Title · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...