🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
Web TippsGoogle Workspace Weekly Recap - September 11, 2026(11.09.2026 um 21:19 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenVorsicht: Android-Malware verschlüsselt Ihre Handys und nimmt heimlich Fotos auf(11.09.2026 um 09:35 Uhr)
🕵️ SicherheitslückenMicrosoft geht endlich eines der nervigsten Probleme von Windows 11 an(11.09.2026 um 11:58 Uhr)
💾 IT Security ToolsSysinternals Suite(11.09.2026 um 12:00 Uhr)
🕵️ SicherheitslückenDefender 0-Day ShieldBreak (CVE-2026-69414) nicht sauber gepatcht - BornCity(11.09.2026 um 12:52 Uhr)
🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
Web TippsGoogle Workspace Weekly Recap - September 11, 2026(11.09.2026 um 21:19 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenVorsicht: Android-Malware verschlüsselt Ihre Handys und nimmt heimlich Fotos auf(11.09.2026 um 09:35 Uhr)
🕵️ SicherheitslückenMicrosoft geht endlich eines der nervigsten Probleme von Windows 11 an(11.09.2026 um 11:58 Uhr)
💾 IT Security ToolsSysinternals Suite(11.09.2026 um 12:00 Uhr)
🕵️ SicherheitslückenDefender 0-Day ShieldBreak (CVE-2026-69414) nicht sauber gepatcht - BornCity(11.09.2026 um 12:52 Uhr)

🔧 Programmierung 🕛 vor 2 Monaten 3 Min Lesezeit
0

Has Anyone Measured How LLM Output Quality Degrades Across Multiple Compactions?

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht




The Observation



After ~70 sessions with DeepSeek V4 (1M context), I noticed something odd. When Claude Code compacts my session, output quality doesn't just go down linearly. There's a moment — usually after the second compaction — where the model briefly gets better. Then it declines and never recovers.



Maybe I'm imagining it. Maybe it's specific to my model, my prompts, my workflow. But I can't shake the thought: what if context compaction has a curve, and nobody has mapped it?






What I Found (Not Much)



I searched for benchmarks that measure multi-round compaction degradation. Here's what exists:





  • RULER: Measures how performance drops as static input grows longer. Nothing about what happens after you compress and re-compress.


  • Context Rot (Chroma 2025): 18 models tested, all degrade with more tokens. Again, static.


  • Multi-turn evaluation: Tests whether models drift across conversation turns. Doesn't touch compaction.



Parameter compression (pruning, quantization) has well-mapped scaling laws. The Lottery Ticket Hypothesis (ICLR 2019) and Compression Laws for LLMs (2025) tell you exactly where the performance peak sits. Context summarization — the thing that happens every time your agent runs /compact — has no such curve.






Why This Might Matter



If the curve is real, you could:




  • Know exactly when to start a fresh session (before the decline hits)

  • Compare models on a new dimension: who maintains quality longest across compactions?

  • Give LLM providers a concrete target: "your compaction quality drops 20% faster than competitor X"



Right now, none of the major benchmark suites (MMLU, HELM, BigBench, RULER) include a "compaction persistence" metric. If context windows keep growing and sessions keep getting longer, this gap gets bigger every year.






What I'm Asking



I built a tiny monitor ( — 50 lines of Python, 10 benchmark tasks, 0-5 rubric. It's not polished. It's a starting point.



What I'd love:




  1. Someone with a Claude Opus / GPT-5 / Gemini account to try reproducing this

  2. Feedback on whether the methodology makes sense or is fundamentally flawed

  3. If this is a real thing, ideas for how to measure it properly



I don't have the compute or the stats background to do this alone. But if enough people contribute data points across different models, we might find out whether this curve exists — and if it does, maybe it's useful to more people than just me.






References




  • Frankle & Carbin, "The Lottery Ticket Hypothesis" (ICLR 2019)

  • "Compression Laws for Large Language Models" (2025)

  • RULER: What's the Real Context Size of Your Long-Context Language Models? (COLM 2024)

  • Chroma Research, "Context Rot" (2025)

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
How to Evaluate Live & Voice Agents in ADK
1 Quelle
Autonomous LLM post-training with Tunix on TPUs
1 Quelle
Lawyer fined $5K over AI-hallucinated witnesses in a murder case
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Has Anyone Measured How LLM Output Quality Degrades Across Multiple Compactions?

Thematisch verwandte Begriffe: Anyone, Measured, Output, Quality · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...