📰 IT Security NachrichtenHuawei Shows Next-Generation Data Centre Optics Amid Standards Push(14.09.2026 um 09:01 Uhr)
📰 IT Security NachrichtenRobot Start-Up Skild AI Hits $100m Revenue Run-Rate(14.09.2026 um 09:31 Uhr)
📰 IT Security NachrichtenA week in security (September 7 – September 13)(14.09.2026 um 09:31 Uhr)
📰 IT Security NachrichtenIT Security News Hourly Summary 2026-09-14 10h : 8 posts(14.09.2026 um 10:00 Uhr)
📰 IT Security NachrichtenHackers Abuse AutoIt to Inject AsyncRAT Into Microsoft-Signed Windows Process(14.09.2026 um 10:02 Uhr)
📰 IT Security NachrichtenUK Military Spends Millions On SpaceX Satellite Services(14.09.2026 um 10:02 Uhr)
🔧 AI Nachrichten OpenAI, Claude and Other Chatbot Outages Reported(03.09.2026 um 20:00 Uhr)
📰 IT Security NachrichtenHuawei Shows Next-Generation Data Centre Optics Amid Standards Push(14.09.2026 um 09:01 Uhr)
📰 IT Security NachrichtenRobot Start-Up Skild AI Hits $100m Revenue Run-Rate(14.09.2026 um 09:31 Uhr)
📰 IT Security NachrichtenA week in security (September 7 – September 13)(14.09.2026 um 09:31 Uhr)
📰 IT Security NachrichtenIT Security News Hourly Summary 2026-09-14 10h : 8 posts(14.09.2026 um 10:00 Uhr)
📰 IT Security NachrichtenHackers Abuse AutoIt to Inject AsyncRAT Into Microsoft-Signed Windows Process(14.09.2026 um 10:02 Uhr)
📰 IT Security NachrichtenUK Military Spends Millions On SpaceX Satellite Services(14.09.2026 um 10:02 Uhr)
🔧 AI Nachrichten OpenAI, Claude and Other Chatbot Outages Reported(03.09.2026 um 20:00 Uhr)

🔧 Programmierung 🕛 vor 3 Monaten 3 Min Lesezeit
0

Memory beats full context on LongMemEval — and the wins we don't get

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

A common objection to agent memory is that you don't need it: context windows are huge now, so just put the whole history in the prompt. We wanted a real answer, not a vibe, so we ran two public long-term-memory benchmarks against a full-context baseline. Here's what we found — including the case where the baseline wins.






The setup



We compared two configurations on the same questions. The full-context baseline stuffs the entire conversation history into the prompt. Eidentic memory ingests the history into its four-tier engine and retrieves only what each question needs. Both use the same model and the same LLM judge. We ran the full sets — no sampling — and we're publishing wins and losses together.






LongMemEval: memory wins across the board



LongMemEval uses long histories — roughly 115k tokens across ~50 sessions, 500 questions. This is where memory should help, and it does: 55.2% overall vs 41.0% for full context, a 14.2-point gain, winning all six question types.
















































Question type Full context Eidentic memory
Single-session · user 67.1% 84.3%
Single-session · assistant 73.2% 92.9%
Single-session · preference 3.3% 26.7%
Multi-session 27.8% 42.1%
Temporal reasoning 20.3% 34.6%
Knowledge update 66.7% 70.5%
Overall 41.0% 55.2%


The cost difference is the other half of the story. Memory answers each question with about 2,550 tokens of retrieved context; the baseline spends about 99,435 re-reading the whole history every time — up to ~39× fewer tokens for the better score. Retrieval isn't just more accurate here, it's dramatically cheaper.






LoCoMo: where full context still wins



LoCoMo has a much smaller haystack. When the entire history comfortably fits in the window, brute force is hard to beat: the model can see everything at once, and single- and multi-hop questions don't need retrieval. Here the full-context baseline comes out 7.8 points ahead. Memory still uses far fewer tokens (~893 vs ~19,030), but on a small history that trade-off doesn't pay for itself on accuracy.




The larger the history, the more memory wins — on accuracy and on cost. On small histories, full context stays competitive. We'd rather you know both numbers than just the flattering one.







What this means in practice



If your agent's conversations are short and bounded, you may not need a memory engine at all — and we'll tell you that. But the moment histories grow past what you want to pay to re-read on every turn, retrieval-based memory wins twice: better answers, far fewer tokens. That crossover arrives quickly in real products.



Full methodology, the harness, and the raw per-question records are in the . Reproduce it, and tell us where we're wrong.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
2 Quellen
Microsoft bringt Emoji 17.0 auf Windows 11
1 Quelle
Neue Android-Malware schreit Sie an, wenn Sie nicht zahlen
1 Quelle
Handy: Wer diese App installiert hat, sollte sein Gerät besser zurücksetzen
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Memory beats full context on LongMemEval — and the wins we don't get

Thematisch verwandte Begriffe: Memory, beats, full, context · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...