We have SWE-bench for coding and GAIA for reasoning. We have BFCL for function calling and LoCoMo for long-term memory. But ask a simple question — can the agent remember its own name after context compaction? — and no benchmark has an answer.
The benchmarks we have test impressive things: resolving real GitHub issues, navigating websites,...
🛡️ VERIFIED CYBER INTELLIGENCE ID: #3323044