The general shape of the problem is that every public LLM benchmark is on a saturation clock that runs from the moment of its publication to the moment a model's training corpus has eaten it. The clock has been running, on the visible benchmarks of the last five years, for somewhere between twelve and thirty months before each one is no longer useful for differentiating frontier models. The benchmarks are not failing. They are doing exactly what they were designed to do, in the order they were designed to do it, and the field has been running through them faster than the people designing them anticipated.
I want to put numbers on the saturation pattern, walk through what the contamination evidence actually says, and then sit with the question of what an honest benchmark would have to look like in 2026 — because the "private held-out eval" answer that the labs are converging on has economics that are worth examining carefully before any of us salute it as the solution.
The saturation timeline, with numbers
(Hendrycks et al., September 2020). 57 subjects, ~14,000 multiple-choice questions, taken from publicly-available test prep and academic sources. The problem with MMLU is not that it's saturated in the same way HumanEval is — top scores are in the high 80s rather than against the ceiling — but that the benchmark was built from public sources that ended up in training corpora. The contamination evidence is concrete: a accepted at ACL 2025, was a contamination-free reconstruction; on it, model rankings shift considerably. Lifespan from publication to demonstrated contamination: about 36 months.
and now recommends SWE-bench Pro (1,865 multi-language tasks, not in the same training-corpus blast radius). Lifespan from Verified's August 2024 publication to OpenAI walking away from it in February 2026: about 18 months.
(Epoch AI, November 2024). The benchmark designed explicitly to resist saturation: tiers 1–3 cover undergraduate through early-postdoc mathematics, tier 4 is research-level. Hundreds of original problems, vetted by working mathematicians, never published in answerable form. At launch in late 2024, no tested model exceeded 2% on the full benchmark. By the end of 2025, frontier reasoning models were solving substantial fractions of tiers 1–3, and Epoch's own framing changed from "a benchmark current AI cannot do" to "a benchmark current AI is starting to crack." Lifespan from publication to first significant scores: about 12 months.
and (LMSYS) attempts the user-generated test items property — every prompt comes from a real user interaction, so the test distribution is open-ended and not authorable in advance. Each gets one or two of the falsifiability properties above. None of them gets all six.
The summary that matches the data
If I tabulate what's actually visible in the saturation evidence, the picture is unambiguous.
| Benchmark | Published | Top score 2025–2026 | Saturation lifespan | Primary contamination concern |
|---|---|---|---|---|
| HumanEval | 2021-07 | 96.3% (o1-preview) | ~36 months | Direct: 164 problems publicly indexed since release |
| MMLU | 2020-09 | mid-90s | ~36 months | Direct: documented test-slot reproduction in 2023 |
| GPQA Diamond | 2023-11 | 94.1% (Gemini 3.1 Pro Preview) | ~30 months | Indirect: scientific literature is in training corpora |
| SWE-bench Verified | 2024-08 | 93.9% (Claude Mythos Preview) | ~18 months (OpenAI walked away Feb 2026) | Indirect: training corpus includes the source repos |
| FrontierMath | 2024-11 | non-trivial fraction by end-2025 | ~12 months to first signal | Designed against direct contamination; indirect risk via mathematics literature |
| ARC-AGI-2 | 2025-05 | 84.6% public / 24% Kaggle-constrained | 12 months and counting | Designed to resist scaling; the public-vs-constrained gap is the data point |
A few things stand out reading this table. The lifespan column is shrinking. The "top score" column has hit the high 80s or above on every benchmark in the table that isn't FrontierMath or ARC-AGI-2 — and even those are starting to move. The contamination-concern column has no row that's clean; even benchmarks designed against direct contamination inherit indirect contamination from the source domain.
What this means for reading benchmark numbers
The published number on a public benchmark is informative for a specific window after publication and roughly noise after that window closes. HumanEval and MMLU and GPQA Diamond are at the ceiling. FrontierMath and ARC-AGI-2 are still informative, and won't stay informative for as long as their predecessors did.
The honest reading of any 2026 frontier-model release is to look at which benchmarks the lab is reporting and which it is conspicuously not. OpenAI's silence on SWE-bench Verified is more informative than any number OpenAI is still publishing. Labs that report across the full saturated slate are doing well on benchmarks they know to be saturated; labs emphasising FrontierMath, ARC-AGI-2, or in-house held-out evals are differentiating on harder ground. The signal is in the choice.
A benchmark is only useful for as long as it's hard, and hard is the gap between the benchmark's source distribution and the model's training distribution — a gap that shrinks with every new corpus. The appropriate posture toward any score is to ask three things: the benchmark's age, the contamination evidence, the spread among the top ten models. Read together, they tell you whether the headline is information or wallpaper.
SOCIAL SHARE CARD GENERATOR