How I normalize benchmark scores across chip generations (and why raw AnTuTu numbers lie)
Tags: #showdev #webdev #data #seo
The problem
Benchmark scores inflate over time. A top AnTuTu result in 2024 was around 2.4M points; in 2026 the leaders are pushing 4M. So a statement like "this phone scores 3.1M" is meaningless without context — is that flagship-level today, or last year's midrange?
It gets worse when you mix sources. AnTuTu measures the whole system, Geekbench isolates the CPU, 3DMark stresses the GPU, DxOMark rates cameras on a completely different scale. None of them are comparable to each other, and none are stable over time.
I run — two ratings side by side, with the raw benchmarks, specs and prices underneath.
The database updates daily. There are no sponsored placements — rankings are just the data.
What's next
More depth on sustained performance (throttling is the biggest gap between benchmark numbers and real experience) and richer battery data.
If you've dealt with normalizing noisy third-party data at scale — especially handling scale drift over time — I'd genuinely love to hear how you approached it.
SOCIAL SHARE CARD GENERATOR