An LLM benchmark is only useful for as long as it's hard
🔒
https://dev.to
«The general shape of the problem is that every public LLM benchmark is on a saturation clock that runs from the moment of its publication to the moment a model's training corpus has eaten it. The clock has been running, ...»
Automatische Weiterleitung...
1.5s