The one-line version We ran five current frontier models over a set of documented-failure questions, twice each: bare, and wrapped in a thin external layer (retrieved evidence + a rule that lets the model say "I don't know"). We were not trying to make them smarter. We were asking whether confident fabrication can be reduced from outside the model... Weiterlesen
Intelligence View
We didn't make the models smarter. We built the thing that catches them confidently wrong — and it caught us too.
The one-line version We ran five current frontier models over a set of documented-failure questions, twice each: bare, and wrapped in a thin external layer (retrieved evidence + a rule that lets the model say "I don't…
SOCIAL SHARE CARD GENERATOR