LLM Self-Preference Bias: How Anonymized Peer Review Fixes It


The panel had been agreeing with itself for a week before I noticed, and the worst part is that the logs looked healthy the whole time.

I had built what felt like a clean idea. Several frontier models, different families, each one judging a pool of candidate outputs and ranking them...