Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
🔒
https://machinelearning.apple.com
«LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quan...»
Automatische Weiterleitung...
1.5s