Human evaluation is the gold standard for LLM output quality. It is also the bottleneck that kills every scaling plan.

One human reviewer processes 50-100 examples per hour. A single model comparison across 1,000 test cases takes 10-20 hours of human labor. Run that across 5 metrics and 3 model candidates, and you are looking at weeks of work...