Here's something everyone agrees on about few-shot prompting: give the model more examples, it performs better.

I believed that too. Then I measured it.

So I built AdaptGauge, an open-source tool that measures how efficiently LLMs learn from few-shot examples.







What I tested


I evaluated eight models across four tasks designed to...