I've been running model showdowns on Vibes Coder for a while now. Each round has been a little messier than I wanted — different prompts, accidental context leaks, no clean way to compare cost to quality. This one is the first I'd call a fair bakeoff. Two goals going in:



Make the experiment itself rigorous enough that future rounds can build on...