As AI agents become more sophisticated and deploy across critical business functions, teams face a fundamental challenge: how do you thoroughly evaluate agent performance when real-world data is scarce, expensive, or privacy-constrained? This data scarcity problem becomes especially acute when testing edge cases, rare scenarios, or newly designed...