YouTube Video
The release of Android Bench 2.0 provides developers with a new dataset to evaluate AI coding assistance. We designed new long-horizon tasks to test models and agents on complex assignments — like creating applications from scratch, or library migrations, and introduced a nuanced scoring system that penalizes attempts to tamper with tests while rewarding adherence to requirements and visual fidelity.
Explore the long-horizon task results, and our benchmarking methodology to learn more about the benchmarking insights that can help you evaluate different AI models → https://d.android.com/bench
Subscribe to Android Developers → https://goo.gle/AndroidDevs
#Android #AndroidDevelopers #AndroidBench
Speakers: Zoe Lopez-Latorre
Products Mentioned: Android, Android Bench
Explore the long-horizon task results, and our benchmarking methodology to learn more about the benchmarking insights that can help you evaluate different AI models → https://d.android.com/bench
Subscribe to Android Developers → https://goo.gle/AndroidDevs
#Android #AndroidDevelopers #AndroidBench
Speakers: Zoe Lopez-Latorre
Products Mentioned: Android, Android Bench
SOCIAL SHARE CARD GENERATOR