I have read more voice-agent benchmarks than I would like to admit. They all measure the same thing: how many milliseconds from "user stops talking" to "agent starts talking." Stack comparisons, P95 charts, the whole genre. Every one of them treats the conversation as a relay race where only one runner is moving at a time.

Then I shipped...