The latency barrier in speech-to-speech and video generation has shattered, enabling end-to-end conversational AI that rivals human cadence within 150ms while cloning voices from seconds of reference audio. xAI's Grok Imagine rolled out 10-second video generation with synchronized high-fidelity audio, following FlashLabs' open-sourced Chroma 1.0—a...