I wanted to answer one question:


After packed-codebook TurboQuant failed, was there still a credible latency path?


The short answer:


there was a real speed ceiling, but no stable quality-preserving implementation path.





TL;DR



Hardware-friendly int4 K/V passed byte gates but failed real-KV logit quality.
Qwen2.5-7B work reduction...