Google Research published TurboQuant this week — a compression algorithm that reduces LLM Key-Value cache memory by 6× and delivers up to 8× attention speedup, with zero accuracy loss at 3 bits per channel.

The immediate reaction is straightforward: cheaper inference, faster generation, longer context windows. But the second-order effect is more...