You ship an LLM service. p50 latency looks great. Then a user pastes a 40-page contract into the chat, and for the next 400 milliseconds every other user's tokens stop arriving. Their streams freeze, then catch up in a burst. Your dashboards show inter-token latency spikes with no obvious cause. Nothing crashed. Nothing is rate-limited. One long...