Semantic caching promises to make AI systems faster and cheaper by reducing duplicate calls to large language models (LLMs). But what happens when it doesn’t work as expected?

We built a test environment to find out. Through a caching system, we evaluated how semantically similar queries would behave. When the cache worked, response times were...