Your RAG pipeline retrieves 12 chunks because the retrieval score said "maybe". Your LLM reads all of them. You pay for all of them. And the answer quality was decided by chunks 2 and 7 anyway. On September 29, OpenAI launched the Decisions API built on Luna, and on September 15, TypeSafe launched Jev. Both are fast decision models for exactly... Weiterlesen: Cutting 70% of RAG context tokens and keeping the answers identical (…
Intelligence View
⚡ tsecurity.de Intelligence
Cutting 70% of RAG context tokens and keeping the answers identical (measured)
Your RAG pipeline retrieves 12 chunks because the retrieval score said "maybe". Your LLM reads all of them. You pay for all of them. And the answer quality was…