Your agent just called three MCP tools to answer a single user question. One took 50ms, one took 2 seconds, one failed and retried.
You saw token usage spike. You have no idea which tool burned the most tokens or whether it was even necessary. One of those tools had permission to access customer data—did the agent call it? You're not sure. If compliance asks later, you've got no audit trail. And next month, when token prices change or a new cheaper model comes out, you're rewriting all your routing logic by hand.
That's the gap most teams hit when they move from demo agents to production MCP systems.
The MCP Governance Problem Is Real
In April 2026, the
LiteLLM routing benchmarks – 38% p95 latency improvement with intelligent routing
SOCIAL SHARE CARD GENERATOR