YouTube Video
Why does a shorter question cost more on turn three? Language models are stateless, so every turn resends the system prompt and the full conversation history — the logs climb from 34 prompt tokens to 52 to 71 while the questions get shorter. Tool outputs, retrieved RAG documents and hidden metadata all bill as input tokens too, so they can dominate total spend even though input is the cheaper side.
► Full video: https://youtu.be/mB0IyELzjRg
► Get started: https://aka.ms/FoundryTokenomics
#Shorts #Tokenomics #AICosts #AzureAIFoundry
► Full video: https://youtu.be/mB0IyELzjRg
► Get started: https://aka.ms/FoundryTokenomics
#Shorts #Tokenomics #AICosts #AzureAIFoundry