🔧 ProgrammierungGitHub Release: rust-lang/rust v1.98.1 (03.09.2026)(03.09.2026 um 15:14 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.5 (11.09.2026)(11.09.2026 um 07:07 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.6 (11.09.2026)(11.09.2026 um 08:08 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.7 (11.09.2026)(11.09.2026 um 10:17 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.8 (11.09.2026)(11.09.2026 um 12:30 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.9 (11.09.2026)(11.09.2026 um 14:01 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.10 (11.09.2026)(11.09.2026 um 17:56 Uhr)
🔧 ProgrammierungGitHub Release: rust-lang/rust v1.98.1 (03.09.2026)(03.09.2026 um 15:14 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.5 (11.09.2026)(11.09.2026 um 07:07 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.6 (11.09.2026)(11.09.2026 um 08:08 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.7 (11.09.2026)(11.09.2026 um 10:17 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.8 (11.09.2026)(11.09.2026 um 12:30 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.9 (11.09.2026)(11.09.2026 um 14:01 Uhr)
🔧 AI Nachrichten GitHub Release: openai/codex vrust-v0.155.0-alpha.3.10 (11.09.2026)(11.09.2026 um 17:56 Uhr)

🎥 Video | Youtube 🕛 vor 49 Min. 3 Min Lesezeit
0

DigitalOcean: Does context length affect inference cost linearly? We measured it

↗ Quelle (YouTube · DigitalOcean)
🗣️ Stimme:
📺
YouTube · DigitalOcean
10 YouTube-Aufrufe
Every AI pricing page treats context length as linear: ten times the tokens, ten times the cost. I rented one H200 GPU Droplet, served Ministral 3 14B with vLLM, and ran the same test from 2,000-token prompts up to 256,000. The memory cost per token stayed exactly 160 KiB the whole way, and the effective cost per million tokens still ended up 3.84x higher. This video shows where that gap comes from.

The GPU has one fixed pool of memory for prompts, about 637,000 tokens on this card. Every request parks its whole context there, so every doubling of context halves how many requests fit. At 2K, 311 users at once. At 256K, two. Same card, same hourly price, a quarter of the work per hour.

THE NUMBERS:

KV cache per token: 160 KiB, flat at every context length (the linear part is real)
Concurrent requests: 311 at 2K → 2 at 256K
Throughput: 19,089 → 4,967 tokens/sec
Effective cost: $0.065 → $0.25 per 1M tokens (3.84x)
Breakeven vs the flat $0.20/1M serverless rate: 33% GPU utilization at 2K, 125% at 256K. A GPU can't be 125% busy, so at that context the flat rate wins at any traffic level.

TL;DR: long context doesn't make tokens pricier. It makes your GPU serve fewer people at once, and a card billed by the hour charges you for that with zero errors and zero alerts. Your bytes per token are in the model's config. Your pool size is in the vLLM boot log. Divide them before you rent.

WHAT THIS DOESN'T CLAIM: every number here comes from one setup, Mistral 3 14B (BF16) on a single H200 with vLLM v0.27.1. The mechanism generalizes, the specific numbers don't. Prefix caching was off, so this is the no-reuse floor, and production traffic with repeated prefixes sees a smaller penalty. This measures the memory effect only. Prefill compute is a separate cost this test doesn't size. Nothing here says any provider's price is above or below their own cost, which isn't observable from outside. Pricing as of August 2026.

Full write-up with the methodology, raw per-request data
Blog: https://www.digitalocean.com/community/tutorials/does-context-length-affect-inference-cost-linearly

⏱️ TIMESTAMPS
0:00 Does long context cost linearly?
1:08 The setup: one model, one GPU, eight prompt lengths
1:43 The part the calculators get right
2:36 The parking garage: 311 requests fit
3:21 Okay, guess: prompts 128x longer
4:06 Why that costs you money
5:57 When serverless just wins
7:17 Why dashboards stay silent
8:17 Write-up and repo

🚀 Join the Developer Cloud: https://cloud.digitalocean.com/registrations/new

// STAY CONNECTED
🌏 Follow our blog for the latest updates: https://www.digitalocean.com/blog
🦈 Join our Developer Community on Discord: https://discord.com/invite/digitalocean
🐥 Follow us on X/Twitter: https://x.com/digitalocean
👩‍💻 We're Hiring! See open roles: http://grnh.se/aicoph1

#LLM #AIInference #vLLM #GPU #KVCache #LongContext #Mistral #DigitalOcean
Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf youtube.com.
↗ Original-Artikel auf youtube.com lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Keep the Rebel Spirit Alive #TheSAS2026 #kaspersky #cybersecurity
1 Quelle
Exploits and vulnerabilities in Q2 2026
1 Quelle
The Gemini desktop app is now available for Windows