🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 3 Min Lesezeit
0

The AI Feature Is Cheap to Build and Expensive to Run

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

The quote everyone remembers is the build cost. The number that decides whether an AI feature survives is the monthly one, and it tends to show up in month two, right when the trial credits run dry and real traffic arrives.



I budget AI features the way I'd budget a delivery van. Buying it happens once. Fuel, insurance, and the driver run forever. Here's where the fuel actually hides.






Where the money goes



Tokens, including the ones you forget. Everyone counts the user's question. Fewer people count the system prompt, the retrieved context, the few-shot examples, and the model's own output, all billed on every call. A feature carrying a 3,000-token context that looked tiny in testing can run 10x the estimate once every request drags that prompt along.



Retries and retrieval. A retry on failure doubles the cost of that call. A RAG feature also pays to embed every document, store the vectors, and run a similarity search per query. The model bill is one line on a longer receipt.



The machinery around the model. Vector database hosting. Logging and observability, which for AI features is not optional. Egress. The cache you'll add later to stop paying twice for the same answer.



Humans in the loop. If a person reviews flagged outputs, that review time is a running cost of the feature and belongs in the budget, even though no vendor ever invoices you for it.






How we actually budget it



We estimate a cost per action before a line of the feature exists. Average tokens in, average tokens out, times the model's price, times expected volume. It's back-of-envelope, and it usually lands close, because the inputs are knowable.



Then we pick the cheapest model that passes evaluation, not the highest one on the leaderboard. A smaller model that's good enough on your real task can cut the bill 5 to 10x. We send the easy 80% of requests to the cheap model and escalate only the hard ones. Caching repeat queries shaves off another slice.



The last step is a hard spend cap wired in before launch. Per user, per day, per feature. A runaway loop or a scraper pounding your endpoint should trip a limit and page a human, not keep billing until the card declines.






Give the client the real number



When we scope AI work, the client gets both numbers. Build once, and run monthly at your expected volume, with the assumptions written down beside them. A client who signed off on a $600-a-month running cost stays calm when the bill reads $600. A client shown only the build price feels ambushed, and they're right to.



That conversation is unglamorous, and it's the one that stops a project from souring six weeks in. The team at Shanti Infosoft treats the running-cost estimate as part of the quote, not a thing we discover together later, and you can see how we scope it at .



What did an AI feature actually cost you to run each month, and how far off was the first estimate?

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage