Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds ➔
👥 Community & Social
Sichere ProgrammierungWe Built a CLI to Find Out If You’re Overpaying for Claude(24.09.2026 um 04:35 Uhr)
•
Sichere ProgrammierungMy own sandbox was killing my agent's shell, and the exit code hid it(24.09.2026 um 04:38 Uhr)
••
Sichere ProgrammierungHow three OSLabs engineers built a CLI to catch you overpaying Claude(24.09.2026 um 04:45 Uhr)
•
Sichere ProgrammierungBreaking CI Guards on Purpose to Prove They Can Fail(24.09.2026 um 05:00 Uhr)
••••
IT Security NachrichtenLangfristige Updatefähigkeit als Pflicht(24.09.2026 um 05:03 Uhr)
••
Sichere ProgrammierungWe Built a CLI to Find Out If You’re Overpaying for Claude(24.09.2026 um 04:35 Uhr)
•
Sichere ProgrammierungMy own sandbox was killing my agent's shell, and the exit code hid it(24.09.2026 um 04:38 Uhr)
••
Sichere ProgrammierungHow three OSLabs engineers built a CLI to catch you overpaying Claude(24.09.2026 um 04:45 Uhr)
•
Sichere ProgrammierungBreaking CI Guards on Purpose to Prove They Can Fail(24.09.2026 um 05:00 Uhr)
••••
IT Security NachrichtenLangfristige Updatefähigkeit als Pflicht(24.09.2026 um 05:03 Uhr)
••
Intelligence View
⚡ tsecurity.de Intelligence

Our AI coding bill quietly tripled. Here's what we learned fixing it.

A few months ago I opened our cloud bill and had that small stomach drop moment every engineer knows. Our AI coding spend had roughly tripled. Not because anything was broken, but because everything was working. The team had gone all in on…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

A few months ago I opened our cloud bill and had that small stomach drop moment every engineer knows. Our AI coding spend had roughly tripled. Not because anything was broken, but because everything was working. The team had gone all in on Claude Code and Codex, they were shipping faster than ever, and nobody, including me, could say where the money was actually going.



I run engineering at a small software company. We're not Uber. But it turns out the shape of this problem is the same whether you have 15 engineers or 5,000, and the big players hit it first and hardest. Uber reportedly burned through its entire annual AI coding budget in four months. Meta's engineers pushed tens of trillions of tokens in a single month, partly chasing an internal usage leaderboard, and their own CTO had to point out that token usage isn't a measure of impact.



If it can happen to them, it can absolutely happen to you. Here's what actually went wrong for us, and the concrete things that brought it back under control.






The problem isn't the tools. It's the blindness.



The productivity gains from agentic coding are real. I'm not here to be a skeptic. The speedup was immediate and nobody wanted to go back. The problem is a specific and dangerous gap: you can't see the spend until the invoice arrives, and by then it's already spent.



When I finally put real visibility on our usage, the waste wasn't dramatic. There was no villain. It was ordinary, invisible, and constant:




  • One developer was quietly burning more tokens than five others combined. Not maliciously, just a workflow that leaned hard on the most expensive model for everything.

  • Routine tasks like doc generation, cleanup, and commit messages were hitting a frontier model when a model a fraction of the price would have produced identical output.

  • A test and fix loop someone kicked off had been running semi attended for the better part of two weeks.
    Individually, each of these is nothing. Together, they were the entire gap between "the savings AI promised" and the bill I was actually staring at.



The root cause came down to three blind spots:





  1. No usage visibility. The bill is one big number. You can't break it down by person, project, or task, so you can't reason about it.


  2. No control. Nothing stops a runaway process or a bad default until the money is gone.


  3. No model governance. The tools pick the model themselves. You pay for that choice with no say in it.
    ## What actually reduced the cost



Once I could see the problem, fixing it was mostly mechanical. Here's what moved the needle, roughly in order of impact.






1. Prompt caching (the biggest single lever)



This is the one most teams leave on the table, and it's the largest. Agentic coding re sends a huge amount of stable context every turn: the system prompt, tool definitions, the same files. With prompt caching, that repeated prefix reads at a fraction of full input price, and cached tokens don't count against your rate limits either.



The critical gotcha: if you route Claude Code through any kind of proxy, it's dangerously easy to break caching by not forwarding the right headers or by mutating the request body. A misconfigured gateway can silently turn every request back into full price. If you cache correctly, this alone can cut input costs dramatically.



Check: confirm your cache read token counts are non zero on repeat requests. If they're zero, caching isn't actually happening.






2. Right size the model per task



Not everything needs the frontier model. Doc generation, summaries, simple refactors: a cheaper model handles them at identical quality for a fraction of the cost. The spread is enormous. The same prompt across top models can range from cents to several dollars. Codex and Claude Code both have a "small fast model" concept for background tasks, so make sure that's actually pointed at a cheap model, not silently inheriting your expensive default.






3. Cap the loops



The single scariest line item is anything that calls a model in a loop unattended. Agent runs, test fix cycles, background jobs. Put a hard ceiling on them, a spend or token cap that stops the process rather than a memo asking people to be careful. Make it a config, not a policy document.






4. Trim what you send



Beyond caching, a lot of what gets sent to the model is stale context being dragged along turn after turn. Compressing bulky, old tool output before it hits the model can cut token volume substantially on long agent sessions, as long as you leave recent context and reasoning untouched. Be careful here. Aggressive trimming of active context breaks agentic workflows, so this is a "squeeze the old stuff only" move.






5. Make spend visible to the people spending it



This one is behavioral, and it's underrated. The moment developers could see their own per project usage, they self corrected. Nobody wants to be the person burning 5x the team on commit messages. Visibility alone changed behavior before any hard caps kicked in.






Where this led



I ended up building an internal gateway that sits between our developers and the model providers, so every request is attributed to a person and a project, checked against a budget, and routed to an appropriate model, with caching preserved and hard caps enforced. It eventually became a product (Clawgate), but honestly the approach matters more than the tool. You can assemble most of this yourself with an open source proxy like LiteLLM plus some discipline. The important part is that you stop flying blind.



The result for us was straightforward: same shipping speed, materially lower bill, and, maybe more importantly, I can now answer the question "what is our AI spend buying us?" instead of shrugging at it.






The takeaway



Agentic coding is worth every bit of the hype on the productivity side. But the cost side is the least governed line in most engineering budgets right now, and it's growing fast. You don't need to slow your team down to control it. You need to see it, attribute it, cache aggressively, right size the models, and cap the loops.



The gains are real. Just don't let them leak out the bottom as a surprise bill.






How are you all handling AI coding costs on your teams? Curious whether others are seeing the same "invisible until the invoice" pattern, and what's worked for you. Drop it in the comments.

SOC Incident Playbook: Remote Code Execution (RCE) Defense
title: Detect Exploitation - Our AI coding bill quietly tripled. Here's what we learned fixing it.
id: a589bef0-9e41-45bd-9b08-69949b658caf
status: experimental
description: Automatisch generierte SIEM-Erkennungsregel basierend auf CTI Intelligence
references:
  - https://tsecurity.de/
author: iShareStuff CTI Automated Detection Engine
date: 2026-09-24
logsource:
  category: network_connection
  product: any
detection:
  selection:
      CommandLine|contains:
        - 'exploit'
  condition: selection
falsepositives:
  - Legitime administrative Zugriffe oder Penetrationstests
level: high
tags:
  - attack.initial_access
rule CTI_Threat_Indicator {
    meta:
        author = "iShareStuff CTI Automated Detection Engine"
        date = "2026-09-24"
        description = "YARA Signature for "
    strings:
        $str = "Our AI coding bill quietly tri" ascii wide
    condition:
        any of them
}
tsecurity.de Cognitive Threat RAG
Fokus-Vektor:

Kognitive Analyse für identifizierte Bedrohung: Erhöhte Bedrohungslage im Bereich Our AI coding bill quietly tripled. Here.... Basierend auf 368k Vektor-Korrelationen werden sofortige Isolationsmaßnahmen für betroffene Endpunkte empfohlen.

🛡️ Angriffsfläche & Exposure

Netzwerk/Remote-Zugriff ohne Vorauthentifizierung möglich.

⚡ Empfohlene Sofortmaßnahmen
  • 1. Perimeter-Inspektion: Relevante Portfreigaben und exponierte Endpunkte unverzüglich scannen.
  • 2. Patch-Applikation: Hersteller-Hotfix einspielen oder betroffene Daemons in isolierte DMZ-Segmente überführen.
  • 3. Telemetrie & EDR-Alerts: Prozessaufrufe und Child-Processes auf anomale Shell-Spawns überwachen.
🔗 Semantisch verwandte Zero-Days MariaDB 11.7 VEC
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Our AI coding bill quietly tripled. Here's what we learned fixing it.

Thematisch verwandte Begriffe: coding, bill, quietly, tripled · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-96676 | A vulnerability was identified in Fast FAC1900R 20190827_2.0.2. The impa…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel • Rechts: nächster Artikel • unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel TTP ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger • Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick