Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
Sichere ProgrammierungWhy Claude Code keeps writing shell commands that fail on your Mac(20.09.2026 um 21:06 Uhr)
Sichere Programmierungllms.txt v2: What the Spec Says, and What 137,000 Domains Show(20.09.2026 um 21:17 Uhr)
Sicherheitslücken (CVE)NiceTryGPT: Less pattern matching. More actual hacking.(20.09.2026 um 21:19 Uhr)
IT Security VideoActivities BoF (kde2026)(20.09.2026 um 00:00 Uhr)
IT Security Toolsirdoc-app(20.09.2026 um 20:33 Uhr)
Sichere ProgrammierungWhy Claude Code keeps writing shell commands that fail on your Mac(20.09.2026 um 21:06 Uhr)
Sichere Programmierungllms.txt v2: What the Spec Says, and What 137,000 Domains Show(20.09.2026 um 21:17 Uhr)
Sicherheitslücken (CVE)NiceTryGPT: Less pattern matching. More actual hacking.(20.09.2026 um 21:19 Uhr)
IT Security VideoActivities BoF (kde2026)(20.09.2026 um 00:00 Uhr)
IT Security Toolsirdoc-app(20.09.2026 um 20:33 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Per-customer cost attribution without a proxy

Reagiere als Erste:r — dein Feedback zählt!

Most AI cost tracking solutions force you to route all your LLM traffic through their proxy. Tbh, that's an architectural nightmare waiting to happen. You're adding latency, introducing a single point of failure, and giving some third-party service the keys to your entire prompt stream.

If their proxy goes down, your app goes down. If their proxy gets slow, your users think your app is slow. And let's not even talk about the compliance headache of sending sensitive customer data through an intermediary just to track API costs.

You don't need a proxy to figure out which customer is burning your OpenAI budget. You just need proper attribution at the request level, handled asynchronously.

The Problem with LLM Billing

When you look at your billing dashboard on OpenAI or Anthropic, you just see total tokens used and a massive dollar amount at the end of the month. You don't see that user_123 ran a massive batch extraction job that cost you $40 in API calls, while your other 100 users cost $2 combined.

Multi-tenant SaaS apps need unit economics. If you charge a flat $20/mo subscription but a power user is burning $50/mo in Claude 3.5 API costs, you are actively losing money. But to fix it, you need to know exactly who is spending what.

Why Proxies Are a Bad Idea for This

A lot of dev tools in the AI space right now tell you to just swap your base URL from api.openai.com to proxy.theirservice.com.

Here is what happens when you do that:

  1. Every request adds 50-200ms of network overhead.
  2. If the proxy goes down, your production app fails to serve requests.
  3. You are sending raw PII and proprietary data to a vendor just to count tokens.

It's massive overkill. Cost tracking should be out-of-band. It should never be in the critical path of your application's request/response cycle.

The Async Logging Approach

The correct way to handle this is logging costs asynchronously after the request completes. Your app talks directly to the provider (OpenAI, DeepSeek, OpenRouter, Anthropic), gets the token usage from the response, and fires a background job to log it against the customer ID in your own database.

Here is the flow:

  1. User triggers an action.
  2. Your backend calls the LLM provider directly using their official SDK.
  3. Provider responds with the completion and usage stats (prompt_tokens, completion_tokens).
  4. Your backend returns the response to the user immediately.
  5. Your backend fires a non-blocking async event (e.g., using Inngest, BullMQ, or standard background workers) with the user ID, model used, and token count.

This gives you zero added latency. Zero third-party risk. Your app stays fast and reliable even if your cost-tracking database goes down.

Implementing the Calculation

Calculating the cost is straightforward but tedious. You need to maintain a pricing table for every model you support.

For example, if the payload from OpenAI says:
{ "prompt_tokens": 1500, "completion_tokens": 400 }

  1. The background worker calculates the cost based on current model pricing and writes it to your database.

This gives you zero added latency. Zero third-party risk. Your app stays fast and reliable even if your cost-tracking database goes down.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Per-customer cost attribution without a proxy

Thematisch verwandte Begriffe: Percustomer, cost, attribution, without · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-93956 | A flaw has been found in olivier-ls PHP-FTS up to 1.1.2. Affected by thi…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick