Web TippsUse custom web fonts in Google Sheets charts(08.09.2026 um 17:05 Uhr)
Web TippsIntroducing the new 1Password App for Google Chat(08.09.2026 um 18:02 Uhr)
Web TippsUse custom web fonts in Google Sheets charts(08.09.2026 um 17:05 Uhr)
Web TippsIntroducing the new 1Password App for Google Chat(08.09.2026 um 18:02 Uhr)

🔧 Programmierung 🕛 vor 1 Monat 2 Min Lesezeit
0

Has anyone else run into this problem while building an AI SaaS?

↗ Quelle (dev.to)
🗣️ Stimme:

You launch your product, a few customers start using it, and then you realize...



How do you know which tenant is spending your OpenAI/Anthropic/Gemini budget?



And even if you know, how do you stop them before they burn through hundreds of dollars?



I looked around, and most tools focus on observability—they tell you what happened after the API call. I wanted something that could also enforce spending limits before the next request goes out.



So over the past few weeks, I built token-limit, a Python SDK that:



✅ Monkey-patches the official LLM SDKs (OpenAI, Anthropic, Gemini, DeepSeek, OpenRouter)



✅ Automatically tracks every LLM request per tenant



✅ Batches usage events in the background (no changes to your LLM call sites)



✅ Lets you set per-tenant spending limits and raises a LimitExceededException before another paid request is sent



The setup is intentionally simple:




CODE
from token_limit import Meter, MeterConfig

meter = Meter(MeterConfig(api_key="..."))
meter.patch_all()

with meter.for_tenant("acme"):
client.chat.completions.create(...)






I'm still actively improving it, and I'd genuinely love feedback from people building AI products.



A few questions:




  • How are you tracking LLM usage today?

  • Do you enforce budgets per customer, or just monitor costs?

  • Is monkey-patching something you'd be comfortable using in production, or would you prefer another approach?

  • Which provider or framework would you want to see supported next?



If you'd like to try it or review the implementation, here's the repo:



GitHub: https://github.com/AliEzatyar/token-limit



I'd really appreciate any feedback—positive or critical. If there's something that makes you think, "I wouldn't use this because...", I'd especially like to hear it.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
Use custom web fonts in Google Sheets charts
2 Quellen
Introducing the new 1Password App for Google Chat
1 Quelle
Context-aware access controls are available for Gemini Enterprise in the Admin console
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Has anyone else run into this problem while building an AI SaaS?

Thematisch verwandte Begriffe: anyone, else, into, this · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...