Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
Windows Tipps & SecurityGetting Repeated No Caller ID Calls? Here’s What’s Really Going On(22.09.2026 um 22:31 Uhr)
Windows Tipps & SecurityHöllenmaschine: Gaming-Peripherie für gut 1.800 Euro für die HMX 6(23.09.2026 um 10:20 Uhr)
Windows Tipps & SecurityDas nächste große Ding: KI-Agenten(23.09.2026 um 10:30 Uhr)
Sichere ProgrammierungHow AI Is Making Restaurant Menus Easier to Navigate(23.09.2026 um 10:55 Uhr)
Windows Tipps & SecurityGetting Repeated No Caller ID Calls? Here’s What’s Really Going On(22.09.2026 um 22:31 Uhr)
Windows Tipps & SecurityHöllenmaschine: Gaming-Peripherie für gut 1.800 Euro für die HMX 6(23.09.2026 um 10:20 Uhr)
Windows Tipps & SecurityDas nächste große Ding: KI-Agenten(23.09.2026 um 10:30 Uhr)
Sichere ProgrammierungHow AI Is Making Restaurant Menus Easier to Navigate(23.09.2026 um 10:55 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Beyond automation: How much does AI really cost?

The problem nobody budgeted for An anonymous enterprise recently spent $500 million in a single month on Claude AI — not because the technology failed, but because nobody set usage limits before rolling it out to employees. Uber e…

0
↗ Quelle (cio.com)
Reagiere als Erste:r — dein Feedback zählt!








The problem nobody budgeted for





An anonymous enterprise recently spent $500 million in a single month on Claude AI — not because the technology failed, but because nobody set usage limits before rolling it out to employees. Uber exhausted its entire AI budget for 2026 before the first half of the year ended. JPMorgan published a report titled “AI Token Costs Are Eating into Internet Profits.” Shopify, Spotify, ServiceNow and Roku all cited AI as a major source of operational expense pressure in recent earnings calls.





This is not a technology problem. It is a cost modelling problem.





Most organizations ask the right first questions: What work should be AI-enabled? Which deployment approach fits each domain? But there is a third question that is almost never asked before launch: How much will it cost to operate this at scale?





The answer requires understanding three parameters simultaneously — and the interaction between them is deeply counterintuitive.





The deployments that did not produce budget surprises shared one characteristic: token volume was modelled per workflow type before the architecture was finalized.





The 3-parameter cost model





AI operational cost is not simply a function of how complex or sophisticated the task is. It is the product of three variables:





Total AI Cost = Tokens (activity) × Frequency (repetitions) × N (users)





Tokens(activity) measures the cognitive depth of a single session — how much input and output the AI processes to complete one instance of the task.





Frequency(repetitions) measures how often that activity is executed — daily, weekly, per transaction, per customer interaction.





N(users) measures how many individuals or automated processes are executing that activity across the organization.





The critical insight is that these three parameters behave in opposite directions depending on where the work sits in the T–R–M framework — and that inversion is what produces the budget surprises.





A brief recap: The T–R–M framework





In a previous article in this series, we introduced the T–R–M framework as a structured way to analyze how work is internally composed across three dimensions: Task nature (T), Relational density (R), and Human–AI operational mode (M).





The M dimension — human–AI operational mode — describes how work is distributed between humans and AI, ranging from full automation (M0) to human-dominant work where AI has no viable operational role (M4). Most professional roles operate across multiple M modes simultaneously within the same week.





What the framework did not yet address is the economic consequence of that distribution at scale. That is what this article adds.





Token ranges by operational mode





Each M mode has a characteristic token consumption profile per session. These ranges reflect the cognitive depth of the interaction — but they tell only one third of the story.





ModeLabelTokens / sessionFreq. / user / monthCost driver
M0Fully Autonomous AI1,000 – 8,000Hundreds–ThousandsN users × frequency
M1Supervised AI8,000 – 30,000Tens–HundredsVolume at scale
M2Hybrid Chain20,000 – 60,00010–50Collaboration depth
M3Extended Cognition50,000 – 120,000+2–10Session intensity
M4Human-DominantMinimal / zero1–5Negligible




Table 1. Estimated token consumption per session by Human–AI Operational Mode, with scale and cost driver characteristics.





The apparent paradox is immediate: M3 (Extended Cognition) consumes the most tokens per session, yet Goldman Sachs estimates that agentic AI — operating primarily in M0 and M1 — may increase total token demand by 24 times current levels. The reason is the multiplier effect of frequency and users.





An M0 task consuming 5,000 tokens per execution, running 500 times per day across 1,000 users, generates 2.5 billion tokens per month. An M3 session consuming 80,000 tokens, executed 4 times per month by 15 senior professionals, generates 4.8 million tokens. The ratio is roughly 500 to 1 — in favour of the task that costs less per session.





Profile 1: The business relationship manager





In the previous article, we followed a Business Relationship Manager through a single Tuesday. By Friday she had produced one prioritized backlog, two stakeholder briefings, three escalation memos, a renegotiated SLA, and a verbal commitment that quietly reshaped Q3 priorities for forty engineers.





Decomposed through T–R–M, that single week operated simultaneously across M0, M1, M2, M3, and M4. Applying the three-parameter cost model to each layer reveals a profile that is almost the inverse of what most organizations assume when they deploy AI for this role.





ActivityM ModeTokens / sessionSessions / monthUsers (org)Monthly cost index
Consolidating intake ticketsM03,000 – 8,000~200500+🔴 Very high
Drafting status briefingsM110,000 – 25,00040200🟠 High
Translating needs → requirementsM225,000 – 50,0002050🟡 Medium
Alignment in steering meetingsM350,000 – 100,000810🟡 Medium
SLA renegotiation post-incidentM4Minimal25🟢 Low
Hallway verbal commitmentsM4Zero1🟢 Negligible




Table 2. Token economics model for the Business Relationship Manager profile. ‘Monthly cost index’ is qualitative — relative budget exposure across activity layers.





The insight is not that M0 is too expensive to deploy — it is often the layer with the clearest ROI. The insight is that organizations routinely model the cost of M0 as if it were one user running one query. The actual cost is the product of all three parameters. For a BRM function deployed across a 500-person organization, the ticket consolidation layer alone can represent most of the total AI budget for that role.





Meanwhile, the steering meeting preparation — the M3 layer where the BRM synthesizes competing stakeholder positions, interprets political dynamics, and formulates negotiation strategy — consumes high tokens per session but runs infrequently and serves a small number of senior professionals. Its contribution to total cost is comparatively modest.





Organizations consistently overestimate the cost of the work AI does best and underestimate the cost of the work it does most.





Profile 2: The senior consultant





A senior consultant in a professional services firm operates across a different but structurally comparable T–R–M profile. The mix shifts toward M2 and M3 — more cognitive depth per session, lower frequency, smaller user population — but the same three-parameter logic applies.





ActivityM ModeTokens / sessionSessions / monthUsers (firm)Monthly cost index
Translation (short docs)M18,000 – 20,00015–20300🟠 High
Document analysisM1–M215,000 – 40,0008–10200🟡 Medium
Deliverable creationM220,000 – 60,0004–6100🟡 Medium
RFP analysis + Excel sim.M225,000 – 70,0002–450🟡 Medium
Code / automationM2–M325,000 – 80,0003–580🟡 Medium
Framework developmentM350,000 – 120,000+2–410–20🟢 Low at scale
Strategic negotiationM4Minimal1–35🟢 Negligible




Table 3. Token economics model for the Senior Consultant profile. Framework development sessions (M3) are the most token-intensive per session but the least significant at organizational scale.





Two observations stand out. First, translation — often dismissed as a low-cost commodity task — becomes a significant budget line when deployed at scale across a multilingual firm. A translation layer running 15–20 sessions per month per consultant, across 300 consultants, is not a negligible cost. It is a manageable one, but it must be modelled explicitly.





Second, framework development and strategic reasoning — the M3 activities that generate the highest per-session token consumption — are also the activities with the smallest user population and lowest frequency. Firm-wide, they may represent a smaller budget line than routine document analysis, even though each individual session costs significantly more.





The counterintuitive conclusion





ModeCost per sessionScale (users × freq)True budget risk
M0–M1LowMassive🔴 Primary risk
M2MediumModerate🟡 Manageable
M3HighMinimal🟢 Contained
M4NoneIrrelevant✅ No risk




Table 4. The budget risk paradox. The activities that consume the most tokens per session carry the least organizational budget risk. The activities that consume the least tokens per session carry the most.





This has direct implications for how organizations structure their AI governance. Cost controls applied uniformly across all AI usage — token caps, usage limits, model downgrades — will disproportionately affect M3 users, who are typically the professionals generating the highest-value outputs, while leaving largely untouched the M0–M1 volume that drives the actual budget exposure.





Effective AI cost governance requires mode-aware controls: different token budgets, model tiers, and usage policies calibrated to the M mode of the activity, not to the role title of the user.





Three implications for the CIO






  1. Model before you deploy. Before finalizing the architecture for any AI initiative, estimate token volume per workflow type — not per user, but per execution, multiplied by realistic frequency and user count. This calculation takes hours, not weeks, and it is the single most effective cost governance intervention available before deployment.




  2. The budget risk is at the bottom of the stack, not the top. If you need to contain AI spend, look first at M0 and M1 deployments: agent automation, document processing, content generation at scale. These are where token budgets are most likely to be exceeded. Your senior professionals running M3 sessions are almost certainly not your cost problem.




  3. Uniform limits are the wrong instrument. Token caps applied equally across all users will restrict your highest-value AI interactions while leaving your highest-volume interactions — the actual cost drivers — largely unaffected. Cost governance should be calibrated to operational mode, not to headcount.





This article is published as part of the Foundry Expert Contributor Network.
Want to join?


Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Beyond automation: How much does AI really cost?

Thematisch verwandte Begriffe: Beyond, automation, much, does · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-96258 | A vulnerability has been found in onSite internet GmbH Auktion NG Auktio…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick