Zum Hauptinhalt springen
Echtzeit-Radar & Feeds
Alle RSS Feeds ➔
👥 Community & Social
••
Sichere ProgrammierungHow Hindsight Turned Deployment #1017 Into the Fix for #1057(29.09.2026 um 06:25 Uhr)
••
Sichere ProgrammierungSystems foundations should start below the framework(29.09.2026 um 06:30 Uhr)
••
Sichere ProgrammierungHow a game should say no: designing refusals people can act on(29.09.2026 um 06:30 Uhr)
••
Sichere ProgrammierungTuring Sim: Making the USD Schemas Inspector Readable(29.09.2026 um 06:31 Uhr)
••••
Sichere ProgrammierungHow Hindsight Turned Deployment #1017 Into the Fix for #1057(29.09.2026 um 06:25 Uhr)
••
Sichere ProgrammierungSystems foundations should start below the framework(29.09.2026 um 06:30 Uhr)
••
Sichere ProgrammierungHow a game should say no: designing refusals people can act on(29.09.2026 um 06:30 Uhr)
••
Sichere ProgrammierungTuring Sim: Making the USD Schemas Inspector Readable(29.09.2026 um 06:31 Uhr)
••
Intelligence View
⚡ tsecurity.de Intelligence

Over-editing is a token tax: GPT-5.4 ships 6.5x more diff per fix than Claude Opus 4.6, and your bill notices

A model is over-editing if its output is functionally correct but structurally diverges from the original code more than the minimal fix requires. Left unconstrained, the extended reasoning gives models more room to 'improve' code that…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

A model is over-editing if its output is functionally correct but structurally diverges from the original code more than the minimal fix requires. Left unconstrained, the extended reasoning gives models more room to 'improve' code that doesn't need improving.



GPT-5.4 averages 0.395 normalized Levenshtein distance per edit. Claude Opus 4.6 averages 0.060. That is 6.5x more output tokens for the same class of fix, averaged across the benchmark. Pass@1 correctness is similar (0.723–0.912 across models), so the over-editing is paid waste, not paid capability.



What does 6.5x look like on a bill? A 50-engineer org doing 800 agent edits per engineer per month = 40k edits/mo. At average 500 output tokens per minimal fix × $15/M Opus 4.7 output = $300/mo. At 3,250 output tokens per over-edited fix = $1,950/mo. Delta is $1,650/mo per 40k edits, pure output-token waste with no correctness upside. Scale to your actual traffic.



Why 'just use a smaller model' isn't the answer: reasoning models got worse (not better) at minimal editing when given more reasoning budget. So you can't fix over-editing by paying more; you fix it by measuring the ratio and routing around it.



The metric CFOs actually need is over-edit ratio per agent: over_edit_ratio = output_tokens / minimum_required_tokens_to_achieve_green_tests. Infrastructure to compute this: log full diff of every agent edit, run patch-min on the diff offline, diff size ratio = your over-edit score.



Instrument over-edit ratio this quarter, treat it as a first-class SLO per agent (budget for <0.2 average), and route high-stakes "minimal" tasks to models whose published over-edit score is <0.1.



Attribution is the prerequisite for every other cost signal you'll want this year. LLMeter ships per-customer + per-agent attribution today. Over-edit ratio is the first quality-flavored metric where LLMeter's attribution layer is the right home.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Over-editing is a token tax: GPT-5.4 ships 6.5x more diff per fix than Claude Opus 4.6, and your bill notices

Thematisch verwandte Begriffe: Overediting, token, GPT54, ships · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

💬 Kommentare werden geladen…
Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-102367 | mall4j through 4.0 contains an insufficient session expiration vulnerab…
Advisory →
tsecurity.de Icon
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag