🕵️ SicherheitslückenWhat continuous operational resilience looks like under DORA(09.09.2026 um 17:53 Uhr)
🔧 AI Nachrichten OpenAI seeks tougher AI rules. CIOs may feel the ripple effects(10.09.2026 um 12:11 Uhr)
🔧 AI Nachrichten Mistral valued at €21bn after €3bn Series D funding round(08.09.2026 um 10:19 Uhr)
🪟 Windows TippsWindows XP's Cursor Indicator Is Getting a Windows 11 Refresh(25.08.2026 um 13:00 Uhr)
🕵️ SicherheitslückenWhat continuous operational resilience looks like under DORA(09.09.2026 um 17:53 Uhr)
🔧 AI Nachrichten OpenAI seeks tougher AI rules. CIOs may feel the ripple effects(10.09.2026 um 12:11 Uhr)
🔧 AI Nachrichten Mistral valued at €21bn after €3bn Series D funding round(08.09.2026 um 10:19 Uhr)
🪟 Windows TippsWindows XP's Cursor Indicator Is Getting a Windows 11 Refresh(25.08.2026 um 13:00 Uhr)

🔧 Programmierung 🕛 vor 1 Monat 6 Min Lesezeit
0

The 33,000-token tax, a 30-hour star race, and where agents actually fail

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

Originally published in

To do: audit your own overhead once: check your session's token count before typing anything (in Claude Code, /cost after one trivial message). Then prune — every MCP server, skill and tool definition you keep loaded is paid on every session, whether used or not.



A production migration to GPT-5.6, with numbers: 2.2× faster, 27% cheaper. Rare non-benchmark data — a team moved a production agent and published latency and cost deltas instead of vibes (251 points on HN). ) annotated 63,000+ steps across 1,794 CLI coding-agent trajectories, seven frontier models, three harnesses (OpenHands, MiniSWE, Terminus2). Findings: failures usually start in the first few execution steps, stay hidden until recovery is impossible, and are dominated by epistemic errors — the agent not knowing what it doesn't know.

To do: stop logging only pass/fail. Keep full trajectories, and put your one human checkpoint early (after the first commands run), not at the end — by the time the final answer looks wrong, the paper says it was usually unrecoverable long before.



Long-horizon benchmark lands, and nobody passes: 29 of 46 tasks unsolved.

To do: minimum bar even without new tools: run agents under a separate OS user with no SSH keys and no cloud credentials in env. A VM or container is better; your main account is worse.



The verifier business is now a unicorn. PI raised at a $1B+ valuation with $100M ARR, selling verifiers — the machinery that scores whether an agent's output is right. The market priced the bottleneck: evaluation, not generation. ); GPT-5.6 is simultaneously fresh in builders' hands.

To do: rare week where you can benchmark two frontier models on your own workload at low cost. Run the comparison now; both windows close.





Under the radar





  • (9★) — local-first, project-aware memory MCP server for Claude Code/Codex, with Obsidian sync. → One install gives all your agents shared memory across projects, on your disk not a SaaS.


  • — one issue a week, every item ends with something you can do, and if a week is slow I say so.

    Vollständiger Original-Bericht
    Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
    ↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Sam Altman calls GPT-6 Astra rollout ‘messy’ as enterprise users wait for access
1 Quelle
Swiss government explores replacing Microsoft 365 with open-source software
1 Quelle
What continuous operational resilience looks like under DORA
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten The 33,000-token tax, a 30-hour star race, and where agents actually fail

Thematisch verwandte Begriffe: 33000token, 30hour, star, race · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...