🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 3 Min Lesezeit
0

The 20-minute check I run before swapping an agent to a new model

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

Every time a new model ships, the same ritual: change the model string, run the agent, read a few replies, they look good, ship it.



The replies are the one part of an agent that almost never breaks visibly. What breaks is behavior. A tool call quietly disappears, an argument drifts, a refund amount loses its decimal point, and the agent keeps talking like everything is fine. I learned this the hard way when a swap made my agent stop calling cancel_subscription while it kept telling users their subscription was cancelled.



With a new frontier model out this week, a lot of model strings are about to change. This is the check I now run before any swap. It takes about 20 minutes and produces a real diff instead of a vibe check.






1. Record a baseline before you touch anything



This is the step you cannot recover later. Once you swap, the old behavior is gone.



Start a recording proxy and point your agent at it, no code changes:




CODE
npx whatbroke-cli record --out baseline.jsonl









CODE
OPENAI_BASE_URL=http://127.0.0.1:4141/v1     # openai sdk
ANTHROPIC_BASE_URL=http://127.0.0.1:4141 # anthropic sdk






Run your agent through its real scenarios. Agents are nondeterministic, so run each scenario three times and name the runs refund-flow#1, refund-flow#2, refund-flow#3 (an x-whatbroke-run header per request, or --run). Three samples per scenario is enough to tell flaps from real changes.



Pick scenarios where the agent has to do something: call tools, hit an API, write a record. Pure chat scenarios are where nothing ever visibly breaks, so they tell you the least.






2. Swap the model. Only the model.



Change the model string. Nothing else. If you also want to tweak the prompt for the new model, do that as a second swap with its own diff, otherwise you will never know which change caused what.






3. Record the same scenarios again






CODE
npx whatbroke-cli record --out swapped.jsonl






Same scenarios, same names, three runs each.






4. Diff






CODE
npx whatbroke-cli diff baseline.jsonl swapped.jsonl






Read it top down:





  • breaking findings first: dropped tool calls, runs that now fail, outputs that vanished. Any of these means the swap is not a drop-in.


  • changed findings next: argument drift is the sneaky one. The tool still gets called, but with different args. Check every one by hand.

  • The flap rate tells you whether a finding is real. 3/3 means it happens every time. 1/3 on something your baseline also flapped on is just your agent being itself, and the diff demotes those automatically.

  • Cost and latency move on every swap. Regressions past a ratio get flagged (defaults 1.5x latency, 1.25x cost).






5. Keep it in CI






CODE
npx whatbroke-cli diff baseline.jsonl current.jsonl --fail-on breaking






Exit code 1 on breaking changes, --md for a report you can drop into a PR comment.






Already swapped without a baseline?



If your agent runs behind Langfuse, LangSmith, or anything emitting OTel GenAI spans, you already have the baseline, you just have not diffed it yet. Export last week's traces and this week's:




CODE
npx whatbroke-cli import last-week-export.json --run baseline
npx whatbroke-cli import this-week-export.json --run swapped
npx whatbroke-cli diff last-week-export.whatbroke.jsonl this-week-export.whatbroke.jsonl






The tool is deterministic, fully offline, MIT licensed, and your traces never leave your machine: .



If you run this before your next swap and it catches something, I would genuinely love to hear what it was.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten The 20-minute check I run before swapping an agent to a new model

Thematisch verwandte Begriffe: 20minute, check, before, swapping · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...