Claude Opus 4.8 shipped today. The benchmarks are a distraction — here is what actually changes about how your agents run tomorrow.
Anthropic is the 30-minute read.
The 200k context behavior: not just a number
The second change matters in a way the launch post does not explain. The needle-in-a-haystack chart in the . The token-cost reduction from this single change is larger than the entire 4.7→4.8 model upgrade. Total time: 2-3 hours including measuring it.
Audit your eval suite for exact-string assertions on tool call arguments. Replace them with semantic equivalence checks, or accept that you will see flaky tests for the next month. Skipping this step is how you find out about the regression at 11pm on a Thursday from PagerDuty. Total time: half a day, but worth it.
The model upgrade is the easy part. The harness discipline is the part that compounds. Pick one of the three for today.
SOCIAL SHARE CARD GENERATOR