🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 3 Min Lesezeit
0

Which LLM should I actually code with? I built a small benchmark to find out

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht






A small, self-run coding benchmark: 3 models on 14 problems across 3 languages, scored on pass@k, cost, and speed. Last run 11 July 2026.



favicon
peculiarengineer.com






I kept going back and forth on which model to reach for in my actual day job. Every "which LLM is best at code" thread turns into vibes and screenshots, and none of it answered the question I had, which is which one to open when I have real work in the languages I use. So I stopped guessing and built a small benchmark to settle it for myself.



It is deliberately small. 14 problems across Python, C#, and Bash, three models, three attempts each at temperature 0.7 with a 10 second timeout. Every attempt runs in a sandboxed Docker container and gets scored on pass@k, cost, and latency. It is not an authoritative ranking and I am not pretending it splits hairs. It is enough to show the shape.






The thing that surprised me



Accuracy is not the differentiator anymore. All three models solved every problem they were allowed to answer, 100% pass@3. If I only looked at pass rates I would have learned nothing, because they all pass.



The catch hides in "allowed to answer." One model got content filtered out of four problems, so its perfect score covers 10 of the 14, not the whole set. A perfect score on the problems you answered and a perfect score on the whole bench are not the same result, which is why the leaderboard sorts on coverage first.






Where they actually differ



If they all pass, the decision comes down to what you pay and how long you wait. That is where the spread lives.




  • Cost was close, about 1.3x from cheapest to priciest across the suite.

  • Latency was not close. 6.6x between the fastest and the slowest.

  • The cheapest run that also covered all 14 problems came in around $0.028 per solved problem.



So the honest summary is boring in the best way. Pick on speed and price, because accuracy already agrees.






The gotcha worth writing down



One result is not like the others. One model was quick and cheap on Python and C#, a second or two per problem, then fell off a cliff on Bash at 90 to 168 seconds per problem. Four slow Bash problems dragged its average latency up to something that makes it look slow overall when it is really fine everywhere except Bash. If your day is mostly shell, that matters a lot. If it is Python, you would never notice.



That is why a single "which model is fastest" number is a lie. Fastest at what, in which language, is the only version of the question worth asking.






What I took away



For my work it came down to speed and cost per language. The benchmark did the one job I wanted. It turned a running argument in my head into a few numbers I can re-run whenever the models change.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Which LLM should I actually code with? I built a small benchmark to find out

Thematisch verwandte Begriffe: Which, should, actually, code · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...