🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 9 Min Lesezeit
0

We built the benchmark we'd want any AI estimation vendor to pass. Then we failed it.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

Large IT projects run about 45% over budget on average. That figure comes from McKinsey and Oxford looking at more than 5,400 of them, and it is the reason a wave of tooling now promises to read your spec and tell you what it will cost.



We were about to build one of those tools. Before writing the product, we wrote the test.



The claim is falsifiable, which is the useful thing about it. The effort-estimation research community has been publishing datasets with real recorded effort for decades. So we asked the version of the question that can come back "no":




On public data with real logged effort, can a statistical engine beat the standard baselines and the human expert, out of sample, without leakage?




It cannot. Not with any model class we threw at it. , the per-phase results are in .






Why publish this



We wrote a test designed to be hard, ran it on ourselves, and it came back no. Publishing that is cheaper than the alternative, which is shipping a claim about cold-start estimation from a spec and finding out in front of a client.



There is also a shortage of this. Negative results in applied ML mostly do not get written up, so the same hypothesis keeps getting re-tested privately by people who cannot see each other's results. If you are evaluating an effort-estimation vendor, the benchmark is a thing you can point at and ask them to run.



Repo: github.com/NaCode-Studios/metis-benchmark. Issues are open, and disagreement with the method is the most useful kind.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage