Large IT projects run about 45% over budget on average. That figure comes from McKinsey and Oxford looking at more than 5,400 of them, and it is the reason a wave of tooling now promises to read your spec and tell you what it will cost.
We were about to build one of those tools. Before writing the product, we wrote the test.
The claim is falsifiable, which is the useful thing about it. The effort-estimation research community has been publishing datasets with real recorded effort for decades. So we asked the version of the question that can come back "no":
On public data with real logged effort, can a statistical engine beat the standard baselines and the human expert, out of sample, without leakage?
It cannot. Not with any model class we threw at it. , the per-phase results are in .
Why publish this
We wrote a test designed to be hard, ran it on ourselves, and it came back no. Publishing that is cheaper than the alternative, which is shipping a claim about cold-start estimation from a spec and finding out in front of a client.
There is also a shortage of this. Negative results in applied ML mostly do not get written up, so the same hypothesis keeps getting re-tested privately by people who cannot see each other's results. If you are evaluating an effort-estimation vendor, the benchmark is a thing you can point at and ask them to run.
Repo: github.com/NaCode-Studios/metis-benchmark. Issues are open, and disagreement with the method is the most useful kind.
SOCIAL SHARE CARD GENERATOR