Mike Stonebraker has been right before. He built Ingres in 1972, then Postgres in the 1980s, then predicted the death of general-purpose databases in 2005. Every time, the industry caught up years later. Now he is saying something that should make every company betting on "chat with your data" very nervous.
In a recent interview on the Data Renegades podcast, the Turing Award winner revealed that LLMs score 0% accuracy on real-world data warehouse queries. Not 80%, which is what the popular benchmarks report. Zero.
. His message to AI researchers: "If you think you're really good at text-to-SQL, try a real benchmark, not a fake one."
The Oracle playbook, 1980s edition
Stonebraker has seen tech hype cycles before. He built Ingres at Berkeley in 1972 and commercialized it in 1980. The competition was Larry Ellison's Oracle.
His assessment of how Oracle competed is blunt. "Larry Ellison is a fabulous salesman. He made present tense and future tense indistinguishable. He basically lied to customers."
The example he gives is telling. Ingres implemented referential integrity, the database constraint that ensures data consistency (if you fire the last employee in a department, should the department still exist?). Oracle wrote two manual pages defining referential integrity, then added at the bottom: "Not yet implemented."
Customers bought Oracle anyway. The feature was on the roadmap. The documentation existed. The code did not. Sound familiar? It is the same playbook playing out today with AI companies shipping demos of capabilities that do not work in production.
"One size fits none"
In 2004, Stonebraker published a paper arguing that general-purpose databases were doomed. "One Size Fits All: An Idea Whose Time Has Come and Gone" (), offering durable workflows in TypeScript, Java, Go, and Python. Two-thirds of their customers are building agentic AI. The key insight: most agentic AI today is read-only (generate a prediction). But it is moving to read-write, and that is a distributed database problem. You want atomicity, consistency, and transactions. Moving $100 between accounts using two AI agents requires both to commit or both to roll back. That is what databases were built for.
What this means for engineers
Three takeaways from Stonebraker's career and current work:
1. "Chat with your database" is not production-ready. The benchmarks are gamed. Real enterprise data is messy, complex, and full of domain-specific knowledge that LLMs have never seen. If you are building a product that depends on text-to-SQL working reliably, test it on your actual data warehouse, not Spider.
2. Specialized always wins at scale. Postgres is the right choice for getting started. But if you are running a petabyte data warehouse, you need a column store. If you are doing vector search, you need a vector database. If you are processing streams, you need a stream processor. The one-size-fits-all era is over at the high end.
3. Agentic AI needs database fundamentals. When AI agents start writing to databases, you need transactions, consistency, and durability. That is not an AI problem. It is a database problem. The engineers who understand both worlds, AI and database internals, will be the ones building the systems that actually work.
Stonebraker's career is a masterclass in betting against the herd. He was right about specialized databases. He was right about MapReduce being inefficient (Google eventually abandoned it). He was right about eventual consistency being wrong for most use cases (Google abandoned that too with Spanner). His track record on text-to-SQL should make you pause before trusting the benchmark numbers.
Sources: , ,
SOCIAL SHARE CARD GENERATOR