AI startup , with no grounding beyond the YottaGraph itself.
The benchmark, a , it could: “Across 12 investment banking topics judged on a six-dimension, 1–10 rubric, it scored 9.67 mean versus 9.87 for Gemini Deep Research Max (3.1 Pro), at roughly six cents per report instead of seven dollars, and in under five minutes instead of 17 minutes.”
When asked what prompted the decision to embark on the benchmark, , an independent technology analyst, said that he has long been uncomfortable with the industry’s “arms race-like mentality toward AI dominance,” noting, “Lovelace’s benchmarks suggest the industry may have been focusing on the wrong thing all along.”
Efficiency, he pointed out, has been almost invisible in the industry’s rush toward an AI-enabled future: “Bigger, ever more capable models have grabbed the headlines as vendors fight for bragging rights. And while size certainly matters in terms of any AI model’s ability to effectively crunch massive workloads, we seem to have ignored the costs of all that capability, and whether those costs are even worth it.”
Levy pointed out that someone would not use a sledgehammer to tighten a loose fitting on the front porch. “Instead, we’d choose a smaller, more effective tool to do the job,” he said. “The same logic applies in AI, and while the first few years of the AI era have been almost uniquely focused on the biggest, most powerful tools, we’ve failed to match the costs of all the capability to the underlying business issues that are being addressed.”
Enterprises bringing sledgehammers to every engagement, he said, “are probably significantly overspending relative to organizations that use tools specifically sized to a given business need. As metered AI use becomes a more mainstream reality in the enterprise, efficiency will need to become a priority.”
All of this, Levy added, “is as important for individual enterprises looking to control runaway AI compute costs as it is for vendors building the necessary data center and energy infrastructure to power it all. It’s also critical for governments to ask about efficiency before green-lighting an endless lineup of massively scaled infrastructure.”
Time to redesign AI models
Sanchit Vir Gogia, chief analyst at Greyhound Research, described the benchmark as being “meaningful, but not in the way the market will read it. It is not a clean win for small models over large ones. The defensible reading is narrower: for bounded, evidence-heavy research, the architecture around a model can compress the cost of a good answer by two orders of magnitude without degrading it.”
The industry, he said, “has spent three years acting as though intelligence lived inside the model, reaching for a bigger one whenever quality disappointed. That reflex is now meeting enterprise economics with all the grace of a grand piano pushed down a staircase. A brilliant model fed poor grounding is still a very expensive guesser.”
But financial research, said Gogia, “runs on filings and entities that behave like a graph. It is a graph-shaped problem wearing a research suit. The first question is no longer which model is most capable, but which system answers reliably at a defensible cost. Capability, in the enterprise, is now a property of the system, not the model.”
CIOs and CTOs, he added, “should treat the claim as a prompt to redesign their AI operating model, not as an instruction to buy a product. AI value is now determined by four connected disciplines: the quality of the context supplied, the routing of work to the appropriate model, the governance of the workflow, and the observability of what the system costs and does.”
He pointed out that an agent given poor context does not merely return a weak answer; it takes weak action, which is an operational hazard.
“Buyers should demand the cost of a completed workflow, and ask how portable the architecture is, since context lock-in is still lock-in,” he said. “The CIO’s task is to design the safest path from evidence to outcome.”
SOCIAL SHARE CARD GENERATOR