Four days after I published a piece arguing LLMs can't make the jump, , a mathematician who did his PhD on a problem adjacent to Gromov's conjecture, posted the sharpest version of the distinction I was reaching for and didn't quite land. Astra solved difficult problems inside existing conceptual worlds. Calculus, topology, scheme theory did something different — they didn't answer questions sitting inside a framework, they built frameworks new questions could be asked in.
Non-sofic groups existing or not was always a well-posed question inside group theory as it already stood. Astra found the object. It didn't invent group theory. That's the line: solving hard problems inside a conceptual world is not the same act as inventing the world.
Worth naming plainly, because it cuts the other way against overclaiming too: even a Lean certificate that type-checks doesn't confirm the formal statement actually captures the open problem the way mathematicians understood it. Someone still has to judge whether the formalization is asking the right question. That judgment is exactly the kind of move nobody automated here.
A commenter, Terence Tao's public conversation with ChatGPT on a counterexample to the Jacobian Conjecture. Tao's messages are short. The model's outputs, talking to him, are unusually concise — expertise shunts it out of explaining-to-amateurs mode. He pushes back without contradicting directly: "this looks more complex than I was hoping for." And the detail that matters most: Tao makes the leaps himself. He almost never takes the model's suggested next move.
Goedecke's conclusion: the human is the bottleneck, not the model, because the hard part is communicating exactly what kind of solution you want. The information is already in the model. It takes a very smart human to pull it out.
That's my bookmark-time argument, relocated. I've been deciding what's worth saving since 2016, one bookmark at a time, and calling that curation. Tao is doing the same thing in real time, inside a chat window, calling it prompting. Different timescale, same move: supply the frame, let the model fill it.
So the thesis needs updating, not abandoning. Not "LLMs can't jump." Something narrower and, I think, more true. Inside closed, formally verifiable worlds — math, code, games, anything with a Lean checker or a compiler or a scoreboard — the jump is getting crackable by scale and search, and Astra just proved it faster than I expected. Outside those worlds, in anything ambiguous, causally tangled, unverifiable in advance, nobody's shown it yet. Not the actual Einstein case. Not the actual geophysics case. Not the actual "is this bookmark worth keeping" case. And the people getting the most out of these models, Tao included, are the ones still doing that part themselves.
I got four days. Most theses don't get tested this fast, or this publicly. I'd rather be corrected in the open than be right by accident.
SOCIAL SHARE CARD GENERATOR