Mathematics has an annoying habit: it keeps handing you the same idea over and over under different names, and never once admits you've already met. In your first year you get told about the Legendre transform in mechanics (you understand nothing). Then about free energy in thermodynamics (you understand a bit more but see no connection). Then about softmax in neural networks (by this point you're fairly sure it's just an engineering hack someone came up with because it happened to work). And then, at some point — for me it happened while reading a particular pedagogical paper, more on that below — you suddenly realize that all along you'd been shown the same thing. One operation. Just wearing different costumes.
This is a genuinely satisfying feeling, something like the moment in a detective novel when it turns out the butler, the gardener, and the mysterious stranger on the train are the same person. I want to share that feeling, because it's pleasant, and also because it's useful: once you understand one construction, you get thermodynamics, exponential families from statistics, Fisher information, and — out of nowhere — the last layer of every large language model, plus the temperature slider in its API, all for free. Not a bad exchange rate for one idea.
I'll lean on a paper by Zia, Redish, and McKay, "Making Sense of the Legendre Transform" ( — the pedagogical paper this piece leans on: the tangent-line geometry, the symmetric form, and the saddle-point derivation.
SOCIAL SHARE CARD GENERATOR