The question every AI team eventually asks: should we rent GPUs and run models ourselves, or just pay per token through an API?

The answer changed a lot in the last six months. GPU rental prices dropped. API prices dropped faster. New GPU generations shipped. And mixture-of-experts models made the whole calculation messier than it used to...