Author: DigitalOcean - Bewertung: 0x - Views:4 Serverless inference vs running your own GPU: I benchmarked both with the same model (gpt-oss-120b) and measured the part nobody shows you, the cold start.

On a self-hosted AMD MI300X GPU Droplet running vLLM, the cold start took 61 seconds every time (weights already on disk, so no download...