Running large language models locally
gives you privacy, offline capability, and zero API costs.
This benchmark reveals exactly what one can expect from 9 popular
LLMs on Ollama on an RTX 4080.



With a 16GB VRAM GPU, I faced a constant trade-off:
bigger models with potentially better quality, or smaller models with faster inference.



...