Every "run this model locally" guide tells you to grab a Q4 GGUF and move on. That advice is fine right up until you try a long-context run and your machine starts swapping.




The weights are the part everyone budgets for


Quantization maths is straightforward. A model's weight footprint is roughly params x bits /...