I rented an A100 for under an hour to answer one question. Does a single vLLM flag really change what a token costs? The flag was max_num_seqs: how many requests the server works on at once. I set it to 1, which is a deliberately bad setting, and then to 8. Same GPU, same model (Qwen2.5-0.5B), same prompts, eight requests in flight the whole time.... Weiterlesen: I rented an A100 to test one vLLM flag
Intelligence View
⚡ tsecurity.de Intelligence
I rented an A100 to test one vLLM flag
I rented an A100 for under an hour to answer one question. Does a single vLLM flag really change what a token costs? The flag was max_num_seqs: how many…