Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
1–10 of 108 posts
Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#2Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#3Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#4Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#5But I have to say that the comparison is not really fair. Comparison is done with a 2 B model vs frontier models that are likely 100s of times larger. Also taalas with their 15000 tok/s inference are suspiciously missing from the comparison.
We need to see the comparison with this framework and useful models, which at present seems to mean ~30 B.
Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#6Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#7> 8× NVIDIA H200
Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#8For instant code generatio, 400-500 tok/s should be sufficient, though most frontier models give us closer to 70 tok/s.
Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#9> Standard GPUs > 8× NVIDIA H200
Re: Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
#10This looks very interesting. Possible to get those rates without exotic hardware. But I have to say that the comparison is not really fair. Comparison is done with a 2 B model vs frontier models that are likely 100s of times larger. Also taalas with their 15000 tok/s inference are suspiciously missing from the comparison. We need to see the comparison with this framework and useful models, which at present seems to m…
they seem to think it scales up because theyre shortening the stack.