Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
1–10 of 11 posts
Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#2Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#3Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#4Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#5Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#6Is 500 tok/s on Gemma 7B a gamechanger? or is this more just an advertisement for mystic.ai?
Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#7I'll be honest that I've never actually considered tokens per second as something to focus on for my projects, I'm much more concerned with quality of the output then quantity. Is 500 tok/s on Gemma 7B a gamechanger? or is this more just an advertisement for mystic.ai?
Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#8I'll be honest that I've never actually considered tokens per second as something to focus on for my projects, I'm much more concerned with quality of the output then quantity. Is 500 tok/s on Gemma 7B a gamechanger? or is this more just an advertisement for mystic.ai?
Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#9Re: Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
#10I'll be honest that I've never actually considered tokens per second as something to focus on for my projects, I'm much more concerned with quality of the output then quantity. Is 500 tok/s on Gemma 7B a gamechanger? or is this more just an advertisement for mystic.ai?