Gemma 3 QAT Models: Bringing AI to Consumer GPUs
developers.googleblog.com
Gemma 3 QAT Models: Bringing AI to Consumer GPUs
1–10 of 286 posts
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#2Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#3Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#4Could 16gb vram be enough for the 27b QAT version?
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#5Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#6Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#7Available on ollama: https://ollama.com/library/gemma3
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#8Available on ollama: https://ollama.com/library/gemma3
The community getting obsessed with Ollama has done huge damage to the field, as it's ineffecient compared to vLLM. Many people can get far more tok/s than they think they could if only they knew the right tools.
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#9Shouldn’t it fit a 5060 Ti 16GB, for instance?
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#10Am I missing something?
These have been out for a while; if you follow the HF link you can see, for example, the 27b quant has been downloaded from HF 64,000 times over the last 10 days.
Is there something more to this, or is just a follow up blog post?
(is it just that ollama finally has partial (no images right?) support? Or something else?)