That said, the first graph is misleading about the number of H100s required to run DeepSeek r1 at FP16. The model is FP8.
Gemma 3 QAT Models: Bringing AI to Consumer GPUs
21–30 of 286 posts
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#22Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#23? Am I missing something? These have been out for a while; if you follow the HF link you can see, for example, the 27b quant has been downloaded from HF 64,000 times over the last 10 days. Is there something more to this, or is just a follow up blog post? (is it just that ollama finally has partial (no images right?) support? Or something else?)
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#24Earlier quoted context omitted.
QAT “quantization aware training” means they had it quantized to 4 bits during training rather than after training in full or half precision. It’s supposedly a higher quality, but unfortunately they don’t show any comparisons between QAT and post-training quantization.
I understand that, but the qat models (1) are not new uploads. How is this more significant now than when they were uploaded 2 weeks ago? Are we expecting new models? I don’t understand the timing. This post feels like it’s two weeks late. [1] - https://huggingface.co/collections/google/gemma-3-qat-67ee61...
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#25Available on ollama: https://ollama.com/library/gemma3
How many times do I have to say this? Ollama, llamacpp, and many other projects are slower than vLLM/sglang. vLLM is a much superior inference engine and is fully supported by the only LLM frontends that matter (sillytavern). The community getting obsessed with Ollama has done huge damage to the field, as it's ineffecient compared to vLLM. Many people can get far more tok/s than they think they could if only they kne…
It is important to know about both to decide between the two for your use case though.
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#26They keep mentioning the RTX 3090 (with 24 GB VRAM), but the model is only 14.1 GB. Shouldn’t it fit a 5060 Ti 16GB, for instance?
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#27Available on ollama: https://ollama.com/library/gemma3
Make sure you're using the "-it-qat" suffixed models like "gemma3:27b-it-qat"
Re: Gemma 3 QAT Models: Bringing AI to Consumer GPUs
#28Earlier quoted context omitted.
I understand that, but the qat models (1) are not new uploads. How is this more significant now than when they were uploaded 2 weeks ago? Are we expecting new models? I don’t understand the timing. This post feels like it’s two weeks late. [1] - https://huggingface.co/collections/google/gemma-3-qat-67ee61...
8 days is closer to 1 week then 2. And it’s a blog post, nobody owes you realtime updates.
> 17 days ago
Anywaaay...
I'm literally asking, quite honestly, if this is just an 'after the fact' update literally weeks later, that they uploaded a bunch of models, or if there is something more significant about this I'm missing.