Live data from Hacker News

Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

blog.google

111–120 of 138 posts

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#111

Earlier quoted context omitted.

As an aside uvx is so pleasant to use... I wish Nvidia supported it as first-class rather than making folks jump through Docker hoops.

I wish people would stop using python sure ai. It's slow and the PKG resolution is way too flat.

What do you use?

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#112
post #89

Earlier quoted context omitted.

These models aren't products? They are open source ish (open weight I guess), research outputs. While the naming scheme may be confusing, it is relevant and important. I believe it's on you to understand it.

> I believe it's on you to understand it. This is exactly why Google has 10 messenger Apps.

Google released their latest messenger app 9 years ago. https://en.wikipedia.org/wiki/Google_Chat

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#113

Unsloth's collection as well [0], with their results [1]. Looks like they can get very close to 100% accuracy compared to the BF16 model that is unquantized, and Unsloth's quants are better than the original Google's QAT as posted in the article. Personal I'm using the 2B model for web search and structured JSON output back via Unsloth Studio and its API, works very well for that even with the model embedded on phone…

Is this [0] saying that unsloth's versions of Google's QAT models are better than Google's own QAT models? Or am I not understanding it correctly?

[0] https://unsloth.ai/docs/models/gemma-4/qat#qat-analysis

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#114
post #94
post #80

Earlier quoted context omitted.

The E2B ones? Or what do you mean by specialized drafters?

They have -assistant in the name, so e.g.: https://huggingface.co/google/gemma-4-31B-it-assistant

Thanks

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#116

Noob q: can advancements like this targeted at local inference have bonus effects for cloud inference? Presumably if you can get great results on cheaper hardware that also equates to less resource usage on cutting edge hardware, and less power draw? Will advancements like this ultimately reduce the carbon footprint of AI?

Consumer and server hardware are quite different, especially Google's TPUs. They notably have much larger mixture-of-experts ratios and more complex caching systems. At such scale and inference budgets, they are incentivised to optimize as much as possible.

Also Google Deepmins has a six month embargo on strategic papers, so I bet the juiciest quantization tech isn't public yet.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#117
post #107

Earlier quoted context omitted.

I'll be honest with you. My main ask for on device AI is that when I am typing "Going out for a quick j" it corrects to "jog" and not "Jonathan". I don't think it needs that many gigabytes.

Who doesn't enjoy a quick Jonathan now and then. But seriously, wouldn't productive text on a 90s cell phone pass this test?

The autocomplete of a decade ago is better than what we have now.

It’s harder now because emojis and draw-to-type as well as pen input. We didn’t have these things 14 years ago when “I’ll be right back” could be expanded from “I’ll b ri ba”

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#118
post #39
post #3

had a good run with Gemma 4 E2B Unsloth 4Q: https://youtube.com/shorts/XLsAnz5aAAI The E4B model doesn’t fit on my phone TPU, so it swaps to RAM, the QAT version means more accuracy, good!

How do you know it swaps to ram vs on the TPU? Would be interested in testing this on my pixel.

Because TPU has 2GB and weight + context needs more

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#119
post #71

Earlier quoted context omitted.

However, you didn't actually get what I meant down, so you ended up inadvertently Straw Manning me. My disinterest is in sharing my intellectual IP. Most people up to now, have never shared this much of their intellectual IP with a company. Name one product through human history before that got this much data and insight into human thinking and now can use your most intimate conversations, ideas and needs for non-tra…

Intellectual "property" is not real property. While I disagree with the parent on many things as my comments show, IP is not one of them. Information should be free, for anyone and everyone.

Another straw man.

"real" property or not. You agree that we have some right to our own outputs, right? Is that not dignity, to say "I want my outputs protected".

Seems like you think that your ideas should be free, as you called it information. How about you back that up with action... please send me all your most intimate, valuable ideas. Oh no, you don't feel comfortable? Then why are you sharing it with companies?

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#120
post #73

Earlier quoted context omitted.

Apple is a good example of ethical services. They still give you privacy and ownership of your data, you keep your dignity and data. Google is a horrible model for this - it matches the whole thing about unethical, abusive, gaslighting relationships I described.

The same Apple that takes 30% of all transactions? In reality no corporation is in the public interest.

straw man. I'm talking about data.
Post reply on HN