Earlier quoted context omitted.
As an aside uvx is so pleasant to use... I wish Nvidia supported it as first-class rather than making folks jump through Docker hoops.
I wish people would stop using python sure ai. It's slow and the PKG resolution is way too flat.
Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
111–120 of 138 posts
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#112Earlier quoted context omitted.
These models aren't products? They are open source ish (open weight I guess), research outputs. While the naming scheme may be confusing, it is relevant and important. I believe it's on you to understand it.
> I believe it's on you to understand it. This is exactly why Google has 10 messenger Apps.
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#113Unsloth's collection as well [0], with their results [1]. Looks like they can get very close to 100% accuracy compared to the BF16 model that is unquantized, and Unsloth's quants are better than the original Google's QAT as posted in the article. Personal I'm using the 2B model for web search and structured JSON output back via Unsloth Studio and its API, works very well for that even with the model embedded on phone…
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#114Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#115I wish they would release the base (non instruction tuned) models for use with pattern completion.
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#116Noob q: can advancements like this targeted at local inference have bonus effects for cloud inference? Presumably if you can get great results on cheaper hardware that also equates to less resource usage on cutting edge hardware, and less power draw? Will advancements like this ultimately reduce the carbon footprint of AI?
Also Google Deepmins has a six month embargo on strategic papers, so I bet the juiciest quantization tech isn't public yet.
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#117Earlier quoted context omitted.
I'll be honest with you. My main ask for on device AI is that when I am typing "Going out for a quick j" it corrects to "jog" and not "Jonathan". I don't think it needs that many gigabytes.
Who doesn't enjoy a quick Jonathan now and then. But seriously, wouldn't productive text on a 90s cell phone pass this test?
It’s harder now because emojis and draw-to-type as well as pen input. We didn’t have these things 14 years ago when “I’ll be right back” could be expanded from “I’ll b ri ba”
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#118had a good run with Gemma 4 E2B Unsloth 4Q: https://youtube.com/shorts/XLsAnz5aAAI The E4B model doesn’t fit on my phone TPU, so it swaps to RAM, the QAT version means more accuracy, good!
How do you know it swaps to ram vs on the TPU? Would be interested in testing this on my pixel.
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#119Earlier quoted context omitted.
However, you didn't actually get what I meant down, so you ended up inadvertently Straw Manning me. My disinterest is in sharing my intellectual IP. Most people up to now, have never shared this much of their intellectual IP with a company. Name one product through human history before that got this much data and insight into human thinking and now can use your most intimate conversations, ideas and needs for non-tra…
Intellectual "property" is not real property. While I disagree with the parent on many things as my comments show, IP is not one of them. Information should be free, for anyone and everyone.
"real" property or not. You agree that we have some right to our own outputs, right? Is that not dignity, to say "I want my outputs protected".
Seems like you think that your ideas should be free, as you called it information. How about you back that up with action... please send me all your most intimate, valuable ideas. Oh no, you don't feel comfortable? Then why are you sharing it with companies?
Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
#120Earlier quoted context omitted.
Apple is a good example of ethical services. They still give you privacy and ownership of your data, you keep your dignity and data. Google is a horrible model for this - it matches the whole thing about unethical, abusive, gaslighting relationships I described.
The same Apple that takes 30% of all transactions? In reality no corporation is in the public interest.