Live data from Hacker News

Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

blog.google

51–60 of 138 posts

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#51

I don't get this obsession with smaller models. I've been using Claude and GPT models for years and have had zero issues with them. I see absolutely no benefit to me as a end user for a local model which is going to take up more of my CPU and memory and slow down my machine. I almost always have Internet and if I don't then not having access to a AI model is the least of my concerns.

I don't like the gaslighting of paying Anthropic or Open(Closed)AI and it being said its unsustainable for them to take my payment while simultaneously they take my data (edit: which is incredibly valuable) and I cannot opt out of that.

The obsession is for leaving hostile and abusive entities, the corporations or the people who fund them that have a horrible track record in regards to ethicality, rights and respect & human dignity.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#52
post #42

I don't get this obsession with smaller models. I've been using Claude and GPT models for years and have had zero issues with them. I see absolutely no benefit to me as a end user for a local model which is going to take up more of my CPU and memory and slow down my machine. I almost always have Internet and if I don't then not having access to a AI model is the least of my concerns.

I like using my computer.

Exactly, thank you, we are on the same page! It's great to be able to use our own devices and not have their compute coopted by a third party.

I'd rather not have intensive compute needed shifted onto my personal machine which I want to use for something else.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#53
post #42

Earlier quoted context omitted.

I like using my computer.

Exactly, thank you, we are on the same page! It's great to be able to use our own devices and not have their compute coopted by a third party. I'd rather not have intensive compute needed shifted onto my personal machine which I want to use for something else.

I am not a "third party" on my own computer.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#54
post #51

I don't get this obsession with smaller models. I've been using Claude and GPT models for years and have had zero issues with them. I see absolutely no benefit to me as a end user for a local model which is going to take up more of my CPU and memory and slow down my machine. I almost always have Internet and if I don't then not having access to a AI model is the least of my concerns.

I don't like the gaslighting of paying Anthropic or Open(Closed)AI and it being said its unsustainable for them to take my payment while simultaneously they take my data (edit: which is incredibly valuable) and I cannot opt out of that. The obsession is for leaving hostile and abusive entities, the corporations or the people who fund them that have a horrible track record in regards to ethicality, rights and respect…

My view is, if you're going to use the service - you should give the data.

It's like using Gmail and expecting them not to train their AI models on your data - how can you expect that when they're giving you a secure, reliable, highly functional email client completely for free?

The digital economy only works if everyone pays their fair share. If you don't want to give your data then you are really harming everyone by slowing down AI development for everyone else.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#55
post #42

Earlier quoted context omitted.

I like using my computer.

Exactly, thank you, we are on the same page! It's great to be able to use our own devices and not have their compute coopted by a third party. I'd rather not have intensive compute needed shifted onto my personal machine which I want to use for something else.

By that logic, any software you run that isn't fully built by yourself is "third party" therefore you shouldn't run anything at all on your machine, thus obviating the need for it entirely.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#56

Earlier quoted context omitted.

Exactly, thank you, we are on the same page! It's great to be able to use our own devices and not have their compute coopted by a third party. I'd rather not have intensive compute needed shifted onto my personal machine which I want to use for something else.

By that logic, any software you run that isn't fully built by yourself is "third party" therefore you shouldn't run anything at all on your machine, thus obviating the need for it entirely.

But practically AI inference requires substantial local computing resources. It's not some web app, it's a order of magnitude more compute needed

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#57

I don't get this obsession with smaller models. I've been using Claude and GPT models for years and have had zero issues with them. I see absolutely no benefit to me as a end user for a local model which is going to take up more of my CPU and memory and slow down my machine. I almost always have Internet and if I don't then not having access to a AI model is the least of my concerns.

The entire universe of automation projects that can be run effectively for free relative to SoTA models? I don't think many realize that most LLM embedded automation, pipelines, products will soon be able to run extremely cheaply on models Frontier models will be used for coding/creation use cases, yes. But for all the pseudo-deterministic, pipeline, analysis style things there will be no practical benefit to running…

The 26B model is really surprising, and it is impressively concise — it spends a lot less time dithering than Qwen3.6.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#58
post #51

Earlier quoted context omitted.

I don't like the gaslighting of paying Anthropic or Open(Closed)AI and it being said its unsustainable for them to take my payment while simultaneously they take my data (edit: which is incredibly valuable) and I cannot opt out of that. The obsession is for leaving hostile and abusive entities, the corporations or the people who fund them that have a horrible track record in regards to ethicality, rights and respect…

My view is, if you're going to use the service - you should give the data. It's like using Gmail and expecting them not to train their AI models on your data - how can you expect that when they're giving you a secure, reliable, highly functional email client completely for free? The digital economy only works if everyone pays their fair share. If you don't want to give your data then you are really harming everyone b…

Because we pay for the models.

If I pay you for a service, what implicit right should you have to then continue to profit in perpetuity by storing the data I paid you to process?

If LLMs were free your Gmail analogy might hold up. They aren’t, and so it doesn’t.

AI development can continue with the data folks opt into, or with the data AI companies incessantly scrape with reckless disregard for polite system loads. AI development does not require retaining all user inputs forever.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#59

Earlier quoted context omitted.

By that logic, any software you run that isn't fully built by yourself is "third party" therefore you shouldn't run anything at all on your machine, thus obviating the need for it entirely.

But practically AI inference requires substantial local computing resources. It's not some web app, it's a order of magnitude more compute needed

[deleted]

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#60

Earlier quoted context omitted.

By that logic, any software you run that isn't fully built by yourself is "third party" therefore you shouldn't run anything at all on your machine, thus obviating the need for it entirely.

But practically AI inference requires substantial local computing resources. It's not some web app, it's a order of magnitude more compute needed

Hopefully now you understand why people want smaller models.
Post reply on HN