Live data from Hacker News

Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

blog.google

81–90 of 138 posts

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#81

Earlier quoted context omitted.

you misunderstand what that chart shows - it shows BF16 QAT Q4_0, not BF16 regular. meaning Google quantized the model to 4 bit and stored the result in BF16 format for compatibility and convenience to downstream packers. Like storing small 8 bit numbers in full 32 bit integers. So it's not close to 100% of unquantized BF16. I'm curious if anybody can explain why Google released 4 bit QAT Q4_0 is not exactly 100% of…

> meaning Google quantized the model to 4 bit and stored the result in BF16 format for compatibility and convenience to downstream packers. You also misunderstand what is happening. Google did not do that. Google further trained the original model with an objective of minimizing error when quantized to 4-bit. The BF16 QAT is not an upscaled 4-bit model. When quantized to 4-bit, it should lose less accuracy than a typ…

So what we want now is unsloth (or anyone) to release 4/6-bit quantized models of these releases?

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#82

Earlier quoted context omitted.

> meaning Google quantized the model to 4 bit and stored the result in BF16 format for compatibility and convenience to downstream packers. You also misunderstand what is happening. Google did not do that. Google further trained the original model with an objective of minimizing error when quantized to 4-bit. The BF16 QAT is not an upscaled 4-bit model. When quantized to 4-bit, it should lose less accuracy than a typ…

So what we want now is unsloth (or anyone) to release 4/6-bit quantized models of these releases?

Yep, Unsloth already did, as linked in the comment at the top of this thread

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#83
post #3

had a good run with Gemma 4 E2B Unsloth 4Q: https://youtube.com/shorts/XLsAnz5aAAI The E4B model doesn’t fit on my phone TPU, so it swaps to RAM, the QAT version means more accuracy, good!

How were you getting anything useful out of that? We found the (unquantized!) E2B model to be completely useless at even the simplest real-world classification tasks.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#84
post #27

Earlier quoted context omitted.

Google already did https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-un...

This is safetensors. Is there any way to run these on a Mac paired with the MLX QAT? (Pardon my ignorance; this stuff moves so fast)

https://huggingface.co/lmstudio-community/gemma-4-26B-A4B-it...

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#86

Earlier quoted context omitted.

you misunderstand what that chart shows - it shows BF16 QAT Q4_0, not BF16 regular. meaning Google quantized the model to 4 bit and stored the result in BF16 format for compatibility and convenience to downstream packers. Like storing small 8 bit numbers in full 32 bit integers. So it's not close to 100% of unquantized BF16. I'm curious if anybody can explain why Google released 4 bit QAT Q4_0 is not exactly 100% of…

> meaning Google quantized the model to 4 bit and stored the result in BF16 format for compatibility and convenience to downstream packers. You also misunderstand what is happening. Google did not do that. Google further trained the original model with an objective of minimizing error when quantized to 4-bit. The BF16 QAT is not an upscaled 4-bit model. When quantized to 4-bit, it should lose less accuracy than a typ…

Are there evidence that this approach helps maintain "accuracy" performance when quantized? It sounds a bit like mxfp4 with gpt-oss, which was a confusing model upon release.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#89

Earlier quoted context omitted.

It's super annoying when you have products that utilize these because there's...4? releases in 3 weeks? - Gemma 4 2B/4B/27BE3B/31B - Gemma 4 2B/4B/27BE3B/31B x "assistant" / MTP drafter models (i.e. multitoken prediction) - Gemma 4 12B (2 days ago? 1?) - Gemma 4 QAT 2B/4B/12B/27BE3B/31B x "assistant" models (i.e. multitoken prediction) It probably sounds silly and really whiny in the abstract. It just causes a ton of…

These models aren't products? They are open source ish (open weight I guess), research outputs. While the naming scheme may be confusing, it is relevant and important. I believe it's on you to understand it.

> I believe it's on you to understand it.

This is exactly why Google has 10 messenger Apps.

Re: Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

#90
post #72

Earlier quoted context omitted.

Argentina vice presidents span from 2007 to 2023. Knowledge cutoff cant explain getting all 5 of them wrong.

What did it say were the presidents from those years?

It can answer presidents fine. It fails for vice presidents.

-----------

As of my last update, here are the five most recent individuals to have served as Vice President of Argentina:

Sergio Massa (Served as Vice President from 2019 to 2023)

Martín Lousteau (Served as Vice President from 2015 to 2019)

Cristina Fernández de Kirchner (Served as Vice President from 2007 to 2015)

Néstor Kirchner (Served as Vice President from 2003 to 2007)

Eduardo Duhalde (Served as Vice President from 1999 to 2003)

Note on the list: The term "most recent" can be interpreted in two ways:

Most recent to have served: This list follows that interpretation, showing the last five people who held the office.

Most recent current officeholders: If you are asking for the current Vice President, that position is currently held by Juan Manuel Moreno (who was appointed in 2024).

If you are looking for the current Vice President, please let me know!

Post reply on HN