Live data from Hacker News

Accelerating Gemma 4: faster inference with multi-token prediction drafters

blog.google

81–90 of 345 posts

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#81

I’m starting to think that googles strategy is a bit different then the other frontier providers. Focusing more on performance to compute efficiency over pure performance. And maybe that’s why Gemini is (seemingly) lagging behind? Other providers hitting capacity and hitting the limits subsidising their inference. Google strategy seems to be about scaling and distributing these models to their existing billions of us…

Isn't that where everyone's strategy is shifting?

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#82

I recently set up the 26B A4B model up on vLLM on an RTX3090 (4-bit) after a hiatus from local models. Just completely blown away by the speed and quality you can get now for sub-$1k investment. I tried first with Qwen but it was unstable and had ridiculously long thinning traces!

Some of the early quants for qwen3.6 were broken. It's still finicky but with a little hand holding it's crazy.

Local models are the future it's awesome

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#84
post #64

if someone wants to work with gemma and dont deal with ollama or configs - there is (my baby) https://airplane-ai.franzai.com/ Beta but useable

LM Studio (for example) is free, can you pitch me on your USP vs. it?

easiness of install (one download), zero configuration, zero online access by design - there will never we websearch, never any kind of tracking, your prompts stay on your device - you can totally put in user data, confident contracts, ...

plus over time the harness - coming version has a hotkey for screen capture, next release will have support for native excel, docx export

there is value in being offline by design

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#85
post #2

I find it puzzling Google doesn’t actively promote its own cloud for inference of Gemma 4. Open source is great, love it. But shouldn’t Google want me to be able to use and pay for it through Gemini and vertex?

Makes me wonder about the partnership with apple to use gemini. safe to assume apple has a preference for on-device, and the best open model (for consumer hardware at least) is a google property with an apache 2 license. Interesting dynamic and seemingly a bright spot in the market

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#86

Watching the computer write text sort of reminds me of using a modem to call a BBS in the old days. This seems like going from 300 baud to 1200 - a significant improvement, but still pretty slow, and someday we will wonder how we put up with it.

There was a startup posted here which built custom hardware that let the AI respond instantly. Thousands of tokens per second.

Taalas. A sibling comment of yours posted the chat demo URL -

https://chatjimmy.ai/

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#87
post #43

Google is singlehandedly carrying western open source models. Gemma 4 31B is fantastic. However, it is a little painful to try to fit the best possible version into 24GB vram with vision + this drafter soon. My build doesn't support any more GPUs and I believe I would want another 4090 (overpriced) for best performance or otherwise just replace it altogether.

Qwen is still better that Gemma though. Also you can tune it more for different tasks, which means that you can prioritize thinking and accuracy versus inference speed.

Yes I would just go with qwen.

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#88
post #59

I don't see it talked about much, but Gemma (and gemini) use enormously less tokens per output than other models, while still staying within arms reach of top benchmark performance. It's not uncommon to see a gemma vs qwen comparison, where qwen does a bit better, but spent 22 minutes on the task, while gemma aligned the buttons wrong, but only spent 4 minutes on the same prompt. So taken at face value, gemma is now…

Anecdotally the 15/month basic Gemini plan allows coding all day. I'm not hitting the limits or needing to upgrade to 100/month plans like other people are doing with Claude or Codex. Caveat: Gemini has been dumbed down a few times over the last year. Rate limits tightened up too. So it might not be this good in the future.

no 15/month does not enough all day? pls dont share wrong info, 3.1 pro CLI sometimes wait 20-30 min thinking sometimes, it's by far worse compared to others.It finishes with few hours of work mostly, but in openai they give you 6 times of that in 24 hours, gemini resets one time a day. It is literally lazy and so many times does half work. I'm a power user for all top models in top 3 AI companies, only Gemini 3.1 waits so long and it's so slow. Even Gemini pro 3 and pro 2.5 was not like this at all

Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters

#90
post #86

Earlier quoted context omitted.

There was a startup posted here which built custom hardware that let the AI respond instantly. Thousands of tokens per second.

Taalas. A sibling comment of yours posted the chat demo URL - https://chatjimmy.ai/

Woah. How is this working? It's stupid fast.
Post reply on HN