I’m starting to think that googles strategy is a bit different then the other frontier providers. Focusing more on performance to compute efficiency over pure performance. And maybe that’s why Gemini is (seemingly) lagging behind? Other providers hitting capacity and hitting the limits subsidising their inference. Google strategy seems to be about scaling and distributing these models to their existing billions of us…
Accelerating Gemma 4: faster inference with multi-token prediction drafters
81–90 of 345 posts
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#82I recently set up the 26B A4B model up on vLLM on an RTX3090 (4-bit) after a hiatus from local models. Just completely blown away by the speed and quality you can get now for sub-$1k investment. I tried first with Qwen but it was unstable and had ridiculously long thinning traces!
Local models are the future it's awesome
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#83Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#84if someone wants to work with gemma and dont deal with ollama or configs - there is (my baby) https://airplane-ai.franzai.com/ Beta but useable
LM Studio (for example) is free, can you pitch me on your USP vs. it?
plus over time the harness - coming version has a hotkey for screen capture, next release will have support for native excel, docx export
there is value in being offline by design
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#85I find it puzzling Google doesn’t actively promote its own cloud for inference of Gemma 4. Open source is great, love it. But shouldn’t Google want me to be able to use and pay for it through Gemini and vertex?
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#86Watching the computer write text sort of reminds me of using a modem to call a BBS in the old days. This seems like going from 300 baud to 1200 - a significant improvement, but still pretty slow, and someday we will wonder how we put up with it.
There was a startup posted here which built custom hardware that let the AI respond instantly. Thousands of tokens per second.
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#87Google is singlehandedly carrying western open source models. Gemma 4 31B is fantastic. However, it is a little painful to try to fit the best possible version into 24GB vram with vision + this drafter soon. My build doesn't support any more GPUs and I believe I would want another 4090 (overpriced) for best performance or otherwise just replace it altogether.
Qwen is still better that Gemma though. Also you can tune it more for different tasks, which means that you can prioritize thinking and accuracy versus inference speed.
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#88I don't see it talked about much, but Gemma (and gemini) use enormously less tokens per output than other models, while still staying within arms reach of top benchmark performance. It's not uncommon to see a gemma vs qwen comparison, where qwen does a bit better, but spent 22 minutes on the task, while gemma aligned the buttons wrong, but only spent 4 minutes on the same prompt. So taken at face value, gemma is now…
Anecdotally the 15/month basic Gemini plan allows coding all day. I'm not hitting the limits or needing to upgrade to 100/month plans like other people are doing with Claude or Codex. Caveat: Gemini has been dumbed down a few times over the last year. Rate limits tightened up too. So it might not be this good in the future.
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#89Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#90Earlier quoted context omitted.
There was a startup posted here which built custom hardware that let the AI respond instantly. Thousands of tokens per second.
Taalas. A sibling comment of yours posted the chat demo URL - https://chatjimmy.ai/