Accelerating Gemma 4: faster inference with multi-token prediction drafters
1–10 of 345 posts
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#2Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#3Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#4I find it puzzling Google doesn’t actively promote its own cloud for inference of Gemma 4. Open source is great, love it. But shouldn’t Google want me to be able to use and pay for it through Gemini and vertex?
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#5I find it puzzling Google doesn’t actively promote its own cloud for inference of Gemma 4. Open source is great, love it. But shouldn’t Google want me to be able to use and pay for it through Gemini and vertex?
Might be easier to chuck it over the fence and let other providers handle it as it'll run in almost any commercial grade card?
Also speculating, but I wonder if it might also create a bit of a pricing problem relative to Gemini flashlight depending on serving cost and quality of outputs?
As a comparison, despite being SotA for their size, the smallest qwen models on openrouter (27b and 35b) are not at all worth using, as there are way bigger and better models for less oricemon a per token basis
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#6I find it puzzling Google doesn’t actively promote its own cloud for inference of Gemma 4. Open source is great, love it. But shouldn’t Google want me to be able to use and pay for it through Gemini and vertex?
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#7Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#8Has anyone managed to get this to work in LM Studio? They've got a option in the UI, but it never seems to allow me to enable it.
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#9The performance uplift on local/self-hosted models in both quality and speed has been amazing in the last few months.
Re: Accelerating Gemma 4: faster inference with multi-token prediction drafters
#10I find it puzzling Google doesn’t actively promote its own cloud for inference of Gemma 4. Open source is great, love it. But shouldn’t Google want me to be able to use and pay for it through Gemini and vertex?
I wonder if for a model that small with a permissive license it might not be worth their time to host a commercial grade inference stack? Might be easier to chuck it over the fence and let other providers handle it as it'll run in almost any commercial grade card? Also speculating, but I wonder if it might also create a bit of a pricing problem relative to Gemini flashlight depending on serving cost and quality of ou…