Gemma 4 12B: A unified, encoder-free multimodal model
51–60 of 421 posts
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#52I do enjoy the immediate out of touch signaling with the "runs on your 16gb vram laptop" line. Because everyone has a laptop with 16gb vram, or can just pop out and buy a new one, right?
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#53The big story here is the encoder-free part, which I still don't fully understand. > Vision: We replaced Gemma 4’s vision encoder with a lightweight embedding module consisting of a single matrix multiplication, positional embedding and normalizations. That's technically encoding, just without using a dedicated model for it like SigLIP? The Developer's Guide elaborates, it's still a 35M layer which I am curious is ro…
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#54A model that comfortably fits in 16GB of VRAM (allowing room for context) is a welcome upgrade.
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#55What are the use cases for these small models? Is there anyone using models of this scale in their daily life who could share their experience?
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#56Re: Gemma 4 12B: A unified, encoder-free multimodal model
#57What's Google's business case for releasing open models? Don't get me wrong, I am grateful and appreciative of these releases. I'm trying to understand how it fits into their bigger picture as a for profit company? Are they not helping competitors build on the novel technology they have developed? Is it simply goodwill and/or marketing? Or am I missing something strategic?
Eventually the local model is not enough, and you'll upgrade to the big ones.
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#58What are the use cases for these small models? Is there anyone using models of this scale in their daily life who could share their experience?