Earlier quoted context omitted.
MLX is quite literally macOS-specific technology, for other platforms you want non-MLX. I was sure "MLX" stood for "Metal-something-something" but can't find any reference to that somehow, anywho, "Metal" is hardware-accelerated graphics on Apple platforms FWIW. Edit: about the actual release on Ollama, if you're on non-Apple hardware you probably want the NVFP4 variant ("gemma4:12b-nvfp4") which was uploaded 45 minu…
I realize this is a little confusing; we're working w/ the MLX team to bring MLX to other platforms, but we're not quite there yet. The `gemma4:12b-nvfp4` model is specifically for the MLX engine. For the GGUF 4bit variant (i.e. non-macs) you'll need `gemma4:12b-it-q4_K_M` which I just pushed. You'll also need to upgrade to version 0.30.4 which we're just about to release (it's in prerelease and we're running through…
Gemma 4 12B: A unified, encoder-free multimodal model
351–360 of 421 posts
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#352Earlier quoted context omitted.
It was almost certainly not trained for coding, as it's got both audio and vision input, is only 12B, and nowhere in the announcement is coding mentioned. It will likely not have good performance on coding in general, compared to other small models like Qwen 3.6 35B A3B, Gemma 4 26B A4B, Nvidia Nemotron 3 Nano 30B-A3B, gpt-oss-20b. For 16GB laptops, Qwen 3.5 9B is the undisputed champ. Gemma 4 31B is the top dog at s…
I find ram crazy. My thinkpad has 32G of ram, it's a t470 that's nearly a decade old Why do people with modern laptops have such little amounts of ram?
Probably doesn't matter these days with all-day batterys, but now the demand-supply curve is lopsided.
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#353Earlier quoted context omitted.
I would contend that the actual big story is the gallery app: https://developers.google.com/edge/gallery Anyone with a 16GB Mac — that is quite a lot of journalists, surely — can download that, install a model into it, and play. Surely journalists have to start asking questions at least about OpenAI's consumer revenue projections now. I am a major, major AI cynic, but I decided to be an informed cynic so I've been pl…
Is the story that it's now also available outside of android? I've had this app on my phone for I believe about a year.
The combination of these things, though, I still think is significant. It’s a product from an old-fashioned (!) FAANG that installs as easily as Chrome, downloads a model as easily as it could be, combines a chat interface with audio and video analysis/transcription, has a customisable system prompt, MTP, agent skills support etc.
Now, it is from Google so they could kill it when they get bored! But clearly this is local AI packaged in a really accessible format, and the model seems quite capable for its size. It is something Microsoft could do when they can really point to easy consumer hardware that can do it well. It’s certainly something Apple could do better with their distillations of Gemini under the Google deal.
I think a sane line of enquiry for a tech journalist is: 1) doesn’t this threaten the appeal of consumer-tier subscriptions to ChatGPT (which is a big part of OpenAI’s revenue plans), and 2) is it therefore not questionable that the buy-and-hold economics of DRAM, SSD and GPU products that OpenAI benefits from having provoked into causing ridiculous price increases is fundamentally anti-consumer?
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#354Earlier quoted context omitted.
I've got a home-built dictation app that uses a local model to clear up the text and fix grammar. It was super easy to build. I’m extending it to capture meeting notes and summarise too. All on-device. I saw a little app the other day, I think someone posted on here, that looks at your screenshot and renames the file based off the contents of the file. There's tons of little examples like that. For a lot of use cases…
That's a great user case. Am sorry using parakeet but sometimes it garbles up things. Can you open source it?
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#355Earlier quoted context omitted.
> It roughly compares with GPT-4.1 (!!), released 14 months ago I think the mayor win for coding was reasoning. That's why such a small model can match GPT-4.1 in coding, but I suspect that GPT-4.1 still wins in general world knowledge due to bigger size.
> I suspect ... still wins in general world knowledge due to bigger size Encyclopedic knowledge matters relatively little in perspective, given the expectable future developments: even the more knowledgeable of us will use that knowledge for reasoning and intuition (and we will have absorbed the intellectual keys during our training), but under our professional hat we should in theory be ready to go "I stand correcte…
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#356Earlier quoted context omitted.
It was almost certainly not trained for coding, as it's got both audio and vision input, is only 12B, and nowhere in the announcement is coding mentioned. It will likely not have good performance on coding in general, compared to other small models like Qwen 3.6 35B A3B, Gemma 4 26B A4B, Nvidia Nemotron 3 Nano 30B-A3B, gpt-oss-20b. For 16GB laptops, Qwen 3.5 9B is the undisputed champ. Gemma 4 31B is the top dog at s…
> For 16GB laptops, Qwen 3.5 9B is the undisputed champ. You seem like the guy to ask. For a laptop with 12GB VRAM (RTX 5070) and 32 GB system RAM, what is a good multilingual (English, Hebrew, Greek) model for conversing with personal notes in Org mode format? I don't care how long updating the model or rag takes, and even inference can be reasonably slow, but the results of the query as they relate to my personal n…
Qwen models are always good. The 35B A3 model is a MoE model which means it has higher performance in RAM constrained environments compared to the 27B dense model (which is better at coding).
I don't have experience to rate it's Hebrew or Greek performance but apparently it's not bad.
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#357Re: Gemma 4 12B: A unified, encoder-free multimodal model
#358Re: Gemma 4 12B: A unified, encoder-free multimodal model
#359Scorched earth tactics to make anthropic and openai IPO fail?
Re: Gemma 4 12B: A unified, encoder-free multimodal model
#360I don't understand why Google does this. If I can run this locally, why would I need a subscription or use any inference provider, including Google..? Scorched earth tactics to make anthropic and openai IPO fail?