Live data from Hacker News

Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU

xcancel.com

21–24 of 24 posts

Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU

#21

Earlier quoted context omitted.

some of us knew the cloud was unreliable and chose a better path.

Well I would say you just chose ANOTHER sensible path with tradeoffs just like frontier API/subscriptions. Yes you get 100% control, cheaper inference, not having to stare at status.claude.com for like 2 hours per week, complete privacy (assuming local hosting or self hosting on servers). But, you can’t get Fable level performance. OSS has reliably trailed the frontier by like 4-7 months for years now

Exactly, and insert meme "why not both?" I run local models for plenty of things, but for work there's a lot of complexity and Fable handled it so much better than any open model has so far. It's not really a choice for some domains currently.

Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU

#22
post #14
post #4

Earlier quoted context omitted.

They're kind of at the mercy of the US government on this, and the government seems to have them in the position you describe.

It's so great that the US is against AI regulation and gives corporations freedom to innovate /s

yep, and while Anthropic is now spending their time dealing with regulation, the Chinese models have some time to catch up.

I do wonder though, when those models are "mythos class" (whatever that means), will China do the same thing restricting it from export? If they get a model that's better than US companies, I fully expect them to stop open sourcing them for the world (but I hope I'm wrong about that).

Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU

#23
post #14

Earlier quoted context omitted.

It's so great that the US is against AI regulation and gives corporations freedom to innovate /s

yep, and while Anthropic is now spending their time dealing with regulation, the Chinese models have some time to catch up. I do wonder though, when those models are "mythos class" (whatever that means), will China do the same thing restricting it from export? If they get a model that's better than US companies, I fully expect them to stop open sourcing them for the world (but I hope I'm wrong about that).

China is benefiting a lot from releasing the models open-source and those benefits immediately end if they start doing closed-source releases. It would be very short-sighted

Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU

#24

That's very impressive. What's the best way to run these kernels natively on a Mac? I saw that there's a way to plug Claude into Apple's Foundation Models framework, and there's a CLI tool that can access models via that framework. It might be useful to have something so fast and good available via a small CLI tool for various purposes, especially when connected with a small suite of tools I have for things like file…

Either ollama or omlx, both are pretty dang performant. Omlx lets you run Claude code locally though as long as you bootstrap it with the right model

Omlx is really nice, thanks for the recommendation!
Post reply on HN