Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
1–10 of 24 posts
Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#2Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#3More of a meta comment, but I really wish anthropic would say something about their plans for Fable. We're all just kind of left here floating and aimless, with no idea of what to expect
Hard to imagine where things go from here. GLM-5.3 will be released some day, with Fable class capabilities, and the (MAGA) US government will still be faffing around in their alt-reality cinematic bullshitiverse.
Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#4More of a meta comment, but I really wish anthropic would say something about their plans for Fable. We're all just kind of left here floating and aimless, with no idea of what to expect
Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#5For comparison, the current agent swarm challenge on HF is at 508 tok/s on a A10G GPU:
https://huggingface.co/spaces/gemma-challenge/gemma-dashboar...
Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#6Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#7That's very impressive. What's the best way to run these kernels natively on a Mac? I saw that there's a way to plug Claude into Apple's Foundation Models framework, and there's a CLI tool that can access models via that framework. It might be useful to have something so fast and good available via a small CLI tool for various purposes, especially when connected with a small suite of tools I have for things like file…
Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#8More of a meta comment, but I really wish anthropic would say something about their plans for Fable. We're all just kind of left here floating and aimless, with no idea of what to expect
They're kind of at the mercy of the US government on this, and the government seems to have them in the position you describe.
Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#9That's very impressive. What's the best way to run these kernels natively on a Mac? I saw that there's a way to plug Claude into Apple's Foundation Models framework, and there's a CLI tool that can access models via that framework. It might be useful to have something so fast and good available via a small CLI tool for various purposes, especially when connected with a small suite of tools I have for things like file…
Either ollama or omlx, both are pretty dang performant. Omlx lets you run Claude code locally though as long as you bootstrap it with the right model
Re: Fable 5 pushed Gemma 4 to 255 tok/s on WebGPU
#10More of a meta comment, but I really wish anthropic would say something about their plans for Fable. We're all just kind of left here floating and aimless, with no idea of what to expect