Live data from Hacker News

Granite 4.1: IBM's 8B Model Matching 32B MoE

firethering.com

71–80 of 223 posts

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#72

Earlier quoted context omitted.

So it’s just like, your opinion, man? edit: It was a play on The Big Lebowski, folks.

the (dead) internet is full of opinions exactly like this

you tried qwen3.6 and you think it is not good?

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#74
Very impressive series of SLM by IBM here.

I have been using it with their Chunkless RAG concept and it is fitting very well! (for curious https://github.com/scub-france/Docling-Studio)

I convinced that SLM are a real parto of solution for true integrated AI in process...

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#75
post #51

Earlier quoted context omitted.

Have you tried the Gemma 4 series, out of curiosity? I haven’t run a local model in a while, but the benchmarks look good. I’d take a free local tool-use model if it was relatively consistent.

Qwen 3.6 burns it to the ground. it was not even a challenge. Gemma4 seriously fails at toolcalls and agentic works. It got all messed up after 2-3 turns of Vibecoding.

How do you run it? vllm? llama.cpp?

Can you share some parameters you enable tool calling and agentic usage?

Or, higher level, some philosophies on what approaches you are using for tuning to get better tool calling and/or agentic usage?

I'm having surprisingly good success with unsloth/Qwen3.6-27B-GGUF:Q4_K_M (love unsloth guys) on my RTX3090/24GB using opencode as the orchestrator.

It concocts some misleading paths, but the code often compiles, and I consider that a victory.

You have to watch it like you would watch a 14 year old boy who says he is doing his homework but you hear the sound effects of explosions.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#76
post #51

Earlier quoted context omitted.

Have you tried the Gemma 4 series, out of curiosity? I haven’t run a local model in a while, but the benchmarks look good. I’d take a free local tool-use model if it was relatively consistent.

Qwen 3.6 burns it to the ground. it was not even a challenge. Gemma4 seriously fails at toolcalls and agentic works. It got all messed up after 2-3 turns of Vibecoding.

Gemma 4 31b was working ok for me; but it was consuming tons of memory on SWA checkpoints, I had to turn them way down, and as a 31b dense model is fairly slow on a Strix Halo. I did have a lot of tool calling issues on 26b-a4b, though.

The Qwen models are quite solid though.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#77

sounds interesting. Here's hoping they release a 32B model, thats a pretty good sweet spot for feasibility of home setups. edit: I just realised they do actually have a 30b release alongside this. Haven't tried it yet.

Try qwen 3.6. it will knock your socks off

[dead]
Post reply on HN