Earlier quoted context omitted.
Yea, No doubt Qwen 3.6 open weights are far more strong
Why no doubt?
Granite 4.1: IBM's 8B Model Matching 32B MoE
41–50 of 223 posts
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#42I test drove it yesterday. It's pretty impressive at 8b. Runs on commodity hardware quickly. Qwen3.6 35b a3b is still my local champion but I may use this for auto complete and small tasks. Granite has recent training data which is nice. If the other small models got fine tuned on recent data I don't know if I would use this at all, but that alone makes it pretty decent. The 4b they released was not good for my needs…
Have you tried the Gemma 4 series, out of curiosity? I haven’t run a local model in a while, but the benchmarks look good. I’d take a free local tool-use model if it was relatively consistent.
The 4b was okay. It didn't get all of my small math questions right, it didn't know about some of the libraries I use, but it was able to do some basic auto complete type stuff. For microscopic models I like the llama 3.2 3b more right now for what I do, it's a little faster and seems a little stronger for what I do. But everyone is different and I don't think I'll use it anymore this past month has been crazy for local model releases.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#43Earlier quoted context omitted.
Because Qwen 3.6 pushes way above its weight. Granite 8B is impressive, but Qwen still wins on raw capability, especially for coding.
You just asserted the same thing again. Why do you say this is the case?
Qwen3.6 raises the bar for models of its size. There really isn't a comparison in my opinion.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#44Earlier quoted context omitted.
You just asserted the same thing again. Why do you say this is the case?
Having tried it. Qwen is really good. Also, generally, it makes sense. 8B models are generally not very good^. That this 8B model is decent is impressive, but that it could perform on par with a good model 4 times as large is a daydream. ^ - To be polite. The small models + tool use for coding agents are almost universally ass. Proof: my personal experience. Ive tried many of them.
edit: It was a play on The Big Lebowski, folks.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#45> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#46Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#47sounds interesting. Here's hoping they release a 32B model, thats a pretty good sweet spot for feasibility of home setups. edit: I just realised they do actually have a 30b release alongside this. Haven't tried it yet.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#48> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#49Earlier quoted context omitted.
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
The language of drama and import without meaningful substance. Words statistically likely to be used in a segue, regardless of the preceding or subsequent point. Particularly effective when it seems like you’re getting let in on a secret. Really fatiguing to read A writing teacher once excoriated me for saying that something was important. “Don’t tell me it’s important, show me, and let me decide, and if you do your…
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#50Interesting to see a pivot away from MoE by both IBM and mistral while the larger classes of SOTA of models all seem to be sticking to it. Quick vibe check of it- 8B @ Q6 - seems promising. Bit of a clinical tone, but can see that being useful for data processing and similar. You don't really want a LLM that spams you with emojis sometimes...