show me.
Granite 4.1: IBM's 8B Model Matching 32B MoE
61–70 of 223 posts
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#62Earlier quoted context omitted.
The language of drama and import without meaningful substance. Words statistically likely to be used in a segue, regardless of the preceding or subsequent point. Particularly effective when it seems like you’re getting let in on a secret. Really fatiguing to read A writing teacher once excoriated me for saying that something was important. “Don’t tell me it’s important, show me, and let me decide, and if you do your…
Maybe the solution is to cull the bad, cliché writing from the training data.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#63Earlier quoted context omitted.
Have you tried the Gemma 4 series, out of curiosity? I haven’t run a local model in a while, but the benchmarks look good. I’d take a free local tool-use model if it was relatively consistent.
I tried the Gemma 4 I think 2 and 4b. The 2b was not useful for me at all. A little too weak for my use cases The 4b was okay. It didn't get all of my small math questions right, it didn't know about some of the libraries I use, but it was able to do some basic auto complete type stuff. For microscopic models I like the llama 3.2 3b more right now for what I do, it's a little faster and seems a little stronger for wh…
curious how people are leveraging these models
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#64The real "sleeper" might be https://huggingface.co/ibm-granite/granite-vision-4.1-4b if the benchmarks hold up for such a small model against frontier models for table & semantic k:v extraction.
Woah, is this part of the future of models? Basically little models you can use as tools.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#65Interesting to see a pivot away from MoE by both IBM and mistral while the larger classes of SOTA of models all seem to be sticking to it. Quick vibe check of it- 8B @ Q6 - seems promising. Bit of a clinical tone, but can see that being useful for data processing and similar. You don't really want a LLM that spams you with emojis sometimes...
I never want LLM to span me with emojis. What is the use case for that? I find it highly annoying.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#66The real "sleeper" might be https://huggingface.co/ibm-granite/granite-vision-4.1-4b if the benchmarks hold up for such a small model against frontier models for table & semantic k:v extraction.
Woah, is this part of the future of models? Basically little models you can use as tools.
Regardless, the people in the 80s capable of pruning programs to fit on small devices is likely happening now. I'd bet most of the Chinese firms are doing it because of the US's silly GPU games among other constraints.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#67Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#68Earlier quoted context omitted.
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
Ugh, you're making me remember the last time I listened to NPR. It's so bad.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#69Earlier quoted context omitted.
Having tried it. Qwen is really good. Also, generally, it makes sense. 8B models are generally not very good^. That this 8B model is decent is impressive, but that it could perform on par with a good model 4 times as large is a daydream. ^ - To be polite. The small models + tool use for coding agents are almost universally ass. Proof: my personal experience. Ive tried many of them.
So it’s just like, your opinion, man? edit: It was a play on The Big Lebowski, folks.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#70The real "sleeper" might be https://huggingface.co/ibm-granite/granite-vision-4.1-4b if the benchmarks hold up for such a small model against frontier models for table & semantic k:v extraction.
Woah, is this part of the future of models? Basically little models you can use as tools.