Earlier quoted context omitted.
Yea, No doubt Qwen 3.6 open weights are far more strong
Why no doubt?
Granite 4.1: IBM's 8B Model Matching 32B MoE
21–30 of 223 posts
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#22> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#23> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#24> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left
So are we saying it's fine that the article is written by an LLM as long as it doesn't have the tell-tale signs of LLMs?
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#25If you really think about why MoE came into existence, its to save significant cost during training, I don't think there was any concrete evidence of performance gains for comparable MoE vs dense models. Over the years, I believe all the new techniques being employed in post training have made the models better.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#26Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#27Earlier quoted context omitted.
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
I think this kind of language predates widespread LLM use, and has been picked up from that kind of writing. It's a "and here's where it gets interesting" pattern that people like Malcolm Gladwell and Freakonomics have used, even if the same thing could be said in a way that makes it sound much less intriguing.
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#28If you really think about why MoE came into existence, its to save significant cost during training, I don't think there was any concrete evidence of performance gains for comparable MoE vs dense models. Over the years, I believe all the new techniques being employed in post training have made the models better.
But I don’t think it necessarily saved training cost; if it did, I’d be interested to learn how!
Re: Granite 4.1: IBM's 8B Model Matching 32B MoE
#29> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left
Third line in to the article: "But there’s one result in the benchmarks I keep coming back to." I hear this sort of thing all the time now on YouTube from media/news personalities: “And that’s the part nobody seems to be talking about.” "And here's what keeps me up at night." “This is where the story gets complicated.” “Here’s the piece that doesn’t quite fit.” “And this is where the usual explanation starts to break…
A writing teacher once excoriated me for saying that something was important. “Don’t tell me it’s important, show me, and let me decide, and if you do your job I’ll agree”
I don’t know how a completion can tell when it needs to do this. Mostly so far it doesn’t seem capable