Live data from Hacker News

Granite 4.1: IBM's 8B Model Matching 32B MoE

firethering.com

11–20 of 223 posts

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#12
post #10
post #9

> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left

So are we saying it's fine that the article is written by an LLM as long as it doesn't have the tell-tale signs of LLMs?

It's more about curating the things you're publishing. Why would I bother reading what you couldn't bother to read?

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#13
post #10
post #9

> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left

So are we saying it's fine that the article is written by an LLM as long as it doesn't have the tell-tale signs of LLMs?

I don't really see reason to complain about tool use, so long as the result is cohesive, accurate and that ultimately means a human has at least read their own output before publishing. It's a bit like receiving a supposedly personal letter that starts "Dear [INSERT_FIRST_NAME_FIELD]," are you really going to read such a thing?

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#14
post #9

> Full stop. Why people don't edit out obvious sloppification and expect to still have readers left

Third line in to the article: "But there’s one result in the benchmarks I keep coming back to."

I hear this sort of thing all the time now on YouTube from media/news personalities:

“And that’s the part nobody seems to be talking about.”

"And here's what keeps me up at night."

“This is where the story gets complicated.”

“Here’s the piece that doesn’t quite fit.”

“And this is where the usual explanation starts to break down.”

“Here’s what I can’t stop thinking about.”

“The part that should worry us is not the obvious one.”

“And that’s where the real problem begins.”

“But the more interesting question is the one no one is asking.”

“And this is where things stop being simple.”

It doesn't really worry me but I think its interesting that LLM speak sounds so distinctive, and how willing these media personalities are to be so obvious in reading out on TV what the LLM spat out.

I've never studied what LLMs say in depth is it is interesting that my brain recognises the speech pattern so easily.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#16

Earlier quoted context omitted.

Yea, No doubt Qwen 3.6 open weights are far more strong

Why no doubt?

Because Qwen 3.6 pushes way above its weight. Granite 8B is impressive, but Qwen still wins on raw capability, especially for coding.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#18
If you really think about why MoE came into existence, its to save significant cost during training, I don't think there was any concrete evidence of performance gains for comparable MoE vs dense models. Over the years, I believe all the new techniques being employed in post training have made the models better.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#19

Earlier quoted context omitted.

Why no doubt?

Because Qwen 3.6 pushes way above its weight. Granite 8B is impressive, but Qwen still wins on raw capability, especially for coding.

Way above its weights.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#20

Earlier quoted context omitted.

Why no doubt?

Because Qwen 3.6 pushes way above its weight. Granite 8B is impressive, but Qwen still wins on raw capability, especially for coding.

You just asserted the same thing again. Why do you say this is the case?
Post reply on HN