Live data from Hacker News

Granite 4.1: IBM's 8B Model Matching 32B MoE

firethering.com

131–140 of 223 posts

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#131

On the topic of local models, is there a good equivalent to something like Claude's chat interface? I've recently started transitioning to open models after getting fed up with Claude's usage limits (I'm not in a position to drop $200/month), and for coding tasks Kimi 2.6 has been about the same as Sonnet in my experience. The only thing I've found myself missing is a nice interface to ask it questions and have it he…

llama-server from the llama.cpp package has a local web interface.

yes. I've used it a lot. its very simple and good

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#132

People complain a lot about LLM-written articles, but the human comments here on HN are far worse. Mostly a bunch of people extremely proud of themselves for not reading an LLM-written article, and then a bunch of people who take it at face value and make the model seem almost useful, and one comment that actually looked at other benchmarks. Good 'ol humanity, good at.. being emotional... and not doing analysis.....…

The pro LLM rant is weird, LLMs "hallucinate" in creating detailed elaborate lies, the frontier models still do this egregiously, an LLM written article by default has 0 value since every single line could be true or it could be a convincingly crafted lie, every line has to be fact checked I'm using Gemini 3.1 pro to help me research my thesis, it still with search enabled and on pro mode, invents entire papers that…

If you are asking an LLM to cite it's sources you are wasting your time and degrading the quality of the response. LLMs have no inherent mechanism for "knowledge source tracking", because that isn't at all how they work. We're trying to get there with agentic stacks, but it's still too new.

For sparse knowledge tasks, where you know that the model can't possibly have much training because even humans themselves don't have much knowledge there, use it as a brainstorming partner, not as a source. Or put relevant papers in it's context to help you eval those papers in relation to your work. But it's just going to hurt itself in confusion trying to tie fuzzy ideas to sparse sources embedded in pages upon pages of mildly related google search results.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#133

Earlier quoted context omitted.

The pro LLM rant is weird, LLMs "hallucinate" in creating detailed elaborate lies, the frontier models still do this egregiously, an LLM written article by default has 0 value since every single line could be true or it could be a convincingly crafted lie, every line has to be fact checked I'm using Gemini 3.1 pro to help me research my thesis, it still with search enabled and on pro mode, invents entire papers that…

No, you're being weird (why are you calling people weird anyway, not helpful). You're complaining about facts that have been true since words have been written on paper. If you read the article with the same criticality you read any other article you wont have the problem you complain about. The reality is, you're only complaining because you hate ai. Cool, but dont dress it up and resort to name calling to browbeat…

If I read something and cannot tell that it is AI generated, then there's no problem.

If it has AI tells then I wont bother to continue reading because it was either written by an AI or it was written by someone who can't tell the difference.

Either way it's probably a poor piece of writing.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#134
The Granite 4.1 3B model is only 2GB from Unsloth: https://huggingface.co/unsloth/granite-4.1-3b-GGUF

I ran it in LM Studio and got a pleasingly abstract pelican on a bicycle (genuinely not bad for a tiny 3B model - it can at least output valid SVG): https://gist.github.com/simonw/5f2df6093885a04c9573cf5756d34...

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#135

On the topic of local models, is there a good equivalent to something like Claude's chat interface? I've recently started transitioning to open models after getting fed up with Claude's usage limits (I'm not in a position to drop $200/month), and for coding tasks Kimi 2.6 has been about the same as Sonnet in my experience. The only thing I've found myself missing is a nice interface to ask it questions and have it he…

I've been mostly using LM Studio for this recently. Ollama has an OK chat UI now too. 'brew install llama.cpp' gets you 'llama-server' which provides quite a good web UI.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#136

On the topic of local models, is there a good equivalent to something like Claude's chat interface? I've recently started transitioning to open models after getting fed up with Claude's usage limits (I'm not in a position to drop $200/month), and for coding tasks Kimi 2.6 has been about the same as Sonnet in my experience. The only thing I've found myself missing is a nice interface to ask it questions and have it he…

With Ollama* you can use Claude Code with `ollama launch claude`

* https://docs.ollama.com/integrations/claude-code

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#137

People complain a lot about LLM-written articles, but the human comments here on HN are far worse. Mostly a bunch of people extremely proud of themselves for not reading an LLM-written article, and then a bunch of people who take it at face value and make the model seem almost useful, and one comment that actually looked at other benchmarks. Good 'ol humanity, good at.. being emotional... and not doing analysis.....…

>Mostly a bunch of people extremely proud of themselves for not reading an LLM-written article

I'm not sure it's proud as much as people voicing displeasure with the uncertainty about what went into the LLM prompt. This may have been a 1 sentence prompt, or it may have been some well researched background that simply reformatted it. Why waste minutes-hours on verifying it if it's possible someone could have spent 10 second on it? It's very easy to see their point.

People seem to indicate people they disagree with voicing their opinion about anything lately is some auto-fellatio, I wonder what causes them to think this way.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#138
post #76
post #51

Earlier quoted context omitted.

Qwen 3.6 burns it to the ground. it was not even a challenge. Gemma4 seriously fails at toolcalls and agentic works. It got all messed up after 2-3 turns of Vibecoding.

Gemma 4 31b was working ok for me; but it was consuming tons of memory on SWA checkpoints, I had to turn them way down, and as a 31b dense model is fairly slow on a Strix Halo. I did have a lot of tool calling issues on 26b-a4b, though. The Qwen models are quite solid though.

What are you using to run it vllm, llama.cpp or other?

Can you share your switches and approach for using tools?

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#139
post #118

Earlier quoted context omitted.

"The article makes some good points about model design" But how can I tell if those are good points or not? I don't want to invest time in reading something if the presence of those "good points" depends on a roll of the dice.

even calling it roll of the dice is an assumption. Can you point anything you find as mistake?

You expect people to read every single excretion, which can be generated faster than I can read,just to find the rare gem that might exist?

The problem is that in the past it took multiple times more effort and hours to write something than it took to read. That served two purposes:

1. Lazy people just looking for an audience were effectively gatekept from drowning the world with their every vapid thought.

2. Because supply was many times slower than consumption it was viable to give most articles a chance: the author could not drown me in a deluge even if they wanted to.

Having the criteria now that the author should spend at least as much effort creating the piece as they expect the reader expend reading it is a damn useful bar: instead of reading 1000 AI articles just to find the one good one, I can simply read 10 human authored articles and be certain that 9 of them have something worthwhile.

Re: Granite 4.1: IBM's 8B Model Matching 32B MoE

#140

People complain a lot about LLM-written articles, but the human comments here on HN are far worse. Mostly a bunch of people extremely proud of themselves for not reading an LLM-written article, and then a bunch of people who take it at face value and make the model seem almost useful, and one comment that actually looked at other benchmarks. Good 'ol humanity, good at.. being emotional... and not doing analysis.....…

> people complain a lot about LLM-written articles, but the human comments here on HN are far worse.

No, they aren't.

You are comparing writing produced with little to no effort to writing produced with the minimal effort required to communicate.

It's reasonable for people to complain that they are presented material that not even the author thought was worth the effort.

Post reply on HN