Live data from Hacker News

VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

arxiv.org

221–226 of 226 posts

Re: VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

#221

I still cannot trust evaluations and benchmarks. How can you prove that the test datasets are truly unseen examples? I think the only way to prove that these models are truly as good as they claim is to wait and see if they are getting adopted in practice.

> the only way to prove that these models are truly as good as they claim

It would be actually to progress towards the solution of the "black box" problem, the goal of "transparency".

You have to implement a reasoner (etc.), you conceive the best architecture for it - then implement and test it.

Re: VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

#222

GRPO skips the value network that makes PPO expensive — it scores candidates relative to each other within a group. that's what makes verifiable-reward training practical at 3B scale

Why was this downvoted.

Re: VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

#223
post #171

Beats Opus 4.5 on reasoning you say? Prompt: If A goes to B who then goes to C, can A send something to C? Response: We need to interpret best. The phrase "If A goes to B who then goes to C, can A send something to C?" could be a puzzle about the concept of sending something (like passing a ball) and the relationships. Scenario: A gives something to B, and B passes it on to C. Question: Can A also give the same thing…

I am a human and I don't know how to interpret this prompt.

> If A goes to B who then goes to C, can A send something to C?

"If John travels to the location of Mary, and later Mary reaches Paul, is John enabled to have an item delivered to Paul"

Or what did you mean.

Re: VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

#224
post #16

There is some base level of intelligence any model needs to be useful, even in narrow tasks. Could you teach a 5 year old to drive a car? A 10 year old? A 12 year old? To drive a car requires being able to read, to have judgement about ice or rainy conditions, to anticipate a child running after a ball. By the time a human in in their mid teens they have acquired the base knowledge... Small models need to have enough…

As far as just the general driving skill: https://www.youtube.com/watch?v=RZ_0ImDYrPY

Re: VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

#225
post #120

Earlier quoted context omitted.

It feels sometimes like optimizations are only starting.

I’m beginning to suspect the closed SOTA labs were doing all these optimisations, keeping quiet about it, and just charging us out the yinyang for inference.

Also as much landgrab as possible for data centres, infrastructure, that can kepe it running for the next 5-15 years.

Why does an M1 Max continue to remain capable, if not more capable with every passing year with LM Studio? :)

Re: VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO

#226
post #191

I tried actually talking to it. It reminded me of GPT-2.

It's not supposed to be a chat model. It is crazy good at math.

Yeah, I thought so too. It's supposedly good at coding, too. But, if its English responses are insane, why should I trust it to understand my programming-related instructions?

Well, only one way to find out! I mean, maybe it only says insane things and goes in circles if you ask non-coding questions, but it doesn't exactly inspire confidence.

Post reply on HN