Live data from Hacker News

Understanding Reasoning LLMs

magazine.sebastianraschka.com

51–60 of 196 posts

Re: Understanding Reasoning LLMs

#51
post #35
post #30

Earlier quoted context omitted.

> If you would listen to most of the people critical of LLMs saying they're a "stochastic parrot" - it should be impossible for them to do better than random on any out of distribution problem. Even just changing one number to create a novel math problem should totally stump them and result in entirely random outputs, but it does not. You don't seem to understand how they work, they recurse their solution meaning if…

> You don't seem to understand how they work I don't think anyone understands how they work- these type of explanations aren't very complete or accurate. Such explanations/models allow one to reason out what types of things they should be capable of vs incapable of in principle regardless of scale or algorithm tweaks, and those predictions and arguments never match reality and require constant goal post shifting as t…

> I don't think anyone understands how they work

Yes we do, we literally built them.

> We understand how we brought them about via setting up an optimization problem in a specific way, that isn't the same at all as knowing how they work.

You're mistaking "knowing how they work" with "understanding all of the emergent behaviors of them"

If I build a physics simulation, then I know how it works. But that's a separate question from whether I can mentally model and explain the precise way that a ball will bounce given a set of initial conditions within the physics simulation which is what you seem to be talking about.

Re: Understanding Reasoning LLMs

#52
Are there any websites that show the results of popular models on different benchmarks, which are explained in plain language? As an end user, I'd love a quick way to compare different models suitability for different tasks.

Re: Understanding Reasoning LLMs

#53

Great post, but every time I read something like this I feel like I am living in a prequel to the Culture.

Is that bad? The Culture is pretty cool I think. I doubt the real thing would be so similar to us but who knows.

Oh no, I’d live on an Orbital in a heartbeat. No, it’s just that all of these kinds of posts make me feel like we’re about to live through “The Bad Old Days”.

Re: Understanding Reasoning LLMs

#54
post #19

Nice article. >Whether and how an LLM actually "thinks" is a separate discussion. The "whether" is hardly a discussion at all. Or, at least one that was settled long ago. "The question of whether a computer can think is no more interesting than the question of whether a submarine can swim." --Edsger Dijkstra

It's interesting if you're asking the computer to think, which we are.

It's not interesting if you're asking it to count to a billion.

Re: Understanding Reasoning LLMs

#55

Earlier quoted context omitted.

ie, the people that AI is dumb? Or you are saying I'm in a cult for being pro it - I'm definitely part of that cult - the "we already have agi and you have to contort yourself into a pretzel to believe otherwise" cult. Not sure if there is a leader though.

You think we have AGI? What makes you think that?

By knowing what each of the letters stand for

Re: Understanding Reasoning LLMs

#56

I like Raschka's writing, even if he is considerably more optimistic about this tech than I am. But I think it's inappropriate to claim that models like R1 are "good at deductive or inductive reasoning" when that is demonstrably not true, they are incapable of even the simplest "out-of-distribution" deductive reasoning: https://xcancel.com/JJitsev/status/1883158738661691878 They are certainly capable of doing is a wi…

"researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation" - Rich Sutton

Re: Understanding Reasoning LLMs

#57
post #50
post #44

Earlier quoted context omitted.

Well we do know pretty much exactly what they do, don't we? What surprises us is the behaviors coming out of that process. But surprise isn't magic, magic shouldn't even be on the list of explanations to consider.

Magic wasn’t mentioned here. We don’t understand the emerging behavior, in the sense that we can’t reason well about it and make good predictions about it (which would allow us to better control and develop it). This is similar to how understanding chemistry doesn’t imply understanding biology, or understanding how a brain works.

Exactly, we don't understand, but we want to believe it's reasoning, which would be magic.

Re: Understanding Reasoning LLMs

#58

Earlier quoted context omitted.

You think we have AGI? What makes you think that?

By knowing what each of the letters stand for

Well that’s disappointing. It was an extraordinary claim that really interested me.

Thought I was about to be learn!

Instead, I just met an asshole.

Re: Understanding Reasoning LLMs

#59
post #49

Earlier quoted context omitted.

anyone saying an LLM is a stochastic parrot doesn't understand them... they are just parroting what they heard.

A good literary production. I would have been proud of it had I thought of it, but it's a path to observe a strong "whataboutery" element that if we use "stochastic parrot" as shorthand and you dislike the term, now you understand why we dislike the constant use of "infer", "reason" and "hallucinate" Parrots are self aware, complex reasoning brains which can solve problems in geometry, tell lies, and act socially or…

I don't like wading into this debate when semantics are very personal/subjective. But to me, it seems like almost a sleight of hand to add the stochastic part, when actually they're possibly weighted more on the parrot part. Parrots are much more concrete, whereas the term LLM could refer to the general architecture.

The question to me seems: If we expand on this architecture (in some direction, compute, size etc.), will we get something much more powerful? Whereas if you give nature more time to iterate on the parrot, you'd probably still end up with a parrot.

There's a giant impedance mismatch here (time scaling being one). Unless people want to think of parrots being a subset of all animals, and so 'stochastic animal' is what they mean. But then it's really the difference of 'stochastic human' and 'human'. And I don't think people really want to face that particular distinction.

Re: Understanding Reasoning LLMs

#60
post #59
post #49

Earlier quoted context omitted.

A good literary production. I would have been proud of it had I thought of it, but it's a path to observe a strong "whataboutery" element that if we use "stochastic parrot" as shorthand and you dislike the term, now you understand why we dislike the constant use of "infer", "reason" and "hallucinate" Parrots are self aware, complex reasoning brains which can solve problems in geometry, tell lies, and act socially or…

I don't like wading into this debate when semantics are very personal/subjective. But to me, it seems like almost a sleight of hand to add the stochastic part, when actually they're possibly weighted more on the parrot part. Parrots are much more concrete, whereas the term LLM could refer to the general architecture. The question to me seems: If we expand on this architecture (in some direction, compute, size etc.),…

I'm sure both of you know this, but "stochastic parrot" refers to the title of a research article that contained a particular argument about LLM limitations that had very little to do with parrots.
Post reply on HN