Live data from Hacker News

The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

ml-site.cdn-apple.com

211–220 of 276 posts

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#211
post #77

Earlier quoted context omitted.

I saw your comment and counted — in May I took a Waymo thirty times.

Waymo is a popular argument in self-driving cars, and they do well. However, Waymo is Deep Blue of self-driving cars. Doing very well in a closed space . As a result of this geofencing, they have effectively exhausted their search space, hence they work well as a consequence of lack of surprises. AI works well when search space is limited, but General AI in any category needs to handle a vastly larger search space, a…

I suspect that Waymo car's could operate in a lot more areas than they do. The issue is that Waymo are trying to sell the service of safe travel and not a car with an addon you can pay for which doesn't actually work.

In other words, since they accept liability for their cars it's not in their interest to roll out the service too fast. It makes more sense to do it slow and steady.

It's not really a strong argument that their technology is incapable of working in general areas.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#212
post #92

Earlier quoted context omitted.

> I figure it's hard to argue that that is not at least somewhat intelligent? The fact that this technology can be very useful doesn't imply that it's intelligent. My argument is about the language used to describe it, not about its abilities. The breakthroughs we've had is because there is a lot of utility from finding patterns in data which humans aren't very good at. Many of our problems can be boiled down to this…

>The machine has no semantic understanding of what the data represents. How do you define "semantic understanding" in a way that doesn't ultimately boil down to saying they don't have phenomenal consciousness? Any functional concept of semantic understanding is captured to some degree by LLMs. Typically when we attribute understanding to some entity, we recognize some substantial abilities in the entity in relation t…

You don't need phenomenal consciousness. You need consistency.

LLMs are not consistent. This is unarguable. They will produce a string of text that says they have solved a problem and/or done a thing when neither is true.

And sometimes they will do it over and over, even when corrected.

Your last paragraph admits this.

Tokenisation on its own simply cannot represent reality accurately and reliably. It can be tweaked so that specific problems can appear solved, but true AI would be based on a reliable general strategy which solves entire classes of problems without needing this kind of tweaking.

It's clear we're nowhere close to that.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#213
post #208
post #202

All "reasoning" models hit a complexity wall where they completely collapse to 0% accuracy. No matter how much computing power you give them, they can't solve harder problems. This research suggests we're not as close to AGI as the hype suggests. Current "reasoning" breakthroughs may be hitting fundamental walls that can't be solved by just adding more data or compute. Apple's researchers used controllable puzzle env…

I think this might be part of the reason Apple is “behind” on generative AI … LLMs have not really proven to be useful outside of relatively niche areas such as coding assistants, legal boiler plate and research, and maybe some data science/analysis which I’m less familiar with Other “end user” facing use cases have so far been comically bad or possibly harmful, and they just don’t meet the quality bar for inclusion…

You are being too charitable to Apple. They are just being the fox from the Sour Grapes fable.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#215
Isn't it a matter of training? This is the way I think about LLMs. They just "learned" so much context that they spit out tokens one after the other based on that context. So if it hallucinates it's because it understood the context wrong or doesn't grasp the nuance in the context. The more complicated the task the higher chance of hallucinations. Now I don't know if this can be improved with more training but that is the only tool we have.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#216
post #202

All "reasoning" models hit a complexity wall where they completely collapse to 0% accuracy. No matter how much computing power you give them, they can't solve harder problems. This research suggests we're not as close to AGI as the hype suggests. Current "reasoning" breakthroughs may be hitting fundamental walls that can't be solved by just adding more data or compute. Apple's researchers used controllable puzzle env…

No matter how much computing power you give them, they can't solve harder problems.

Why would anyone ever expect otherwise?

These models are inherently handicapped and always will be in terms of real world experience. They have no real grasp of things people understand intuitively --- like time or money or truth ... or even death.

They only *reality* they have to work from is a flawed statistical model built from their training data.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#217
post #208
post #202

All "reasoning" models hit a complexity wall where they completely collapse to 0% accuracy. No matter how much computing power you give them, they can't solve harder problems. This research suggests we're not as close to AGI as the hype suggests. Current "reasoning" breakthroughs may be hitting fundamental walls that can't be solved by just adding more data or compute. Apple's researchers used controllable puzzle env…

I think this might be part of the reason Apple is “behind” on generative AI … LLMs have not really proven to be useful outside of relatively niche areas such as coding assistants, legal boiler plate and research, and maybe some data science/analysis which I’m less familiar with Other “end user” facing use cases have so far been comically bad or possibly harmful, and they just don’t meet the quality bar for inclusion…

Apple prides itself on building high quality user experiences. One can argue over whether that’s true anymore or ever was but it’s very clear they pride themselves on that. This is why Apple tends to “be late” on many features. I think like you’re saying it’s becoming even more clear they aren’t seeing a UX they are willing to ship to customers. We’ll see with WWDC if they found something in the last year to make it better but this paper seems to indicate they haven’t.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#218

Earlier quoted context omitted.

>The machine has no semantic understanding of what the data represents. How do you define "semantic understanding" in a way that doesn't ultimately boil down to saying they don't have phenomenal consciousness? Any functional concept of semantic understanding is captured to some degree by LLMs. Typically when we attribute understanding to some entity, we recognize some substantial abilities in the entity in relation t…

You don't need phenomenal consciousness. You need consistency. LLMs are not consistent. This is unarguable. They will produce a string of text that says they have solved a problem and/or done a thing when neither is true. And sometimes they will do it over and over, even when corrected. Your last paragraph admits this. Tokenisation on its own simply cannot represent reality accurately and reliably. It can be tweaked…

Consistency is a strange criteria seeing as humans aren't very consistent either. Intelligent beings can make mistakes.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#219
Reasoning models are just wrappers over the base model. It was pretty obvious it wasn’t actually reasoning but rather just refining the results using some kind of reasoning like heuristic. At least that’s what I assumed when they were released and you couldn’t modify the system prompt.
Post reply on HN