Live data from Hacker News

A Knockout Blow for LLMs?

garymarcus.substack.com

21–30 of 49 posts

Re: A Knockout Blow for LLMs?

#21
post #2

In other news, water is wet. I don't think anybody who uses LLMs professionally day-to-day thinks that it can reason like human beings... If some people thought this, they fundamentally do not understand how LLMs work under the hood.

Oh buddy, step our of your bubble. There are people out there who swear by LLM being a modern day mesahiah. And no, this are not just SV VCs trying to sell their investments.

Sure but VCs always exaggerate. Remember the dot com bubble? Or "there's an app for that"?

The internet and smartphones were still extremely useful. There's no need to refute VC exaggeration. It's like writing articles to prove that perfume won't catch you Bradd Pitt. Nobody literally believes the adverts but that doesn't mean perfume is a lie.

I'm not saying that VCs only push good ideas - e.g. flying cars & web3 aren't going to work. Just that their claims are obviously exaggerated and can be ignored, even for useful ideas.

Re: A Knockout Blow for LLMs?

#22
From a quick glance it seems to be about spatial reasoning problems. I think there is good reasons for why it's tricky to become extremely good at these from being trained on text and static images. Future models being further multimodally trained with video and then physics simulators should deal with this much better I think.

There's a recent talk about this by Jim Fan from Nvidia https://youtu.be/_2NijXqBESI

Re: A Knockout Blow for LLMs?

#23
post #16

Oh, another LLM skepticism paper from Apple. This paper from last year doesn't age well due to rapid proliferation of reasoning models. https://machinelearning.apple.com/research/gsm-symbolic

Is apple putting out these papers just to justify their seeming inability to properly integrate them into their software?

Re: A Knockout Blow for LLMs?

#24
post #9

The first figure in the paper with Accuracy vs Complexity makes the whole point moot. The authors find that the performance of Claude 3.7 collapses around complexity 3 while Claude 3.7 thinking collapsed around complexity 7. A massive improvement in the complexity horizon that can be dealt with. It's real, it's quantitative, so what's the point of philosophical atguments about whether it is truly "reasoning" or not.…

> bigger models trained on bigger data with bigger reasoning posttraining and better distillation will push the horizons further and further

There is no evidence this is the case.

We could be in an era of diminishing returns where bigger models do not yield substantial improvements in quality but instead they become faster, cheaper and more resource efficient.

Re: A Knockout Blow for LLMs?

#25
post #17
post #5

"They're super expensive pattern matchers that break as soon as we step outside their training distribution" - I find it really weird that things like these are seen as some groundbreaking endgame discovery about LLMs LLMs have a real issues with polarisation. It's probably smart people saying all this stuff about knockout blows, and LLM uselessness, but I find them really useful. Is there some emperor's new clothes…

Marcus’s writing is from a scientific perspective, it’s not general artificial intelligence and probably not a meaningful path it GAI. But the ivory tower misses the point of how LLM improved the ability of regular people to interact with information and technology. While it might not be the grail they were seeking, it’s still a useful thing what will improve life and in turn be improved.

Marcus's writing is from the perspective of someone who is situated in the branch of AI that didn't work out - symbolic systems - and has a bit of an axe to grind against LLMs.

He's not always wrong, and sometimes useful as a contrarian foil, but not a source of much insight.

Re: A Knockout Blow for LLMs?

#26
post #7

Earlier quoted context omitted.

I think that these kind of papers are necessary to ground people back into reality - the hype machine is too strong to be left unguarded.

All the papers I’ve seen show models have limits. This is just an attention grab by lazy “researchers” cashing in on their Apple credentials.

Samy Bengio is the co-author of Torch. His credentials speak for themselves.

Re: A Knockout Blow for LLMs?

#27
Citing a few points to justify my own conclusion:

> Many (not all) humans screw up on versions of the Tower of Hanoi with 8 discs.

> LLMs are no substitute for good well-specified conventional algorithms.

> will continue have their uses, especially for coding and brainstorming and writing

> But anybody who thinks LLMs are a direct route to the sort AGI that could fundamentally transform society for the good is kidding themselves.

I agree with the assessment but disagree with the conclusion:

Being good at coding, writing, etc is precisely the sort of labor that is both “general intelligence” and will radically change society when clerical jobs are mechanized — and their ability to write (and interface with) classical algorithms to buttress their performance will only improve.

This is like when machines came for artisans.

Re: A Knockout Blow for LLMs?

#28
post #9

The first figure in the paper with Accuracy vs Complexity makes the whole point moot. The authors find that the performance of Claude 3.7 collapses around complexity 3 while Claude 3.7 thinking collapsed around complexity 7. A massive improvement in the complexity horizon that can be dealt with. It's real, it's quantitative, so what's the point of philosophical atguments about whether it is truly "reasoning" or not.…

> bigger models trained on bigger data with bigger reasoning posttraining and better distillation will push the horizons further and further There is no evidence this is the case. We could be in an era of diminishing returns where bigger models do not yield substantial improvements in quality but instead they become faster, cheaper and more resource efficient.

I would claim that o1 -> o3 is evidence of exactly that, and supposedly in half a year we will have even better reasoning models (further complexity horizon), so what could that be besides what I am describing.

Re: A Knockout Blow for LLMs?

#29
post #2

In other news, water is wet. I don't think anybody who uses LLMs professionally day-to-day thinks that it can reason like human beings... If some people thought this, they fundamentally do not understand how LLMs work under the hood.

There are quite a lot of options out at the extremes and both ends seem to presuppose things about consciousness that even specialist's in the field have debated for years.

I'm ok with thinking it's possible that some subset of consciousness might exist in LLMs while also being well aware of their limitations. Cognitive science has plenty of examples of mental impairments that show that there are individuals that lack some of the things that LLMs also lack that. We would hardly deny those individuals are conscious. The distinction for what is thought is lower but no less complex.

Before we had machines pushing at these boundaries, there were very learned people debating these issues, but it seems like now some gut instincts from people who have chatted to a bot for a bit are carrying the day.

Re: A Knockout Blow for LLMs?

#30
post #28

Earlier quoted context omitted.

> bigger models trained on bigger data with bigger reasoning posttraining and better distillation will push the horizons further and further There is no evidence this is the case. We could be in an era of diminishing returns where bigger models do not yield substantial improvements in quality but instead they become faster, cheaper and more resource efficient.

I would claim that o1 -> o3 is evidence of exactly that, and supposedly in half a year we will have even better reasoning models (further complexity horizon), so what could that be besides what I am describing.

Is there some breakthrough in reasoning between o1 and o3 that we are all missing.

And no one cares what we may have in the future. OpenAI etc already have an issue with credibility.

Post reply on HN