Live data from Hacker News

Seven replies to the viral Apple reasoning paper and why they fall short

garymarcus.substack.com

71–80 of 331 posts

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#71
post #36

AI hype-bros like to complain that real AI experts are too much concerned about debunking current AI then improving it - but the truth is that debunking bad AI IS improving AI. Science is a process of trial and error which only works by continuously questioning the current state.

Indeed. I completely agree with this. My objection to the whole thing is the AI hype bros, which is really the funding solicitation facade over everything rather the truth, only has one outcome and that is that it cannot be sustained. At that point all investor confidence disappears, the money is gone and everyone loses access to the tools that they suddenly built all their dependencies on because it's all proprietar…

A lot of the best internet services came around in the decade after the dot-com crash. There is a chance Anthropic or OpenAI may not survive when funding suddenly dries up, but existing open weight models won't be majorly impacted. There will always be someone willing to host DeepSeek for you if you're willing to pay.

And while it will be sad to see model improvements slow down when the bubble bursts there is a lot of untapped potential in the models we already have. Especially as they become cheaper and easier to run

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#72

Earlier quoted context omitted.

I don't understand what people mean when they say that AI is being hyped. AI is at the point where you can have a conversation with it about almost anything, and it will answer more intelligently than 90% of people. That's incredibly impressive, and normal people don't need to be sold on it. They're just naturally impressed by it.

I don’t need a tool that’s right maybe 70% of the time (and that’s me being optimistic). It needs to be right all the time or at least tell you when it doesn’t know for sure, instead of just making up something. Comparing it to going out in the streets and asking random people random questions is not a good comparison.

> I don’t need a tool that’s right maybe 70% of the time (and that’s me being optimistic).

Where are you getting this from? 70%?

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#73

AI hype-bros like to complain that real AI experts are too much concerned about debunking current AI then improving it - but the truth is that debunking bad AI IS improving AI. Science is a process of trial and error which only works by continuously questioning the current state.

To hammer one point though, you have to understand that researcher are desensitized to minor novel improvement that translate to great value products. While obviously studying and assessing the limitations of AI is crucial, to the general public its capabilities are just so amazing, they can't fathom why we should think about limitations. Optimizing what we have is bette than rethinking the whole process.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#74

The key insight is that LLMs can 'reason' when they've seen similar solutions in training data, but this breaks down on truly novel problems. This isn't reasoning exactly, but close enough to be useful in many circumstances. Repeating solutions on demand can be handy, just like repeating facts on demand is handy. Marcus gets this right technically but focuses too much on emotional arguments rather than clear explanat…

That’s the opposite of reasoning tho. Ai bros want to make people believe LLM are smart but they’re not capable of intelligence and reasoning. Reasoning mean you can take on a problem you’ve never seen before and think of innovative ways to solve it. LLM can only replicate what is in its data, it can in no way think or guess or estimate what will likely be the best solution, it can only output a solution based on a p…

You're assuming we're saying LLMs can't reason. That's not what we're saying. They can execute reasoning-like processes when they've seen similar patterns, but this breaks down when true novel reasoning is required. Most people do the same thing. Some poeple can come up with novel solutions to new problems, but LLMs will choke. Here's an example:

Prompt: "Let's try a reasoning test. Estimate how many pianos there are at the bottom of the sea."

I tried this on three advanced AIs* and they all choked on it without further hints from me. Claude then said:

    Roughly 3 million shipwrecks on ocean floors globally
    Maybe 1 in 1000 ships historically carried a piano (passenger ships, luxury vessels)
    So ~3,000 ships with pianos sunk
    Average maybe 0.5 pianos per ship (not all passenger areas had them)
    Estimate: ~1,500 pianos
*Claude Sonnet 4, Google Gemini 2.5 and GPT 4o

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#75

I'm glad to read articles like this one, because I think it is important that we pour some water on the hype cycle If we want to get serious about using these new AI tools then we need to come out of the clouds and get real about their capabilities Are they impressive? Sure. Useful? Yes probably in a lot of cases But we cannot continue the hype this way, it doesn't serve anyone except the people who are financially i…

I don't understand what people mean when they say that AI is being hyped. AI is at the point where you can have a conversation with it about almost anything, and it will answer more intelligently than 90% of people. That's incredibly impressive, and normal people don't need to be sold on it. They're just naturally impressed by it.

AI, in the form of LLMs, can be a useful tool.

It is still being vastly overhyped, though, by people attempting to sell the idea that we are actually close to an AGI "singularity".

Such overhype is usually easy to handwave away as like not my problem. Like, if investors get fooled into thinking this is anything like AGI, well, a fool and his money and all that. But investors aside this AI hype is likely to have some very bad real world consequences based on the same hype-men selling people on the idea that we need to generate 2-4 times more power than we currently do to power this godlike AI they are claiming is imminent.

And even right now there's massive real world impact in the form of say, how much grok is polluting Georgia.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#76
post #35

The key insight is that LLMs can 'reason' when they've seen similar solutions in training data, but this breaks down on truly novel problems. This isn't reasoning exactly, but close enough to be useful in many circumstances. Repeating solutions on demand can be handy, just like repeating facts on demand is handy. Marcus gets this right technically but focuses too much on emotional arguments rather than clear explanat…

I’m so tired of hearing this be repeated, like the whole “LLMs are _just_ parrots” thing. It’s patently obvious to me that LLMs can reason and solve novel problems not in their training data. You can test this out in so many ways, and there’s so many examples out there. ______________ Edit for responders, instead of replying to each: We obviously have to define what we mean by "reasoning" and "solving novel problems"…

[deleted]

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#77

AI hype-bros like to complain that real AI experts are too much concerned about debunking current AI then improving it - but the truth is that debunking bad AI IS improving AI. Science is a process of trial and error which only works by continuously questioning the current state.

Can you please make your substantive points without name-calling or swipes? This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#78
post #5

In case anyone else missed the original paper (and discussion): https://news.ycombinator.com/item?id=44203562

Thanks! Macroexpanded:

The Illusion of Thinking: Strengths and limitations of reasoning models [pdf] - https://news.ycombinator.com/item?id=44203562 - June 2025 (269 comments)

Also this: A Knockout Blow for LLMs? - https://news.ycombinator.com/item?id=44215131 - June 2025 (48 comments)

Were there others?

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#79

Earlier quoted context omitted.

How actually useful are they though? We've had more than a year now of saying these things 10X knowledge workers and creatives, so.... where is the output? Is there a new office suite I can try? 10 times as many mobile apps? A huge new library of ebooks? Is this actually in practice producing things beyond Ghibli memes and RETVRN nostalgia slop?

I think it largely depends on what you’re writing. I’ve had it reply to corporate emails which is good since I need to sound professional not human. If I’m coding it still needs a lot of baby sitting and sometimes I’m much faster than it.

So this would be an interesting output to measure but I have no idea how we would do that: has the volume of corporate email gone up? Or the time spent creating it gone down?

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#80
post #25

Earlier quoted context omitted.

Gary Marcus isn't about "getting real", it's making a name for himself as a contrarian to the popular AI narrative. This article may seem reasonable, but here he's defending a paper that in his previous article he called "A knockout blow for LLMs". Many of his articles seem reasonable (if a bit off) until you read a couple dozen a spot a trend.

> Gary Marcus isn't about "getting real", it's making a name for himself as a contrarian to the popular AI narrative. That's an odd standard. Not wanting to be wrong is a universal human instinct. By that logic, every person who ever took any position on LLMs is automatically untrustworthy. After all, they made a name for themselves by being pro- or con-. Or maybe a centrist - that's a position too. Either he makes g…

Marcus' points routinely fail to pass scrutiny, nobody in the field takes him seriously. If you seek real scientifically interesting LLM criticism, read François Chollet and his Arc AGI series of evals.
Post reply on HN