Earlier quoted context omitted.
I’m so tired of hearing this be repeated, like the whole “LLMs are _just_ parrots” thing. It’s patently obvious to me that LLMs can reason and solve novel problems not in their training data. You can test this out in so many ways, and there’s so many examples out there. ______________ Edit for responders, instead of replying to each: We obviously have to define what we mean by "reasoning" and "solving novel problems"…
> It’s patently obvious that LLMs can reason and solve novel problems not in their training data. Would you care to tell us more ? « It’s patently obvious » is not really an argument, I could say just as well that everyone know LLM can’t resonate or think (in the way we living beings do).
Seven replies to the viral Apple reasoning paper and why they fall short
131–140 of 331 posts
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#132Earlier quoted context omitted.
That’s the opposite of reasoning tho. Ai bros want to make people believe LLM are smart but they’re not capable of intelligence and reasoning. Reasoning mean you can take on a problem you’ve never seen before and think of innovative ways to solve it. LLM can only replicate what is in its data, it can in no way think or guess or estimate what will likely be the best solution, it can only output a solution based on a p…
You're assuming we're saying LLMs can't reason. That's not what we're saying. They can execute reasoning-like processes when they've seen similar patterns, but this breaks down when true novel reasoning is required. Most people do the same thing. Some poeple can come up with novel solutions to new problems, but LLMs will choke. Here's an example: Prompt: "Let's try a reasoning test. Estimate how many pianos there are…
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#133Earlier quoted context omitted.
I don’t read GM as saying that LLMs “don’t work” in a practical sense. He acknowledges that they have useful applications. Indeed, if they didn’t work at all, why would he be advocating for regulating their use? He just doesn’t think they’re close to AGI.
The funny thing is, if you asked “what is AGI” 5 years ago, most people would describe something like o3.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#134Earlier quoted context omitted.
Agree. Both sides of the argument are unsatisfying. They seem like quantitative answers to a qualitative question.
"Have we created machines that can do something qualitatevely similar to that part of us that can correlate known information and pattern recognition to produce new ideas and solutions to problems -- that part we call thinking?" I think the answer to this question is certainly "Yes". I think the reason people deny this is because it was just laughably easy in retrospect. In mid-2022 people were like. "Wow this GPT3 t…
I cannot afford to consider whether you are right because I am a slave to capital, and therefore may as well be a slave to capital's LLMs. The same goes for you.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#135Earlier quoted context omitted.
"I don't think today's systems can invent, you know, do true invention, true creativity, hypothesize new scientific theories. They're extremely useful, they're impressive, but they have holes." Demis Hassabis On The Future of Work in the Age of AI (@ 2:30 mark) https://www.youtube.com/watch?v=CRraHg4Ks_g
Yes, this one. Thanks
Again, for all I know maybe he does believe that transformer-based LLMs as such can't be truly creative. Maybe it's true, whether he believes it or not. But that interview doesn't say it.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#136Earlier quoted context omitted.
Agree. Both sides of the argument are unsatisfying. They seem like quantitative answers to a qualitative question.
"Have we created machines that can do something qualitatevely similar to that part of us that can correlate known information and pattern recognition to produce new ideas and solutions to problems -- that part we call thinking?" I think the answer to this question is certainly "Yes". I think the reason people deny this is because it was just laughably easy in retrospect. In mid-2022 people were like. "Wow this GPT3 t…
It is unequivocally "No". A good joint distribution estimator is always by definition a posteriori and completely incapable of synthetic a priori thought.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#137Earlier quoted context omitted.
Agreed. But also his point about AGI is incorrect. AI that will perform on the level of average human in every task is AGI by definition.
That very much depends on which AGI definition you are using. I imagine there are a dozen or so variants out there. See also "AI" and "agents" and (apparently) "vibe coding" and pretty much every other piece of jargon in this field.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#138Earlier quoted context omitted.
That’s the opposite of reasoning tho. Ai bros want to make people believe LLM are smart but they’re not capable of intelligence and reasoning. Reasoning mean you can take on a problem you’ve never seen before and think of innovative ways to solve it. LLM can only replicate what is in its data, it can in no way think or guess or estimate what will likely be the best solution, it can only output a solution based on a p…
You're assuming we're saying LLMs can't reason. That's not what we're saying. They can execute reasoning-like processes when they've seen similar patterns, but this breaks down when true novel reasoning is required. Most people do the same thing. Some poeple can come up with novel solutions to new problems, but LLMs will choke. Here's an example: Prompt: "Let's try a reasoning test. Estimate how many pianos there are…
I gave your prompt to o3 pro, and this is what I got without any hints:
Historic shipwrecks (1850 → 1970)
• ~20 000 deep water wrecks recorded since the age of steam and steel
• 10 % were passenger or mail ships likely to carry a cabin class or saloon piano
• 1 piano per such vessel 20 000 × 10 % × 1 ≈ 2 000
Modern container losses (1970 → today)
• ~1 500 shipping containers lost at sea each year
• 1 in 2 000 containers carries a piano or electric piano
• Each piano container holds ≈ 5 units
• 50 year window 1 500 × 50 / 2 000 × 5 ≈ 190
Coastal disasters (hurricanes, tsunamis, floods)
• Major coastal disasters each decade destroy ~50 000 houses
• 1 house in 50 owns a piano
• 25 % of those pianos are swept far enough offshore to sink and remain (50 000 / 50) × 25 % × 5 decades ≈ 1 250
Add a little margin for isolated one offs (yachts, barges, deliberate dumping): ≈ 300
Best guess range: 3 000 – 5 000 pianos are probably resting on the seafloor worldwide.Re: Seven replies to the viral Apple reasoning paper and why they fall short
#139Earlier quoted context omitted.
I’m so tired of hearing this be repeated, like the whole “LLMs are _just_ parrots” thing. It’s patently obvious to me that LLMs can reason and solve novel problems not in their training data. You can test this out in so many ways, and there’s so many examples out there. ______________ Edit for responders, instead of replying to each: We obviously have to define what we mean by "reasoning" and "solving novel problems"…
I've done this excercise dozens of times because people keep saying it, but I can't find an example where this is true. I wish it was. I'd be solving world problems with novel solutions right now. People make a common mistake by conflating "solving problems with novel surface features" with "reasoning outside training data." This is exactly the kind of binary thinking I mentioned earlier.
Can you reason? Yes? Then why haven't you cured cancer? Let's not have double standards.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#140You have a choice: master these transformative tools and harness their potential, or risk being left behind by those who do.
Pro tip: Endless negativity from the same voices won't help you adapt to what's coming—learning will.