Live data from Hacker News

Seven replies to the viral Apple reasoning paper and why they fall short

garymarcus.substack.com

241–250 of 331 posts

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#241

Earlier quoted context omitted.

If that's the point, shouldn't they ask the model to explain the principle for any number of discs? What's the benefit of a concrete application?

Because that would prove absolutely nothing. There are numerous examples of tower of Hanoi explanations in the training set.

How do you check that a human understood it and not simply memorised different approaches?

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#242
post #114

Earlier quoted context omitted.

I don’t need a tool that’s right maybe 70% of the time (and that’s me being optimistic). It needs to be right all the time or at least tell you when it doesn’t know for sure, instead of just making up something. Comparing it to going out in the streets and asking random people random questions is not a good comparison.

It might not fit your work, but there are tons of areas where “good enough” can still provide a lot of value. I’m sure you’d be thrilled with a tool that could correctly tell you if Apple’s stock was going up or down tomorrow 70% of the time.

I work in a mail room sending hard copy letters to customers. If I got my job right only 70% of the time then I’d be causing massive privacy breaches daily by sending the wrong personal information to the wrong customers.

Would you trust an AI that gets your banking transactions right only 70% of the time?

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#243

> 1. Humans have trouble with complex problems and memory demands. True! But incomplete. We have every right to expect machines to do things we can’t. [...] If we want to get to AGI, we will have to better. I don't get this argument. The paper is about "whether RLLMs can think". If we grant "humans make these mistakes too", but also "we still require this ability in our definition of thinking", aren't we saying "thin…

Agreed. But also his point about AGI is incorrect. AI that will perform on the level of average human in every task is AGI by definition.

The Hanoi Towers example demonstrates that SOTA RLMs struggle with tasks a pre-schooler solves.

The implication here is that they excel at things that occur very often and are bad at novelty. This is good for individuals (by using RLMs I can quickly learn about many other aspects of human body of knowledge in a way impossible/inefficient with traditional methods) but they are bad at innovation. Which, honestly, is not necessarily bad: we can offload lower-level tasks[0] to RLMs and pursue innovation as humans.

[0] Usual caveats apply: with time, the population of people actually good at these low-level tasks will diminish, just as we have very few Assembler programmers for Intel/AMD processors.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#244

Earlier quoted context omitted.

If the student could reference notes a fraction of the size of the LLM then I would not be convinced.

LLMs are (suspected) a few TB in size. Gemma 2 27B, one of the top ranked open source models, is ~60GB in size. LLama 405B is about 1TB. Mind you that they train on likely exabytes of data. That alone should be a strong indication that there is a lot more than memory going on here.

I'm not convinced by this argument. You can fit a bunch of books covering up to MSc level maths on less than 100MB. After that point, more books will mostly be redundant information so it doesn't need much more space for maths beyond that.

Similarly TBs of Twitter/Reddit/HN add near zero new information per comment.

If anything you can fit an enormous amount of information in 1MB - we just don't need to do it because storage is cheap.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#245
post #136

Earlier quoted context omitted.

>I think the answer to this question is certainly "Yes". It is unequivocally "No" . A good joint distribution estimator is always by definition a posteriori and completely incapable of synthetic a priori thought.

The human mind is an estimator too. The fact that the human mind can think in concepts, images AND words, and then compresses that into words for transmission, wheras LLMs think directly in words, is no object. If you watch someone reach a ledge, your mind will generate, based on past experience, a probabilistic image of that person falling. Then it will tie that to the concept of problem (self-attention) and start g…

Do you think language is sufficient to model reality (not just physical, but abstract) here?

I think not, we can get close, but there exists problems and situations beyond that, especially in mathematics and philosophy. And I don't a visual medium or combination of is sufficient either, there's a more fundamental, underlying abstract structure that we use to model reality.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#246

We built planes—critics said they weren't birds. We built submarines—critics said they weren't fish. Progress moves forward regardless. You have a choice: master these transformative tools and harness their potential, or risk being left behind by those who do. Pro tip: Endless negativity from the same voices won't help you adapt to what's coming—learning will.

Is there a name for this authorial voice and cadence? I see midwits posting exactly like this on twitter and linkedin; it's insufferable.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#247

We built planes—critics said they weren't birds. We built submarines—critics said they weren't fish. Progress moves forward regardless. You have a choice: master these transformative tools and harness their potential, or risk being left behind by those who do. Pro tip: Endless negativity from the same voices won't help you adapt to what's coming—learning will.

Indeed. Anyone who has built things with Claude Code (Opus 4) and/or something more than one-shot with o3 should be feeling the AGI at this point. Certainly there’s still many limitations, but progress is undoubtedly moving forward.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#248
post #135
post #64

Earlier quoted context omitted.

Yes, this one. Thanks

He doesn't say "that LLMs based on transformers can't create anything truly novel". Maybe he thinks that, maybe not, but what he says is that "today's systems" can't do that. He doesn't make any general statement about what transformer-based LLMs can or can't do; he's saying: we've interacted with these specific systems we have right now and they aren't creating genuinely novel things. That's a very different claim,…

[deleted]

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#249
post #45

Earlier quoted context omitted.

They can't create anything novel and it's patently obvious if you understand how they're implemented. But I'm just some anonymous guy on HN, so maybe this time I will just cite the opinion of the DeepMind CEO, who said in a recent interview with The Verge (available on YouTube) that LLMs based on transformers can't create anything truly novel.

Since when is reasoning synonymous with invention? All humans with a functioning brain can reason, but only a tiny fraction have or will ever invent anything.

Read what OP said "It’s patently obvious to me that LLMs can ... solve novel problems", this is what I was replying to. I see everyone is smarter here than researchers at DeepMind, without any proofs or credentials to back their claims.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#250

Earlier quoted context omitted.

> I’d expect a smart human to just say “that’s too much” or “that’s beyond my abilities” rather than do a best effort faulty answer)? That's what the models did. They gave the first 100 steps, then explained how it was too much to output all of it, and gave the steps one would follow to complete it. They were graded as "wrong answer" for this. --- Source: https://x.com/scaling01/status/1931783050511126954?t=ZfmpSxH..…

Why should we trust a guy with the following twitter bio to accurately replicate a scientific finding? >lead them to paradise >intelligence is inherently about scaling >be kind to us AGI Who even is this guy? He seems like just another r/singularity-style tech bro.

Not to be that guy but... clearly Ad Hominem.
Post reply on HN