Live data from Hacker News

Seven replies to the viral Apple reasoning paper and why they fall short

garymarcus.substack.com

181–190 of 331 posts

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#181
post #136

Earlier quoted context omitted.

"Have we created machines that can do something qualitatevely similar to that part of us that can correlate known information and pattern recognition to produce new ideas and solutions to problems -- that part we call thinking?" I think the answer to this question is certainly "Yes". I think the reason people deny this is because it was just laughably easy in retrospect. In mid-2022 people were like. "Wow this GPT3 t…

>I think the answer to this question is certainly "Yes". It is unequivocally "No" . A good joint distribution estimator is always by definition a posteriori and completely incapable of synthetic a priori thought.

The human mind is an estimator too.

The fact that the human mind can think in concepts, images AND words, and then compresses that into words for transmission, wheras LLMs think directly in words, is no object.

If you watch someone reach a ledge, your mind will generate, based on past experience, a probabilistic image of that person falling. Then it will tie that to the concept of problem (self-attention) and start generating solutions, such as warning them or pulling them back etc.

LLMs can do all this too, but only in words.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#182

I think the Apple paper is practically a hack job - the problem was set up in such a way that the reasoning models must do all of their reasoning before outputting any of their results. Imagine a human trying to solve something this way: you’d have to either memorize the entire answer before speaking or come up with a simple pattern you could do while reciting that takes significantly less brainpower - and past a cer…

Forcing reasoning is analogous to requiring a student to show their work when solving a problem if im understanding the paper correctly. > you’d have to either memorize the entire answer before speaking or come up with a simple pattern you could do while reciting that takes significantly less brainpower This part i dont understand. Why would coming up with an algorithm (e.g. a simple pattern) and reciting it be impos…

> The paper doesnt mention the models coming up with the algorithm at all AFAIK.

And that's because they specifically hamstrung their tests so that the LLMs were not "allowed" to generate algorithms.

If you simply type "Give me the solution for Towers of Hanoi for 12 disks" into chatGPT it will happily give you the answer. It will write program to solve it, and then run that program to produce the answer.

But according to the skeptical community - that is "cheating" because it's using tools. Nevermind that it is the most effective way to solve the problem.

https://chatgpt.com/share/6845f0f2-ea14-800d-9f30-115a3b644e...

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#183

Earlier quoted context omitted.

Agreed. But also his point about AGI is incorrect. AI that will perform on the level of average human in every task is AGI by definition.

the average human is good at something, and sucks at almost everything. Human performance at chess and average performance at chess differ by 7 orders of magnitude.

Your standard model of human needs a little bit of fine tuning for most games.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#184

> 1. Humans have trouble with complex problems and memory demands. True! But incomplete. We have every right to expect machines to do things we can’t. [...] If we want to get to AGI, we will have to better. I don't get this argument. The paper is about "whether RLLMs can think". If we grant "humans make these mistakes too", but also "we still require this ability in our definition of thinking", aren't we saying "thin…

Agreed. But also his point about AGI is incorrect. AI that will perform on the level of average human in every task is AGI by definition.

AGI should perform on the level of an experienced professional in every task. The average human is useless for pretty much everything but capable of learning to perform almost any task, given enough motivation and effort.

Or perhaps AGI should be able to reach the level of an experienced professional in any task. Maybe a single system can't be good at everything, if there are inherent trade-offs in learning to perform different tasks well.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#185

Earlier quoted context omitted.

"Have we created machines that can do something qualitatevely similar to that part of us that can correlate known information and pattern recognition to produce new ideas and solutions to problems -- that part we call thinking?" I think the answer to this question is certainly "Yes". I think the reason people deny this is because it was just laughably easy in retrospect. In mid-2022 people were like. "Wow this GPT3 t…

> The key is to not become a slave to them. To realize that our worth consists not in our ability to think. And that we are more than that. I cannot afford to consider whether you are right because I am a slave to capital, and therefore may as well be a slave to capital's LLMs. The same goes for you.

I am not a slave to capital. I am a slave to the harsh nature of the world.

I get too hot in summer and too cold in winter. I die of hunger. I am harassed by critters of all sorts.

And when my bed breaks, to keep my fragile spine from straining at night, I _want_ some trees to be cut, some mattresses to be provisioned, some designers to be provisioned etc. And capital is what gets me that, from people I will never meet, who wouldn't blink once if I died tomorrow.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#186
It's easy to check if a blackbox AI can reason: give it a checkerboard pattern, or something more complex, and see if it can come up with a compact formula that generates this pattern. You can't bullshit your way thru this problem, and it's easy to verify the answer, yet none of these so-called researchers attempt to do this.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#187
post #52

Good article giving some critique to Apple's paper and Gary Marcus specifically. https://www.lesswrong.com/posts/5uw26uDdFbFQgKzih/beware-gen...

  > this is a preprint that has not been peer reviewed.
This conversation is peer review...

You don't need a conference for something to be peer reviewed, you only need... peers...

In fact, this paper is getting more peer review than most works. Conferences are notoriously noisy as reviewers often don't care and are happy to point out criticisms. All works have valid criticisms... Finding criticisms is the easy part. The hard part is figuring out if these invalidate the claims or not.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#188

Earlier quoted context omitted.

Agreed. But also his point about AGI is incorrect. AI that will perform on the level of average human in every task is AGI by definition.

AGI should perform on the level of an experienced professional in every task. The average human is useless for pretty much everything but capable of learning to perform almost any task, given enough motivation and effort. Or perhaps AGI should be able to reach the level of an experienced professional in any task. Maybe a single system can't be good at everything, if there are inherent trade-offs in learning to perfor…

For comparison, the average person can't print Hello World in python. Your average programmer (probably) can.

It's surprisingly simple to be above average in most tasks. Which people often confuse with having expertise. It's probably pretty easy to get into the 80th percentile of most subjects. That won't make you the 80th percentile of people that do the thing, but most people don't. I'd wager 80th percentile is still amateur.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#190

Earlier quoted context omitted.

I think it's very widely accepted definition and there's really no competing definitions either as far as I know. While some people might think AGI means superintelligence, it's only because they've heard the term but never bothered to look up what it means.

Doesn't the "G" in AGI stand for "General" as in "Generally Good at everything"?

I think the G is what really screws things up. I thought it was, as good as the general human, but upon googling it has a defined meaning among researchers. There appears to be confusion all over the place tho.

General-Purpose (Wide Scope): It can do many types of things.

Generally as Capable as a Human (Performance Level): It can do what we do.

Possessing General Intelligence (Cognitive Mechanism): It thinks and learns the way a general intelligence does.

So, for researchers, general intelligence is characterized by: applying knowledge from one domain to solve problems in another, adapting to novel situations without being explicitly programmed for them, and: having a broad base of understanding that can be applied across many different areas.

Post reply on HN