Live data from Hacker News

Seven replies to the viral Apple reasoning paper and why they fall short

garymarcus.substack.com

151–160 of 331 posts

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#151
post #122

Earlier quoted context omitted.

[flagged]

Sorry - what do you mean by yud-cult? Searching google didn’t help me (as far as I can tell) - I view LW from an outside perspective as well, but don’t understand the reference

They're referring to the founder of that website, Eliezer Yudkowsky, who is controversial due to his 2023 Time article that called for a complete halt on the development of AI.

https://en.m.wikipedia.org/wiki/Eliezer_Yudkowsky

https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-no...

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#153
post #40

Most of the objections and their counterarguments seem like either poor objections (e.g. ad hominem against the first listed author) or seem to be subsumed under point 5. It’s annoying that most of this post focuses so much effort on discussing most of the other objections when the important discussion is the one to be had in point 5: I.e. to what extent are LLMs able to reliably make use of writing code or using log…

I don't think most of the objections are poor at all apart from 3, it's this article that seems to make lots of strawmans. Especially the first objection is often heard because people claim "this paper proves LLMs don't reason". The author moves goalposts and is arguing against about whether LLMs lead to AGI, which is already a strawman for those arguments. And in addition, he even seems to misunderstand AGI, thinkin…

It's especially weird argument considering that LLMs are already ahead of humans in Tower of Hanoi

No one cares about Towers of Hanoi. Nor do they care about any other logic puzzles like this. People want AIs that solve novel problems for their businesses. The kind of problems regular business employees solve every single day yet LLMs make a mess of.

The purpose of the Apple paper is not to reveal the fact that LLMs routinely fail to solve these problems. Everyone who uses them already knows this. The paper is an argument for why this happens (lack of reasoning skills).

No number of demonstrations of LLMs solving well-known logic puzzles (or other problems humans have already solved) will prove reasoning. It's not interesting at all to solve a problem that humans have already solved (with working software to solve every instance of the problem).

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#154
post #52

Good article giving some critique to Apple's paper and Gary Marcus specifically. https://www.lesswrong.com/posts/5uw26uDdFbFQgKzih/beware-gen...

What gets me, and the author talks about it in the post, is that people will readily attribute correct answers to "its in the training set" but nobody says anything about incorrect answers that are in the training set. LLMs get stuff in the training set wrong all the time, but nobody uses it as evidence that it probably can't lean too hard on it's memorization for complex questions it does get right.

It puts LLMs in an impossible position; if they are right, they memorized it, if they are wrong, they cannot reason.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#155
post #25

I'm glad to read articles like this one, because I think it is important that we pour some water on the hype cycle If we want to get serious about using these new AI tools then we need to come out of the clouds and get real about their capabilities Are they impressive? Sure. Useful? Yes probably in a lot of cases But we cannot continue the hype this way, it doesn't serve anyone except the people who are financially i…

Gary Marcus isn't about "getting real", it's making a name for himself as a contrarian to the popular AI narrative. This article may seem reasonable, but here he's defending a paper that in his previous article he called "A knockout blow for LLMs". Many of his articles seem reasonable (if a bit off) until you read a couple dozen a spot a trend.

I was very put off by his article "A knockout blow for LLMs?", especially all the fuss he was making about using his own name as a verb to mean debunking AI hype...

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#156
post #35

Earlier quoted context omitted.

I’m so tired of hearing this be repeated, like the whole “LLMs are _just_ parrots” thing. It’s patently obvious to me that LLMs can reason and solve novel problems not in their training data. You can test this out in so many ways, and there’s so many examples out there. ______________ Edit for responders, instead of replying to each: We obviously have to define what we mean by "reasoning" and "solving novel problems"…

So far they cannot even answer questions which are straight up fact checking and search engine like queries. Reasoning means they would be able to work through a problem and generate a proof they way a student might.

So if they have bad memory, then they must be reasoning to get the correct answer for the problems they do solve?

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#157

Earlier quoted context omitted.

Sorry - what do you mean by yud-cult? Searching google didn’t help me (as far as I can tell) - I view LW from an outside perspective as well, but don’t understand the reference

They're referring to the founder of that website, Eliezer Yudkowsky, who is controversial due to his 2023 Time article that called for a complete halt on the development of AI. https://en.m.wikipedia.org/wiki/Eliezer_Yudkowsky https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-no...

Yudkowsky is controversial for much more than an article from 2023.

Yudkowsky lacks credentials and MIRI and its adjacents have proven to be incestuous organizations when it comes to the rationalist cottage industry, one that has a serious problem with sexual abuse and literal cults.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#158
post #100
post #63

Earlier quoted context omitted.

[flagged]

I see the opposite, the wide majority of people commenting on Hacker News seem now very favorable to LLMs.

I think a common opinion is that it's useful as a research and code generation tool, and that it has some really negative effects on the Internet and society in general. Since the discussion on HN is often focused on coding the first aspect is just a bit more visible.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#159
I find it weird that people are taking the original paper to be some kind of indictment against llms. It's not like LLMs failing at doing Hanoi tower problem at higher levels is new, the paper took an existing method that was done before.

It was simply comparing the effectiveness of reasoning and non reasoning models on the same problem.

Re: Seven replies to the viral Apple reasoning paper and why they fall short

#160

> 5. A student might complain about a math exam requiring integration or differentiation by hand, even though math software can produce the correct answer instantly. The teacher’s goal in assigning the problem, though, isn’t finding the answer to that question (presumably the teacher already know the answer), but to assess the student’s conceptual understanding. Do LLM’s conceptually understand Hanoi? That’s what the…

If the student could reference notes a fraction of the size of the LLM then I would not be convinced.

LLMs are (suspected) a few TB in size.

Gemma 2 27B, one of the top ranked open source models, is ~60GB in size. LLama 405B is about 1TB.

Mind you that they train on likely exabytes of data. That alone should be a strong indication that there is a lot more than memory going on here.

Post reply on HN