Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

171–180 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#172
post #87

Earlier quoted context omitted.

I appear to be reasoning at times but I have mostly no idea what I am talking about. I hit a bunch of words and concepts in the given context and thus kind of hallucinate sense. Given a few months of peace of mind and enough money for good enough food, I could actually learn to reason without sounding like a confused babelarian. Reasoning is mostly a human convention supported by human context that would have been a…

Yeah, I think these chatbots are just too sure of themselves. They only really do "system 1 thinking" and only do "system 2 thinking" if you prompt them to. If I ask gpt-4o the riddle in this paper and tell it to assume its reasoning contains possible logical inconsistencies and to come up with reasons why that might be then it does correctly identify the problems with its initial answer and arrives at the correct on…

> After you answer the riddle please review your answer assuming that you have made a logical inconsistency in each step and explain what that inconsistency is. Even if you think there is none do your best to confabulate a reason why it could be logically inconsistent.

LLMs are fundamentally incapable of following this instruction. It is still model inference, no matter how you prompt it.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#173

Citation 40 is the longest list of authors I have ever seen. That is one way to help all your friends get tenure.

I didn't count it, but I think papers from high energy particle physics have it beat. Some have over 5k authors.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#174

Earlier quoted context omitted.

If you really think about what an LLM is you would think there is no way that leads to general purpose AI. At the same time though they are already doing way more than we thought they could. Maybe people were surprised by what OpenAI achieved so now they are all just praying that with enough compute and the right model AGI will emerge.

> If you really think about what an LLM is you would think there is no way that leads to general purpose AI It is an autoregressive sequence predictor/generator. Explain to me how humans are fundamentally different

Even language is not sequential.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#175
post #10

Earlier quoted context omitted.

There must be a name for the new phenomenon, of which your post is an example, of: 1. Someone expresses that an LLM cannot do some trivial task. 2. Another person declares that they cannot do the task, thereby defending the legitimacy of the LLM. As a side note, I cannot believe that the average person who can navigate to a chatgpt prompter would fail to correctly answer this question given sufficient motivation to d…

Many people, especially on this site, really want LLMs to be everything the hype train says and more. Some have literally staked their future on it so they get defensive when people bring up that maybe LLMs aren’t a replacement for human cognition. The number of times I’ve heard “but did you try model X” or “humans hallucinate too” or “but LLMs don’t get sleep or get sick” is hilarious.

Yes. Seems like some users here experience true despair when you suggest that the LLM approach might have a hard limit that means LLMs will be useful but never revolutionary.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#176

Earlier quoted context omitted.

> If you really think about what an LLM is you would think there is no way that leads to general purpose AI It is an autoregressive sequence predictor/generator. Explain to me how humans are fundamentally different

Even language is not sequential.

Tell me more?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#177
The idea that these word problems (and other LLM stumpers) are "easily solvable by humans" needs some empirical data behind it. Computer people like puzzles, and this kind of thing seems straightforward to them. I think the percentage of the general population who would get these puzzles right with the same time constraints LLMs are subjected to is much lower than the authors would expect, and that the LLMs are right in line with human-level reasoning in this case.

(Of course, I don't have a citation either, but I'm not the one writing the paper.)

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#178
post #18

For anyone considering reading the paper and like me don't normally read papers like this, open the PDF and think you don't have time to read it due to its length. The main part of the paper is the first 10 pages and a fairly quick read. On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong co…

The problem is a good chunk of the global population is also not reasoning and thinking in any sense of the word. Logical reasoning is a higher order skill that often requires formal training. It's not a natural ability for human beings.

Seriously? You think individuals are incapable of reasoning without training first?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#179

Earlier quoted context omitted.

Why are so many people so insistent on saying this? I’m guessing you are in denial that we can make a simulated reasoning machine?

People keep saying it because that's literally how LLMs work. They run Montecarlo sampling over a very impressive latent linguistic space. These models are not fundamentally different than the Markov chains of yore except that these latent representations are incredibly powerful. We haven't even started to approach the largest problem which is moving beyond what is essentially a greedy token level search of this ling…

The best compression is some form of understanding

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#180

Earlier quoted context omitted.

> If you really think about what an LLM is you would think there is no way that leads to general purpose AI It is an autoregressive sequence predictor/generator. Explain to me how humans are fundamentally different

AI needs to see thousands or millions of images of a cat before they reliably can identify one. The fact that a child needs to only see one example of a cat to know what a cat is from then on seems to point to humans having something very different.

Humans train on continuous video. Even our most expensive models are, in terms of training set size, far behind what an infant processes in the first year of their life.

EDIT: and it takes human children a couple years to reliably identify a cat. My 2.5 y.o. daughter still confuses cats with small dogs, despite living under one roof with a cat.

Post reply on HN