Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

201–210 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#201

Earlier quoted context omitted.

The existence of an "inner monologue" isn't really a falsifiable claim. Some people claim to have one while other people claim not to, but we can't test the truth of these claims.

In this particular case, is there any reason why we simply can't take their word for it? This is not a case of where if I say "weak" or "strong", most people pick strong because no one wants to be weak, even if the context is unknown (nuclear force for example).

> "why we simply can't take their word for it"?

As someone who was involved in spiritual practice of "stopping internal dialogue" for years, I can tell you that one learns that that dialogue (or monologue, pretty much the same thing) is quite subtle and complex, essentially multi-layered.

Typically, when you think that you "think about nothing at all" it's just the most surface layer that has stopped, and more subtle talking to yourself is still going on. It takes training just to become able to notice and recognize it.

After all, it's just such a constant and monotone hum at the back of one's mind, one learns to completely ignore it.

So no, I would not take a word of people who were not trained to notice their internal monologue that they haven't any :-)

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#202

The idea that these word problems (and other LLM stumpers) are "easily solvable by humans" needs some empirical data behind it. Computer people like puzzles, and this kind of thing seems straightforward to them. I think the percentage of the general population who would get these puzzles right with the same time constraints LLMs are subjected to is much lower than the authors would expect, and that the LLMs are right…

Yeah, as someone with an education background I suspect GPT-4 is relatively close to the general public's performance on this problem. Many people would miss AIW, and almost all would miss AIW+. I'm about as good at this kind of thing as anyone and I'd need a minute with pencil and paper to handle AIW+; it's on par with the most difficult problems found on tests like the GRE. I wonder if these models, trained on data…

I wonder the same thing. If any academic reading this wants a paper idea:

1. Examine papers and other claims that an LLM gets something wrong that a human would have gotten wrong. How many of those claims have any citations about how many humans actually get it wrong? How many of those citations use the general population instead of the population of people who would be uniquely well-suited to answering the question correctly (i.e. people who signed up for the GRE are more likely to get GRE questions right than the general population).

2. For claims that are totally missing citations on human performance, run some tests with humans from the general population (or as close as you can get), and see how the LLMs compare.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#203

Earlier quoted context omitted.

In this particular case, is there any reason why we simply can't take their word for it? This is not a case of where if I say "weak" or "strong", most people pick strong because no one wants to be weak, even if the context is unknown (nuclear force for example).

> is there any reason why we simply can't take their word for it? because if we give them a problem to solve in their head and just give us the answer, they will. By problem I mean planning a trip, a meal, how to pay the mortgage, etc. It's impossible to plan without an internal monologue. Even if some people claim theirs is 'in images'.

> It's impossible to plan without an internal monologue.

Of course it isn't impossible, and this is backed by what we know about paleoanthropology and other instances of cognition in animals - humans were making stone tools millions of years ago, which takes planning in the form of imagining what you want the tool to look like and how you will do it and what it will be used for. It's exceedingly likely we had this ability long before complex speech evolved. Apes also use and make tools, which would require planning, and I don't think they have an internal monologue going on. birds from the corvid family can do some pretty advanced problem solving that requires planning. Cetaceans might be an exception, because they appear to have some form of language, but this is a pretty wild claim not really backed by any kind of science as we understand it today.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#204

Earlier quoted context omitted.

> is there any reason why we simply can't take their word for it? because if we give them a problem to solve in their head and just give us the answer, they will. By problem I mean planning a trip, a meal, how to pay the mortgage, etc. It's impossible to plan without an internal monologue. Even if some people claim theirs is 'in images'.

'It's impossible to plan without an internal monologue.' - Sorry, but I disagree with this. I have no 'internal voice' or monologue - whenever I see a problem, my brain actually and fully models it using images. I believe 25% of the population doesn't have the internal monologue which you're referring to and this has been tested and confirmed. I highly recommend listening to this Lex Friedman podcast episode to get a…

Sure, I do mention thinking in images in my original comment and count it as some type of internal monologue. I personally do not believe it's all images, as that would preclude using highly abstract concepts. But I might be wrong, and it might be 100% images. That being said, it does count as an internal monologue.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#205
post #119

Earlier quoted context omitted.

It’s commonly conjectured that the emergence of human-level reasoning wouldn’t have been possible without the development of language. Personally, I’m able to suppress “word thoughts” in my head (for a short time), but then I lose almost all of my reasoning ability. I could imagine that reasoning is language-based even when it’s not conscious for some people. An internal process being there, and being conscious of it…

Maybe, but symbolic thought can get pretty far away from what we generally call "language." I bet you can reason 1+3x=22 pretty easily without any words whatsoever, or the sound of one ascending octave after another, or the approximate G-force induced on your body if you take the next turn without applying the brakes. All of these forms of reasoning are true and useful calculations: when we talk about "intuition" wha…

>I bet you can reason 1+3x=22 pretty easily without any words whatsoever

I've tried to do it, but I can't. I had to do something like "ok, so we subtract one from both sides and then it's easy, 3*7=21". Maybe I could do 2+8 but I still think the word ten "aloud".

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#206

Earlier quoted context omitted.

> is there any reason why we simply can't take their word for it? because if we give them a problem to solve in their head and just give us the answer, they will. By problem I mean planning a trip, a meal, how to pay the mortgage, etc. It's impossible to plan without an internal monologue. Even if some people claim theirs is 'in images'.

> It's impossible to plan without an internal monologue. How can science make this claim if it can't prove (or disprove) the existence of an internal monologue?

Well, I remember Richard Feynman came up with an interesting experiment. He found he could not count objects when he read aloud some text at the same time. He had to name the numbers, and it was impossible if he was already engaging his speech.

He thought this was universal, but doing this experiment with friends, he discovered a guy who could count while reading aloud. So when Feynman asked him, how he does this, turned out that the guy instead of "pronouncing" numbers was "seeing" colored numbers in his imagination, so his speech was not involved.

I supposed this experiment can be modified and generalized, and at least to shed some light on this problem.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#207

Earlier quoted context omitted.

> is there any reason why we simply can't take their word for it? because if we give them a problem to solve in their head and just give us the answer, they will. By problem I mean planning a trip, a meal, how to pay the mortgage, etc. It's impossible to plan without an internal monologue. Even if some people claim theirs is 'in images'.

> It's impossible to plan without an internal monologue. Of course it isn't impossible, and this is backed by what we know about paleoanthropology and other instances of cognition in animals - humans were making stone tools millions of years ago, which takes planning in the form of imagining what you want the tool to look like and how you will do it and what it will be used for. It's exceedingly likely we had this ab…

Animals can not manipulate abstract concepts nor can they do long-term plans. No crow can plan an international trip spanning a couple of weeks and two change-overs. And some people definitely can't do it start to end, but they can at least plan the first 5-7 steps.

Also, maybe inner monologue is not a binary have/have not, but maybe it is on a continuum.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#209
post #18

For anyone considering reading the paper and like me don't normally read papers like this, open the PDF and think you don't have time to read it due to its length. The main part of the paper is the first 10 pages and a fairly quick read. On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong co…

Its not an AI hype. A hype is defined as something which gets oversold: "promote or publicize (a product or idea) intensively, often exaggerating its benefits." Just yesterday I visited a google cloud summit and one person from bosch told the audiance how they are now able to work with less external agencies like texting, graphicsdesigner and photographers for their materials. It already saves money, has real impacts…

Your post is complete hype, all about people saying things instead of showing things that've actually been done.

For me, 2024 was the LLM exposed as basically pure hype year.

There is no expert of any field I follow online where they're posting up results from AI tooling for any other reason than to show how awful it is. I consider myself an expert in software, and LLMs specifically have only caused me great pain.

Even the one situation where you describe someone describing the ability to work in an absolute vacuum sounds like a huge negative to me. The recent push for DEI policies were even ostensibly about the importance of people of diverse backgrounds and viewpoints working together.

The most important thing you're missing a perspective of scale on is the step you describe as "quality check it". On things I don't know, and have attempted to enlist an LLMs help on, in every case I have had to go back and just actually learn how something works, after time wasted struggling with subtle wrongness in the output.

At least I have the background expertise to do that, however, I have seen a Jr dev's mind get literally rotted by too much time in pure LLM land. Besides the cost of rewriting their code, the company was now the proud owner of a young dev with a mind filled with nonsense.

How do you even weigh the cost of fixing a corrupted human mind?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#210

Earlier quoted context omitted.

Even if this were all true, it points to a fundamental risk of using LLM's for important tasks, which is that it is not at all clear to a user that this prompt would cause a problem. The LLM doesn't say "I'm sorry Dave, I just can't do that", it just complies with it and gets the wrong answer. You can always make excuses for the LLM afterwards, but software with hidden risks like this would not be considered good or…

While LLMs have incredible potential, and are even downright useful in their current format, they have the rather nasty tendency to confidently present bullshit that passes the smell test for most people. When it's wrong, it's not just wrong , by design, it's wrong but sounds plausible. Considering most people with subject matter expertise aren't going to consult such a bot for their own areas of expertise, that mean…

> As soon as I asked "are you giving me instructions for Gemini Advanced or Bard?" it was like "OH LOL WOOPS!! YOU GOT ME BRO! XD I CAN'T DO ANY OF THAT! LOL." Which, honestly, is great. Being able to evaluate its answers to realize it's wrong is really neat.

Did it evaluate its answers, or did your expression of doubt cause the eager-to-please language model to switch from "generate (wrong) instructions because that's what the user asked for" to "acknowledge an error because that's what the user asked for"?

How many times have we seen "Oops, you're right! 2 + 2 is actually 5! I apologize for saying it was 4 earlier!"

Post reply on HN