Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

221–230 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#221
post #161

Earlier quoted context omitted.

With that definition even bacteria have inner monologue.

Can bacteria imagine pictures? Do they have emotions? Why does this matter? Stop being so pedantic. We're talking about a progression of ideas . Talking in your head is one form of ideas, but people can easily solve problems by imagining them.

Initial thesis was - inner monologue is required for reasoning. If you define inner monologue to include everything brains do - the initial thesis becomes a tautology.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#222

Earlier quoted context omitted.

Seriously? You think individuals are incapable of reasoning without training first?

Yes, seriously. Some examples: An individual without training cannot reliably separate cause from effect, or judge that both events A and B may have a common root cause. Similarly, people often confuse conditionals for causation. People often have difficulty reasoning about events based on statistical probabilities. Remember, the average person in North America is far more terrified of a terror attack than an acciden…

You mean without training, people cannot frame answers in the terms you've learned from training. Well, why are you surprised?

If you think reasoning is limited to the frameworks you learned from a book, you live in a small world.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#223

Earlier quoted context omitted.

> is there any reason why we simply can't take their word for it? because if we give them a problem to solve in their head and just give us the answer, they will. By problem I mean planning a trip, a meal, how to pay the mortgage, etc. It's impossible to plan without an internal monologue. Even if some people claim theirs is 'in images'.

'It's impossible to plan without an internal monologue.' - Sorry, but I disagree with this. I have no 'internal voice' or monologue - whenever I see a problem, my brain actually and fully models it using images. I believe 25% of the population doesn't have the internal monologue which you're referring to and this has been tested and confirmed. I highly recommend listening to this Lex Friedman podcast episode to get a…

Can you draw a picture of an example of what you see when you think about something?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#224

The idea that these word problems (and other LLM stumpers) are "easily solvable by humans" needs some empirical data behind it. Computer people like puzzles, and this kind of thing seems straightforward to them. I think the percentage of the general population who would get these puzzles right with the same time constraints LLMs are subjected to is much lower than the authors would expect, and that the LLMs are right…

[deleted]

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#225

This is a cool one, but I know of other such "failures". For example, try to ask (better in Russian), how many letters "а" are there in Russian word "банан". It seems all models answer with "3". Playing with it reveals that apparently LLMs confuse Russian "банан" with English "banana" (same meaning). Trying to get LLMs to produce a correct answer results is some hilarity. I wonder if each "failure" of this kind deser…

No current LLM understands words, nor letters. They all have input and output tokens, that roughly correspond to syllabes and letter groupings. Any kind of task involving counting letters or words is outside their realistic capabilities. LLMs are a tool, and like any other tool, they have strengths and weaknesses. Know your tools.

I understand that, but the article we are discussing points out that LLMs are so good on many tasks, and so good at passing tests, that many people will be tricked into blindly "taking their word for granted" -- even people who should know better: our brain is a lazy machine, and if something works almost always it starts to assume it works always.

I mean, you can ask an LLM to count letters in thousand of words, and pretty much always it will come with the correct answer! So far I don't know of any word other than "банан" that breaks this function.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#226
post #41

Earlier quoted context omitted.

> In many ways, this is very obvious and routine to people who use these systems with a critical understanding of how they work. The last part of that is the problem and why a paper like this is critical. These systems are being pushed onto people who don't understand how they work. CEO's and other business leaders are being pushed to use AI. Average users are being shown it in Google search results. Etc etc. People…

Sure, but even these people... the failures are so common, and often very obvious. Consider a CEO who puts a press briefing in and asks some questions about it, it's not uncommon for those answers to be obviously wrong on any sort of critical reflection. We arent dealing with a technology that is 99.9% right in our most common use cases, so that we need to engineer some incredibly complex problem to expose the flaw.…

> We arent dealing with a technology that is 99.9% right in our most common use cases, so that we need to engineer some incredibly complex problem to expose the flaw. Rather, in most cases there is some obvious flaw. It's a system that requires typically significant "prompt engineering" to provide the reasoning the system otherwise lacks.

Right -- we are a long way from "this is a very nuanced error" being the dominant failure.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#227

Citation 40 is the longest list of authors I have ever seen. That is one way to help all your friends get tenure.

Tenure committees have to write a report detailing every single one of your papers and what your contribution was.

There are always "bean counters" added somewhere in the process. There are many places where the person lists their number of publications and that is all most people will ever see.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#228
post #54

Earlier quoted context omitted.

I honestly thought about this recently when I was trying to see the limits of Claude Opus. Some of the problems I gave it, what if instead of telling it to solve the problem I asked it to write the script and then give me the command and inputs needed to properly run it to get the answer I needed. That way instead of relying on the LLM to do properly analysis of the numbers it just needs to understand enough to write…

I'm not sure what you mean by 'it will not scale well.' When we humans learn that we make a mistake - we make a note and we hold the correct answer in memory - the next time we're prompted with a similar prompt, we can use our old memories to come up with the correct solution. I just did a simple test for this same exact problem using ChatGPT 3.5: 'Can you reformulate the following problem using Prolog? When you exec…

What if you prompt it "You seem to have accidentally included Alice. The correct answer should be 4"?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#229

Earlier quoted context omitted.

Even language is not sequential.

Tell me more?

Language is only sequential in the form it is transmitted (verbally). There is no reason that sequential statements are generated sequentially in the brain. Quite the opposite, really, if you consider rules of grammar.

I really wish most of the LLM folks just took a few courses in linguistics. It would avoid a lot of noise.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#230

Earlier quoted context omitted.

The existence of an "inner monologue" isn't really a falsifiable claim. Some people claim to have one while other people claim not to, but we can't test the truth of these claims.

In this particular case, is there any reason why we simply can't take their word for it? This is not a case of where if I say "weak" or "strong", most people pick strong because no one wants to be weak, even if the context is unknown (nuclear force for example).

> In this particular case, is there any reason why we simply can't take their word for it?

My concern is that if we take their word for it, we're actually buying into two assumptions which (AFAIK) are both unproven:

1. That "Internal Monologues" (not consciously forced by attention) exist in the first place, as opposed to being false-memories generated after-the-fact by our brain to explain/document a non-language process that just occurred. (Similar to how our conscious brains pretend that we were in control of certain fast reflexes.)

2. Some people truly don't have them, as opposed to just not being aware of them.

Post reply on HN