Earlier quoted context omitted.
With that definition even bacteria have inner monologue.
Can bacteria imagine pictures? Do they have emotions? Why does this matter? Stop being so pedantic. We're talking about a progression of ideas . Talking in your head is one form of ideas, but people can easily solve problems by imagining them.
Simple tasks showing reasoning breakdown in state-of-the-art LLMs
221–230 of 393 posts
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#222Earlier quoted context omitted.
Seriously? You think individuals are incapable of reasoning without training first?
Yes, seriously. Some examples: An individual without training cannot reliably separate cause from effect, or judge that both events A and B may have a common root cause. Similarly, people often confuse conditionals for causation. People often have difficulty reasoning about events based on statistical probabilities. Remember, the average person in North America is far more terrified of a terror attack than an acciden…
If you think reasoning is limited to the frameworks you learned from a book, you live in a small world.
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#223Earlier quoted context omitted.
> is there any reason why we simply can't take their word for it? because if we give them a problem to solve in their head and just give us the answer, they will. By problem I mean planning a trip, a meal, how to pay the mortgage, etc. It's impossible to plan without an internal monologue. Even if some people claim theirs is 'in images'.
'It's impossible to plan without an internal monologue.' - Sorry, but I disagree with this. I have no 'internal voice' or monologue - whenever I see a problem, my brain actually and fully models it using images. I believe 25% of the population doesn't have the internal monologue which you're referring to and this has been tested and confirmed. I highly recommend listening to this Lex Friedman podcast episode to get a…
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#224The idea that these word problems (and other LLM stumpers) are "easily solvable by humans" needs some empirical data behind it. Computer people like puzzles, and this kind of thing seems straightforward to them. I think the percentage of the general population who would get these puzzles right with the same time constraints LLMs are subjected to is much lower than the authors would expect, and that the LLMs are right…
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#225This is a cool one, but I know of other such "failures". For example, try to ask (better in Russian), how many letters "а" are there in Russian word "банан". It seems all models answer with "3". Playing with it reveals that apparently LLMs confuse Russian "банан" with English "banana" (same meaning). Trying to get LLMs to produce a correct answer results is some hilarity. I wonder if each "failure" of this kind deser…
No current LLM understands words, nor letters. They all have input and output tokens, that roughly correspond to syllabes and letter groupings. Any kind of task involving counting letters or words is outside their realistic capabilities. LLMs are a tool, and like any other tool, they have strengths and weaknesses. Know your tools.
I mean, you can ask an LLM to count letters in thousand of words, and pretty much always it will come with the correct answer! So far I don't know of any word other than "банан" that breaks this function.
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#226Earlier quoted context omitted.
> In many ways, this is very obvious and routine to people who use these systems with a critical understanding of how they work. The last part of that is the problem and why a paper like this is critical. These systems are being pushed onto people who don't understand how they work. CEO's and other business leaders are being pushed to use AI. Average users are being shown it in Google search results. Etc etc. People…
Sure, but even these people... the failures are so common, and often very obvious. Consider a CEO who puts a press briefing in and asks some questions about it, it's not uncommon for those answers to be obviously wrong on any sort of critical reflection. We arent dealing with a technology that is 99.9% right in our most common use cases, so that we need to engineer some incredibly complex problem to expose the flaw.…
Right -- we are a long way from "this is a very nuanced error" being the dominant failure.
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#227Citation 40 is the longest list of authors I have ever seen. That is one way to help all your friends get tenure.
Tenure committees have to write a report detailing every single one of your papers and what your contribution was.
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#228Earlier quoted context omitted.
I honestly thought about this recently when I was trying to see the limits of Claude Opus. Some of the problems I gave it, what if instead of telling it to solve the problem I asked it to write the script and then give me the command and inputs needed to properly run it to get the answer I needed. That way instead of relying on the LLM to do properly analysis of the numbers it just needs to understand enough to write…
I'm not sure what you mean by 'it will not scale well.' When we humans learn that we make a mistake - we make a note and we hold the correct answer in memory - the next time we're prompted with a similar prompt, we can use our old memories to come up with the correct solution. I just did a simple test for this same exact problem using ChatGPT 3.5: 'Can you reformulate the following problem using Prolog? When you exec…
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#229Earlier quoted context omitted.
Even language is not sequential.
Tell me more?
I really wish most of the LLM folks just took a few courses in linguistics. It would avoid a lot of noise.
Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs
#230Earlier quoted context omitted.
The existence of an "inner monologue" isn't really a falsifiable claim. Some people claim to have one while other people claim not to, but we can't test the truth of these claims.
In this particular case, is there any reason why we simply can't take their word for it? This is not a case of where if I say "weak" or "strong", most people pick strong because no one wants to be weak, even if the context is unknown (nuclear force for example).
My concern is that if we take their word for it, we're actually buying into two assumptions which (AFAIK) are both unproven:
1. That "Internal Monologues" (not consciously forced by attention) exist in the first place, as opposed to being false-memories generated after-the-fact by our brain to explain/document a non-language process that just occurred. (Similar to how our conscious brains pretend that we were in control of certain fast reflexes.)
2. Some people truly don't have them, as opposed to just not being aware of them.