Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

291–300 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#291
post #242

Earlier quoted context omitted.

Not only are they unproven, but are ultimately not provable at all. Some people will say yes, some people will say no. Probably we can take their word for it, but in the simplest case they could just lie (in either direction) and we would have no way to tell. In short, maybe these inner monologues exist and maybe they don't, but science can't comment on that. That said, it is clearly something we are interested in, b…

> but are ultimately not provable at all No, they are potentially falsifiable as we get better at scanning, identifying, intervening in brain activity. Just off the top of my head here, suppose we create a table puzzle problem that (in itself) doesn't require language to understand, like ones we make for certain animals. Have human subjects (silently) solve it. Afterwards, quiz the solvers about their internal monolo…

Sure, there is stuff we can do to tease around the edges (similar problems crop up all the time in psychology and sociology) but we will always have to weaken the claim in order to do experiments relating to it.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#292

Earlier quoted context omitted.

> It's impossible to plan without an internal monologue. Of course it isn't impossible, and this is backed by what we know about paleoanthropology and other instances of cognition in animals - humans were making stone tools millions of years ago, which takes planning in the form of imagining what you want the tool to look like and how you will do it and what it will be used for. It's exceedingly likely we had this ab…

Animals can not manipulate abstract concepts nor can they do long-term plans. No crow can plan an international trip spanning a couple of weeks and two change-overs. And some people definitely can't do it start to end, but they can at least plan the first 5-7 steps. Also, maybe inner monologue is not a binary have/have not, but maybe it is on a continuum.

Not sure. Migratory birds seem to manage this just fine. Not only do they make multiple stops to eat and rest, they also navigate around bad weather and still make it to their intended destination (at least most of the time).

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#293

Earlier quoted context omitted.

I think you’re using “inner monologue” too literally. It could be a progression of pictures, emotions, etc.

To make any progress on this question at all, we need first to come up with some definition of internal monologue. Even if we may need to modify it later, there has to be a starting point. Otherwise, nothing can be established at all, because for any statement there always will be someone's understanding of "internal monologue" for which the statement is true, and someone's else understanding for which the statement…

I'm sure inner monologue just cashes out into the ability to reflect on your own thoughts. And for one to say that they're not having that experience also involves a claim about what they think other people are having which would make me doubly skeptical.

In practice, when you see people arguing about whether they have an "inner monologue" or can "mentally picture objects" on social media, it's more of a contest of who is the most unique in the world rather than anything that sheds clarity on our subjective experience.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#294
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>> I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer:

Remember that the authors of the paper did not find that GPT4-o cannot return the right answer. They found that it can't return the right answer more often than ~60% of the time. So you'd have to repeat the experiment many, many times and aggregate the results (the paper uses a binomial Beta this and that etc etc) before you see similar results as the paper.

You won't replicate the results of the paper unless you really put your back into it.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#295
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

The right answer depends on how Alice identifies I guess? :)

Page 42 of the paper :)

One thing that strikes me is that the model first tries using "inclusive language" in one answer - and literally states so, using this specific term - but seems to interpret it in a more mathematical sense (like set inclusion). Then seamlessly switches to the expected DEI spiel in the next paragraph.

For one thing, it makes me suspect that something with the words "inclusive language" was automatically added to the prompt. But more interesting is how it responds to this demand in two different ways, illustrating a "thought process" that is very much unlike that of a human with normal verbal reasoning ability.

I am not a psychologist, but remember reading that schizophrenic people sometimes confuse different meanings of words in a similar way, jumping from one meaning to another without noticing.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#296

Earlier quoted context omitted.

People keep saying it because that's literally how LLMs work. They run Montecarlo sampling over a very impressive latent linguistic space. These models are not fundamentally different than the Markov chains of yore except that these latent representations are incredibly powerful. We haven't even started to approach the largest problem which is moving beyond what is essentially a greedy token level search of this ling…

The best compression is some form of understanding

That's a fascinating insight and it sound so true!

Can you compress for me Van Gogh's Starry Night, please? I'd like to send a copy to my dear old mother who has never seen it. Please make sure when she decompresses the picture she misses none of the exquisite detail in that famous painting.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#297
post #87

Earlier quoted context omitted.

I appear to be reasoning at times but I have mostly no idea what I am talking about. I hit a bunch of words and concepts in the given context and thus kind of hallucinate sense. Given a few months of peace of mind and enough money for good enough food, I could actually learn to reason without sounding like a confused babelarian. Reasoning is mostly a human convention supported by human context that would have been a…

Yeah, I think these chatbots are just too sure of themselves. They only really do "system 1 thinking" and only do "system 2 thinking" if you prompt them to. If I ask gpt-4o the riddle in this paper and tell it to assume its reasoning contains possible logical inconsistencies and to come up with reasons why that might be then it does correctly identify the problems with its initial answer and arrives at the correct on…

There isn't any evidence that models are doing any kind of "system 2 thinking" here. The model's response is guided by both the prompt and its current output so when you tell it to reason step by step the final answer is guided by its current output text. The second best answer is just something it came up with because you asked, the model has no second best answer to give. The second best answers always seem strange because the model doesn't know what it means to come up with a second best answer; it 'believes' the output it gave is the correct answer and helpfully tries to fulfill your request. Sometimes the second best answer is right but most of the time its completely nonsensical and there is no way to distinguish between the two. If you ask to choose it will be strongly influenced by the framing of its prior response and won't be able to spot logical errors.

Asking it to do lateral thinking and provide examples isn't really helpful because its final output is mostly driven by the step by step reasoning text, not by examples it has generated. At best, the examples are all wrong but it ignores that and spits out the right answer. At worst, it can become confused and give the wrong answer.

I've seen gpt-4 make all kinds of errors with prompts like this. Sometimes, all the reasoning is wrong but the answer is right and vice versa.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#298

Earlier quoted context omitted.

Tell me more?

Language is only sequential in the form it is transmitted (verbally). There is no reason that sequential statements are generated sequentially in the brain. Quite the opposite, really, if you consider rules of grammar. I really wish most of the LLM folks just took a few courses in linguistics. It would avoid a lot of noise.

LLMs don't generate their language sequentially either, they just output it sequentially token by token.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#299
post #244

Earlier quoted context omitted.

Alice has N Brothers, and she has M sisters. How many sisters do Alice’s brothers have? I have not gotten the correct answer to the question as phrased above in one go from Gpt4o yet! (and today was not the first day i tried.) Phrase it as shown above and you'll likely need 5 or more interactions to get it to generate the correct output. With Gemini i could not get it below 8 without feeling like i was cheating. fwiw…

Chat GPT 4o. I was being a bit generous with background information, but still tests ability to interpret: ------ Me: Background facts: Alice is a female human. All sisters are female, and all brothers are male. No one is their own brother or sister. Alice has N brothers, and Alice has M sisters. Now, a few questions based on these facts: How many sisters do Alice’s brothers have? Do Alice's brothers have more sister…

>> Don't forget to consider Alice when counting.

a.k.a. "don't forget to give the LLM the answer when prompting".

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#300
post #16
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

Of course it's going to give an incorrect answer with that prompt. If the instruction fine tuning is neutered like this prompt, it's going to roll over to the foundation model and offer a completion - probably more influenced by the seed than the prompting text. Bad study. Edit - I just skimmed the paper - they do use other more appropriate prompt types for reasoning. My initial response was based on the assumption t…

>> My initial response was based on the assumption that all prompts used that script prompt quoted in the parent.

You, and another 20 or so commenters here. We should really re-examine the guideline about asking people to RTFA.

No offense meant- good on you for correcting your error.

Post reply on HN