Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

1–10 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#3
Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper)

I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer:

https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64caf06c34...

Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong...

https://chatgpt.com/share/7e6a7201-dd2b-43c6-8427-76e5b003ca...

Also worth noting that there are more complex examples where GPT4o seems to fall down such as:

> Alice has 3 sisters. Her mother has 1 sister who does not have children - she has 7 nephews and nieces and also 2 brothers. Alice's father has a brother who has 5 nephews and nieces in total, and who has also 1 son. How many cousins does Alice's sister have?

However I can't honestly say that this is THAT simple or that most people would get this right...

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#6
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

I got it in 5 seconds, am I the singularity ?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#7
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

Seems like an odd thing not to state prominently. ChatGPT 4o solved it every time for me.

Seems to work OK even with the "hands tied behind back" prompt. Wonder if this paper is just inaccurate or if OpenAI adjusted the model -- seems unlikely.

Actually, 4o is reported to solve this riddle 60% of the time. I guess I have to test more, but with my about a dozen tests so far, all were correct.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#8
I'm using chatGPT 4o and it gave me the following :)

ME Alice has 40 brothers and she also has 50 sisters. How many sisters does Alice's brother have?

ChatGPT To determine how many sisters Alice's brother has, we need to use the information provided:

Alice has 40 brothers and 50 sisters. Since Alice's brother is one of the 40 brothers, he shares the same family structure as Alice. This means he has the same number of sisters as Alice.

Therefore, Alice's brother has 50 sisters.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#9
> This breakdown can be considered to be dramatic ... because models tend to express strong overconfidence in reporting their wrong solutions as correct, while often providing confabulations to additionally explain the provided final answer, mimicking reasoning-like tone but containing nonsensical arguments as backup for the equally nonsensical, wrong final answers.

People do that too!

Magical thinking is one example. More tangible examples are found in politics, especially in people who believe magical thinking or politicians' lies.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#10
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

There must be a name for the new phenomenon, of which your post is an example, of: 1. Someone expresses that an LLM cannot do some trivial task. 2. Another person declares that they cannot do the task, thereby defending the legitimacy of the LLM.

As a side note, I cannot believe that the average person who can navigate to a chatgpt prompter would fail to correctly answer this question given sufficient motivation to do so.

Post reply on HN