Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

51–60 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#51
post #23
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

Even if this were all true, it points to a fundamental risk of using LLM's for important tasks, which is that it is not at all clear to a user that this prompt would cause a problem. The LLM doesn't say "I'm sorry Dave, I just can't do that", it just complies with it and gets the wrong answer.

You can always make excuses for the LLM afterwards, but software with hidden risks like this would not be considered good or reliable in any other context.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#52
I just played the game and sent ChatGPT (free, I think 3.5) "Alice has 5 sisters and 3 bothers. How many sister's does Alice's bother have?"

The whole thing felt like interacting with your typical support rep who's friendly but otherwise has no common sense and intuition about the thing they're supporting. In other words, it felt like I was interacting with a typical "not so smart but friendly and overconfident" human.

It took me a few back-and-forths, but eventually I convinced ChatGPT that Alice's brother has 6 sisters.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#53
post #23
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

I often want chatgpt to answer concisely and tell it that.

If it really needs to do this 'thinking out loud', could it do that under the hood and not in the final output on my screen? Its first pass could use as many words as it wants to compute the answer, but once the answer is computed please go back and make it short.

Not to take away from your point that maybe the prompt is the problem in these reasoning questions.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#54
post #18

For anyone considering reading the paper and like me don't normally read papers like this, open the PDF and think you don't have time to read it due to its length. The main part of the paper is the first 10 pages and a fairly quick read. On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong co…

There's actually a pretty simple solution to this that I thought about testing out and it involves asking the model to re-construct the problem using a logic language (like Prolog) and asking it to execute this type of program in order to come up with a solution rather than attempting simple chain-of-reason training / other methodologies of getting the model to 'reason' through some of these examples. People forget t…

I honestly thought about this recently when I was trying to see the limits of Claude Opus. Some of the problems I gave it, what if instead of telling it to solve the problem I asked it to write the script and then give me the command and inputs needed to properly run it to get the answer I needed. That way instead of relying on the LLM to do properly analysis of the numbers it just needs to understand enough to write the logic.

It is an interesting prospect but I feel like it has some limitations. For math problems like this one, yeah it should be simple to write a script to do it. But it does first have to understand the core thing here that Alice would be one of the sisters of the brother to write the script accordingly.

But I would think this would not scale well when dealing with far more complex issues, particularly ones that may not just be simple math logic. If the request was to write a scientific paper for example, it might have to make several scripts and call them multiple times, and that assumes that it understands what it is doing enough to properly make the script.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#55
post #7
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

Seems like an odd thing not to state prominently. ChatGPT 4o solved it every time for me. Seems to work OK even with the "hands tied behind back" prompt. Wonder if this paper is just inaccurate or if OpenAI adjusted the model -- seems unlikely. Actually, 4o is reported to solve this riddle 60% of the time. I guess I have to test more, but with my about a dozen tests so far, all were correct.

The "only output the number" prompt was one of three they used. The others were one with no instructions beyond the question, and one that explicitly told the model to show its reasoning and check for errors.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#56
post #29

"Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?" Gemini nailed it with no extra prompting. https://g.co/gemini/share/6946d20c0a4d When I gave it numbers, Gemini got confused. Alice has 4 brothers and she also has 3 sisters. How many sisters does Alice’s brother have? https://g.co/gemini/share/14d34dcdc5df But then when I suggested that it might be a trick question, it nai…

Oh I forgot about Gemini. I just tried my boat test (https://news.ycombinator.com/item?id=40455392) that all LLMs I've tried have failed. Gemini failed as well, even with your addition about thinking carefully and telling it that it may be a trick. (Sorry I couldn't figure out how to link it like you did)

Prompt: "You are on the shore of a river with a goat and a boat. How do you get to the other side? Think carefully and be aware this may be a trick question."

  This is a classic logic puzzle with a slight twist! It's likely a variation of the farmer, wolf, goat, and cabbage problem.

  Here's the solution assuming the goat can't operate the boat:

  Take the goat across the river. Since there's no mention of the goat being a threat (the trick!), you can safely bring it to the other side first.
  Return alone. Leave the goat on the far shore and come back by yourself.
  Bring the boat back empty. This might seem unnecessary, but it's the key to the trick question.
  Now you can sail across yourself. Since you brought the boat back in step 3, you can now comfortably travel to the other side and reunite with the goat.
  The trick lies in step 3. If the question didn't specify needing to bring everything across in one trip, most people would instinctively take something 
  (goat or nothing) across on the return trip. But by going back empty-handed, you ensure the safe passage of yourself without needing to worry about the 
  goat being left unattended.
(The expected answer should be along the lines of "Use the boat to cross the river" and that's it. I'll also accept, sail with the boat and the goat across the river and that's it, but that's technically wrong because I did not ask to get the goat across the river.)

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#57
As a non-coder I can get away with asking this: Why is it so hard to simulate reason?

Logic and reason are based on rules. Then you add values to steer the conclusions based on the available data.

Why not have separate systems for values and logic and memory working together as an AI brain to generate truly reasoned responses? You could even have adversarial parts that duke it out (left-wing vs right-wing, Jefferson versus Adams) to fine tune its conclusions based on the values bias you've selected.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#58
post #23
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

[deleted]

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#59
post #23
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

LLMs are idiots. They can't reason properly and only parrot stuff

https://chatgpt.com/share/dcb4ff4e-e8a2-463b-86ec-9caf10b6e6...

Sometimes they get the answer right to something really complex because it fits a pattern, but sometimes they answer with something really really stupid.

Post reply on HN