Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

91–100 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#91
post #23
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

I think you're wrong about that — I just tried prompting ChatGPT 4o to show all its working before giving an answer.

It was still incorrect, but when asked to show its working it formatted the answer prettily.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#92
post #29

"Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?" Gemini nailed it with no extra prompting. https://g.co/gemini/share/6946d20c0a4d When I gave it numbers, Gemini got confused. Alice has 4 brothers and she also has 3 sisters. How many sisters does Alice’s brother have? https://g.co/gemini/share/14d34dcdc5df But then when I suggested that it might be a trick question, it nai…

GPT-40 got it right with the abstract puzzle. Gemini got it wrong when I tried it.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#93
post #23
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

>_because that's the only way it can compute anything_

I'm fairly certain we'll soon realize that what's happening here is that the markov chain being run over latent space needs a certain amount of "warmup" before it starts sampling from the optimal region. HMC samplers for Bayesian methods have this same property.

The terms "reasoning", "computing" or "thinking" for this stage should be considered metaphors rather than explanations for what's happening, which is really waiting for a random walk to start sampling from the typical-set.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#94
post #18

For anyone considering reading the paper and like me don't normally read papers like this, open the PDF and think you don't have time to read it due to its length. The main part of the paper is the first 10 pages and a fairly quick read. On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong co…

For reasoning of any complexity, some sort of internal monologue and iteration is needed. For type-2 thinking, we conceptualize possible solutions, arguments, and reasoning paths in our mind, and judge and evaluate them before deciding which one we think is adequate or correct. This can be emulated to some extent by guiding prompts that make that internal thought process external, but we really want it to be internal and automatic.

We don’t know how to do that yet, because what controls the internal thought process is itself not necessarily language-based, and also, since internal thought processes of biological brains are not directly observable, they can’t be used as training data.

Edit: It occurs to me now that there is some parallel between current LLMs and behaviorism [0], and we really need something to which cognitive psychology could be applied instead.

[0] https://en.wikipedia.org/wiki/Behaviorism

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#96
post #29

"Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?" Gemini nailed it with no extra prompting. https://g.co/gemini/share/6946d20c0a4d When I gave it numbers, Gemini got confused. Alice has 4 brothers and she also has 3 sisters. How many sisters does Alice’s brother have? https://g.co/gemini/share/14d34dcdc5df But then when I suggested that it might be a trick question, it nai…

Oh I forgot about Gemini. I just tried my boat test ( https://news.ycombinator.com/item?id=40455392 ) that all LLMs I've tried have failed. Gemini failed as well, even with your addition about thinking carefully and telling it that it may be a trick. (Sorry I couldn't figure out how to link it like you did) Prompt: "You are on the shore of a river with a goat and a boat. How do you get to the other side? Think carefu…

My experience with GPT-4o with the same prompt was pretty solid. With no special system prompt and your same user prompt, I got:

This scenario is intriguing and does seem to carry the potential for a tricky question. In the simplest terms, the presence of a goat doesn't complicate the actual act of crossing the river. If there are no additional constraints or rules involved (like the need to also transport something else that requires careful consideration), you can simply:

1. Place the goat in the boat. 2. Row across the river.

However, if we are to assume there's more depth or a twist to the question, we might consider an outside-the-box answer. Here’s an alternative possibility if there is indeed some trick:

1. Get into the boat. 2. Leave the goat on the shore (if it’s safe) and row across the river.

There might be other interpretations depending on additional constraints that could be implicitly part of the question. If the goal was simply to get yourself across, leaving behind the goat and traversing the river is the straightforward solution.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#97
post #23

Earlier quoted context omitted.

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

Even if this were all true, it points to a fundamental risk of using LLM's for important tasks, which is that it is not at all clear to a user that this prompt would cause a problem. The LLM doesn't say "I'm sorry Dave, I just can't do that", it just complies with it and gets the wrong answer. You can always make excuses for the LLM afterwards, but software with hidden risks like this would not be considered good or…

People really need to stop trying to model an LLM as some kind of magical software component: it all makes a lot more sense if you model it as an under-performing poorly-aligned employee; so like, maybe a distracted kid working for peanuts at your store. You wouldn't trust them to with all of your money and you wouldn't trust them to do a lot of math--if they had to be in charge of checkout, you'd make sure they are only given a point-of-sale terminal and their main job was to, at best, scan the barcodes and compare the total--and yet there are tasks you can imagine handing to them that you'd never give to a robot or computer even though they get it wrong a lot, as not all tasks need to be handled perfectly, they still understand extremely fuzzy tasks, and they are probably cheaper than a qualified adult (certainly cheaper than one who is being paid enough to "give a shit" and pay enough attention to not let you get robbed or even put themselves at some risk for you).

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#99
post #29

"Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?" Gemini nailed it with no extra prompting. https://g.co/gemini/share/6946d20c0a4d When I gave it numbers, Gemini got confused. Alice has 4 brothers and she also has 3 sisters. How many sisters does Alice’s brother have? https://g.co/gemini/share/14d34dcdc5df But then when I suggested that it might be a trick question, it nai…

> that cannot have hundreds of siblings

See this is the problem with claims that humans are a 'general intelligence'. They get confused when encountering out-of-distribution situations. A true general intelligence would simply apply the knowledge that surrogate pregnancies cost around ~$50,000 and recall from historical context their knowledge of IVF. The AGI would then assume that the situation is simply that a billionaire couple has decided to have hundreds of kids and get on with the calculation. The search for intelligent life continues.

content note: i'm sorry

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#100
post #23
post #3

Question is: "Alice has 60 brothers and she also has 212 sisters. How many sisters does Alice’s brother have?" (nb: I have added numbers, it's phrased as X and N in the paper) I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer: https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64ca…

>Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong... Yes this is a common thing I see people who think LLMs are idiots do. The more an LLM talks the smarter it ge…

> The more an LLM talks the smarter it gets

I have a blog post in coming on this topic, but yes, this is right.

My method is to first get the LLM to answer the question, and THEN feed the answer back the LLM to extract the answer using constraints + grammar/logit bias/regex to parse the answer. Previously, I constrained to a single true/false token, which worked, but fails on complex queries.

So I split the decision making into a "justification" portion[0], and a "parsing" portion. I found that even crafting the prompt matters here, if you start with or end with, "It's very important to the response includes 'The answer is:'", then the model will lead with that response or only reply with that response. So I put it in the middle of the prompt, and end with with a request to justify the response. As a result, most models will reason their way to the answer, and then end with 'The answer is:'.

https://github.com/ShelbyJenkins/llm_client/blob/e3c4a860dda...

Post reply on HN