Live data from Hacker News

LLMs don't do formal reasoning

garymarcus.substack.com

101–110 of 128 posts

Re: LLMs don't do formal reasoning

#101
post #11

Earlier quoted context omitted.

> Are you suggesting Humans can't do formal reasoning? most of them can't. they actively vote against their own interests and everybody else around them. if you confront them with facts and figures, they ignore it and resort to emotional appeals. just because you can find one child that can play chess at age 4 doesn't mean that the rest won't just eat the pieces and shit on the board.

I've never been a fan of the argument of how voting against your own interest is some gap in logic. To me it can mean multiple things, one could be not understanding the implications of what that vote could mean, the other is a full understanding but wanting to put "the greater good" above the self. I would not want to live in a society where everyone only votes for their own interests.

it's not the voting that's the question here, it's that it's based on fallacious arguments, to the point that they will still go for it even when it's against their own interests and their neighbours.

Re: LLMs don't do formal reasoning

#102
Neither does Gary Marcus. I'd watch him try to determine truthfulness of the following expression: !!!!...!!true where the number of exclamation marks is chosen at random between 100500 and 500100 without using any external tools.

This would have been an argument against LLMs reasoning if you concede from the above that humans also don't do formal reasoning.

Re: LLMs don't do formal reasoning

#103
post #58

Earlier quoted context omitted.

> The problem is that general problem solving requires potentially arbitrary amounts of moving data Can you expand on this thought?

Forget about solving practical problems for a second. We can just ask the LLM to simulate some arbitrary computation within its context window. But we can in principle require that the output depends on state from arbitrarily many steps in the past. You then need to "carry forward" the required data or otherwise make it available. This is what I mean by moving data. The required associations between data can extend b…

And by extension, the assumption is that animals can carry forward an unlimited amount of information from the past? I.e., humans rely on culture to "carry forward" ideas from the past that are outside their individual contextual window of experience?

Re: LLMs don't do formal reasoning

#104
post #66
post #39

Earlier quoted context omitted.

That's an interesting rebuttal if you can suggest near-future architectures which don't require their own nuclear power plants to reliably calculate 13 x 54.

They're already operating on an architecture that can do that for about a nanojoule. You can also just ask them to write code for you, which appears to be what ChatGPT does now — it has its own python environment, I'm not sure what's in it except matplotlib and pandas, but it's at least that.

I don't know if it's unique to my use case (research), but I haven't had much luck getting ChatGPT to develop useable code. At best, it seems like it's useful for identifying packages to research to solve the problem. Maybe my prompts just need improving.

Re: LLMs don't do formal reasoning

#105
post #2

If anyone is curious, a Meta Data Scientist published a great piece about how the facts about what LLMs are actually doing (and therefore able to do) and how it's papered over by using chat bots. It's a long but very engaging read. https://medium.com/@colin.fraser/who-are-we-talking-to-when-...

this is a good article but very outdated - none of the examples he cites are relevant anymore

Re: LLMs don't do formal reasoning

#106

The paper (published 4 days ago) has this on page 10, and says that o1-mini failed to solve it correctly: Oliver picks 44 kiwis on Friday. Then he picks 58 kiwis on Saturday. On Sunday, he picks double the number of kiwis he did on Friday, but five of them were a bit smaller than average. How many kiwis does Oliver have? I pasted it into ChatGPT and Claude, and all four models I tried gave the correct answer: 4o mini…

Remember. How. It works. Please, please remember how it works. It is generating an answer anew, every single time. It is amazing how often it produces a correct answer, but not at all surprising that it produces inconsistent and sometimes incorrect answers.

Re: LLMs don't do formal reasoning

#108
post #94
post #48

Earlier quoted context omitted.

I struggle to see the use of this comment. Many human beings have jobs where they reason about problems far more complex than this every day. Sure, not every human is great at this. But the interest in using LLMs as agents does kind of require that they can routinely get this right -- the author of the blog post mentions this explicitly.

> Many human beings have jobs where they reason about problems far more complex than this every day. And they hold degrees from decades of education that taught them how to do that. Kids, even smart ones, can't do this reliably. I have two. I'm just saying that 3 years into the AI Revolution is a bit premature to demand that they "routinely get this right" when you yourself took probably 20 years to get to that point…

Quoting from the blog post:

"The inability of standard neural network architectures to reliably extrapolate — and reason formally — has been the central theme of my own work back to 1998 and 2001, and has been a theme in all of my challenges to deep learning, going back to 2012, and LLMs in 2019."

I think he makes a pretty lucid point that people have been questioning this for a long time, and definitely longer than 3 years. If you think there is some particular feature of LLMs that makes this a temporary hurdle, maybe you should make that point.

Re: LLMs don't do formal reasoning

#109
post #104
post #66

Earlier quoted context omitted.

They're already operating on an architecture that can do that for about a nanojoule. You can also just ask them to write code for you, which appears to be what ChatGPT does now — it has its own python environment, I'm not sure what's in it except matplotlib and pandas, but it's at least that.

I don't know if it's unique to my use case (research), but I haven't had much luck getting ChatGPT to develop useable code. At best, it seems like it's useful for identifying packages to research to solve the problem. Maybe my prompts just need improving.

My experience is that the quality varies wildly by task.

As an iOS dev, I certainly wouldn't call it "expert", but it's generally "good enough" to be a starting place whenever I get stuck, and on several occasions has surprised me with a complete bug free solution. Likewise when I ask it for web app stuff, though as that isn't my domain I wouldn't be able to tell you if the answers were "good" or "noob".

For the specific simple multiplication example given previously: https://chatgpt.com/share/6709a090-8934-8011-ae97-139b5758ad...

I do also have custom instructions set, but the critical thing here is the link to the python script, which is linked to at the end of the message, the blue text that reads: [>_]

Re: LLMs don't do formal reasoning

#110

I am not sure who the target audience of Gary Marcus is. Those who know about LLMs are aware that they do not reason, but also know it not very useful to repeat it over and over again and focus on other aspects of research. Those who don't know about LLMs simply learn to use them in a way that's useful in their life.

Maybe his target audience is anyone who might have read (or written) this??? https://openai.com/index/introducing-openai-o1-preview/

- "A new series of reasoning models for solving hard problems. Available now." - "They can reason through complex tasks and solve harder problems than previous models in science, coding, and math." - "In a qualifying exam for the International Mathematics Olympiad (IMO), GPT-4o correctly solved only 13% of problems, while the reasoning model scored 83%." - "But for complex reasoning tasks this is a significant advancement and represents a new level of AI capability." - "As part of developing these new models, we have come up with a new safety training approach that harnesses their reasoning capabilities to make them adhere to safety and alignment guidelines. By being able to reason about our safety rules in context, it can apply them more effectively. " - "These enhanced reasoning capabilities may be particularly useful if you’re tackling complex problems in science, coding, math, and similar fields."

there are a few more in that post, but clearly OpenAI is pushing the reasoning thing A LOT

Post reply on HN