Live data from Hacker News

LLMs don't do formal reasoning

garymarcus.substack.com

111–120 of 128 posts

Re: LLMs don't do formal reasoning

#111
post #7

Earlier quoted context omitted.

Are you suggesting Humans can't do formal reasoning? Because you can easily teach a four year old not to make illegal moves in chess with very little instructions, and by 10 geniuses like Terence Tao were discussing open math problems with Erdos. If anything this article adds further evidence that whatever the architecture of the human brain it is very different to an LLM architecture.

He said most humans. Comparing the LLM to Terence Tao is saying they're already better than almost every human.

The difference in Terence Tao's brain architecture to most humans is extremely minimal, even non-existent.

Re: LLMs don't do formal reasoning

#112
post #103

Earlier quoted context omitted.

Forget about solving practical problems for a second. We can just ask the LLM to simulate some arbitrary computation within its context window. But we can in principle require that the output depends on state from arbitrarily many steps in the past. You then need to "carry forward" the required data or otherwise make it available. This is what I mean by moving data. The required associations between data can extend b…

And by extension, the assumption is that animals can carry forward an unlimited amount of information from the past? I.e., humans rely on culture to "carry forward" ideas from the past that are outside their individual contextual window of experience?

I wasn't thinking in those terms, but yeah I like that. Humans are a kind of superorganism and a part of that is due to the power of culture to shape behavior in ways that are responsive to environmental changes deep in history beyond any individuals lifespan.

Re: LLMs don't do formal reasoning

#113
post #52

Earlier quoted context omitted.

We really really really need to disambiguate the LLM, which is a fixed length, fixed compute time process which takes in an input and produces a token distribution, from the AI system, which takes the output of the LLM and eventually produces something for the user. In this case, all LLMs are fixed-length, but not all AI systems are. An LLM on its own is useless. Current SoTA research includes inserting 'pause' token…

Yes. AIs come in all sorts of flavours. I think the main thing that happened with LLMs was that people anthropomorphise them because they finally understand what's going on. Other AIs might be smarter by solving complicated mathematical problems but most people don't speak that language so they're not impressed. LLM vendors should really make this clear but they don't because a magical thinking machine sells well.

> LLM vendors should really make this clear but they don't because a magical thinking machine sells well.

Hold on though... modern LLM systems, like ChatGPT 4o et al do stop and think. The vendors are not selling LLMs. LLMs are an implementation detail. They're selling AI systems: the LLM in addition to the controlling software.

Re: LLMs don't do formal reasoning

#114
post #7
post #4

You could substitute "LLMs" -> "Humans" and the statement would also be true.

Are you suggesting Humans can't do formal reasoning? Because you can easily teach a four year old not to make illegal moves in chess with very little instructions, and by 10 geniuses like Terence Tao were discussing open math problems with Erdos. If anything this article adds further evidence that whatever the architecture of the human brain it is very different to an LLM architecture.

> Because you can easily teach a four year old not to make illegal moves in chess with very little instructions ...

Have you seen actual four year olds? Most would not only make illegal moves, they will also throw a few pieces away, place their favorite giraffe next to the "horse," and laugh at your frustration. Thus proving, once and for all, that four year olds are not in fact intelligent. /s

Re: LLMs don't do formal reasoning

#115
post #108
post #94

Earlier quoted context omitted.

> Many human beings have jobs where they reason about problems far more complex than this every day. And they hold degrees from decades of education that taught them how to do that. Kids, even smart ones, can't do this reliably. I have two. I'm just saying that 3 years into the AI Revolution is a bit premature to demand that they "routinely get this right" when you yourself took probably 20 years to get to that point…

Quoting from the blog post: "The inability of standard neural network architectures to reliably extrapolate — and reason formally — has been the central theme of my own work back to 1998 and 2001, and has been a theme in all of my challenges to deep learning, going back to 2012, and LLMs in 2019." I think he makes a pretty lucid point that people have been questioning this for a long time, and definitely longer than…

I think I did make the point, but I'll do it again. Teaching a LLM to reason is 100% isomorphic to teaching a child to reason. All the logic being deployed here by the luddite set[1] could be deployed to explain why your grade schooler will never reason correctly. And it's wrong there, and there's no reason to expect that it's wrong here.

Very broadly: you learn to reason by learning to write and run "code" in your head. Can an LLM write and run code? Yes, it can. Do they use it currently to "reason" well? No, because no one has made that work yet. Does that constitute an argument that they CANNOT? Clearly not.

[1] And I'm no LLM booster! See the point about the pendulum upthread.

Re: LLMs don't do formal reasoning

#116
post #77

Earlier quoted context omitted.

> Yes, we should use LLMs to translate human requirements that are ambiguous and have a lot of hidden assumptions (e.g. that football matches should preferably be at times when people are not working & awake [3]), and use them to create formal requirements Why would an LLM trained on human language patterns be good at this? If anything, I would expect it to follow the same pattern that humans do.

It doesn't need to be good at solving the problem. It only needs to be good at translating the problem of "If the unknown x is divided by 3 it has the same value as if I subtracted 9 from it" into "x/3 == x-9 && x is an Real number". The formal method tool will do the rest. Note that if the LLM gets the implicit assumptions wrong, the solution will be unsatisfactory, and the query can be refined. This is exactly what…

The problem is not that writing input for formal method tools is tricky syntactically. The problem is that it is hard to produce the actual semantic content of the input. Humans don't need help with that, especially not the kind of human that is capable of authoring that content. It's much more technical than stuff like "Make me a website with a blue background" or some crap like that. The potential for an LLM to mistranslate the English input is probably unacceptable.

Re: LLMs don't do formal reasoning

#117

Earlier quoted context omitted.

> Large numbers of people and businesses are extracting huge value from the use of LLMs every single day No they aren't. If they really did we would see those number in qtrly reports.

The impact of LLMs for many sectors is deflationary, which means you specifically won't see those numbers in quarterly reports. Doesn't mean that value isn't being extracted with LLMs but rather that it's being eroded elsewhere. What you'll see is companies that don't leverage the benefits gradually being left behind.

Not sure what this comment means. There would be losers but no winners ? You mean there would be a loss of gdp from llms ?

Re: LLMs don't do formal reasoning

#118

Earlier quoted context omitted.

This article is long but doesn't mention key concepts like instruction tuning. I'd suggest the Llama paper as a more worthwhile source.

It does talk about openai explicitly instruction tuning the llm to try to constrain the output and the limitations of such approaches.

ctrl+F 'struction'

0 results

Re: LLMs don't do formal reasoning

#119

The paper (published 4 days ago) has this on page 10, and says that o1-mini failed to solve it correctly: Oliver picks 44 kiwis on Friday. Then he picks 58 kiwis on Saturday. On Sunday, he picks double the number of kiwis he did on Friday, but five of them were a bit smaller than average. How many kiwis does Oliver have? I pasted it into ChatGPT and Claude, and all four models I tried gave the correct answer: 4o mini…

Isn't it because this test has since been spread on the internet and the LLM's picked up on that so now they give the correct answer?

Maybe try a new unique logical question. And not the same question with a few words changed, because that might still match close to data the LLM already scanned.

Re: LLMs don't do formal reasoning

#120

Earlier quoted context omitted.

It does talk about openai explicitly instruction tuning the llm to try to constrain the output and the limitations of such approaches.

ctrl+F 'struction' 0 results

Thanks for demonstrating the depth of your reading.
Post reply on HN