Live data from Hacker News

Procedural knowledge in pretraining drives reasoning in large language models

arxiv.org

91–100 of 104 posts

Re: Procedural knowledge in pretraining drives reasoning in large language models

#91

Earlier quoted context omitted.

I don't think a human could effectively "reason" after being trained on nonsense (I don't think the training would even take). I think believing generative AI is operating on the same kind of meaning we are is a good way to be surprised when they go from writing like a learned professor for paragraphs to suddenly writing in the same tone and confidence but entirely wrong and with a bunch of made-up crap—it's all made…

But how can you be sure? You talk with confidence as if evidence exists to prove what you say but none of this evidence exists. It’s almost as if you’re an LLM yourself making up a claim with zero evidence. Sure you have examples that correlate with your point but nothing that proves your point. Additionally there exists LLM output that runs counter to your point. Explain LLM output that is correct and novel. There e…

I ignored your other response to me because I didn't see anything in the abstract that contradicted my posts, but maybe there's something deeper in the paper that does. I'll read more of it later.

I think, though, the disconnect between us is that I don't see this:

> Explain LLM output that is correct and novel.

As something I need to do for my position to be strong. It would be if I'd made different claims, but I haven't made those claims. I can see parts of the paper's abstract that would also be relevant and tough to deal with if I'd made those other claims, so I'm guessing those are the parts you think I need to focus on, but I'm not disputing stuff like (paraphrasing) "LLMs may produce output that follows the pattern of a form of reasoned argument they were trained on, not just the particulars" from the abstract. Sure, maybe they do, OK.

I don't subscribe to (and don't really understand the motivation for) claims that generative AI can't produce output that's not explicitly in their training set, which is a claim I have seen and (I think?) one you're taking me as promoting. Friggin' Markov chain text generators can, why couldn't LLMs? Another formulation is that everything they output is a result of what they were trained on, which is stronger but only because it's, like, tautologically true and not very interesting.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#92

Earlier quoted context omitted.

> or, at least, not “reasoning” about the things a human would be to solve the problems You can't actually infer that either. Humans have considerable context that LLMs lack. You have no basis to infer how a human would reason given the same context as an LLM, or vice versa.

I don't think a human could effectively "reason" after being trained on nonsense (I don't think the training would even take). I think believing generative AI is operating on the same kind of meaning we are is a good way to be surprised when they go from writing like a learned professor for paragraphs to suddenly writing in the same tone and confidence but entirely wrong and with a bunch of made-up crap—it's all made…

> I don't think a human could effectively "reason" after being trained on nonsense (I don't think the training would even take)

Go talk to a flat Earther or other religious zealot.

> I think the form of the output and the training sets we've chosen are what're making people believe they're doing stuff significantly like thinking, but it's all the same to an LLM.

Yes, but your mistake is asserting that thinking is not just following a set of patterns that lead from premises to conclusions, but that's literally what deductive logic is.

Let me put it this way, your argument is basically saying that computers can't reproduce human thinking because they can be programmed to output all kinds of nonsense that clearly isn't thinking (or at least, this is what it seems like you're saying). Well sure, but if they're programmed to actually think like humans, which should theoretically be possible given the Bekenstein Bound and the Church-Turing thesis, then clearly they are reproducing human thinking despite the fact that they can also be programmed to produce nonsense.

So the core question is whether artficacts of human thinking, like textbooks, poetry, etc. are sufficient for a learning system that can theoretically learn to reproduce arbitrary functions, to learn to reproduce human thinking. So rather than the training sets being some kind of red herring that are confusing people into concluding that LLMs are thinking, they are in fact central to the whole argument for why LLMs might actually be thinking!

We probably agree that LLMs don't have the same understanding of meaning that humans do, but I'm closer to the position that this is because they haven't been exposed to the same datasets we have, and not necessarily because their fundamental operation is so different. I think fundamental operation is probably a red herring because of Turing equivalence.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#93

Going meta a bit: comments so far on this post show diametrically opposing understandings of the paper, which demonstrates just how varied the interpretation of complex text can be. We hold AI to a pretty high standard of correctness, as we should, but humans are not that reliable on matters of fact, let alone on rigor of reasoning.

I find it bizarre that people often criticize LLMs’ capabilities by stating they don’t exhibit human-level performance because they sometimes fail, when occasionally failing is, in fact, a quintessentially human trait. Moreover, how do we reconcile the fact that LLMs perform far better than most humans on all the standardized tests they’ve passed? https://medium.com/@fsndzomga/there-will-be-no-agi-d9be9af44...

Re: Procedural knowledge in pretraining drives reasoning in large language models

#94

Going meta a bit: comments so far on this post show diametrically opposing understandings of the paper, which demonstrates just how varied the interpretation of complex text can be. We hold AI to a pretty high standard of correctness, as we should, but humans are not that reliable on matters of fact, let alone on rigor of reasoning.

[deleted]

Re: Procedural knowledge in pretraining drives reasoning in large language models

#95

Earlier quoted context omitted.

I don't think a human could effectively "reason" after being trained on nonsense (I don't think the training would even take). I think believing generative AI is operating on the same kind of meaning we are is a good way to be surprised when they go from writing like a learned professor for paragraphs to suddenly writing in the same tone and confidence but entirely wrong and with a bunch of made-up crap—it's all made…

> I don't think a human could effectively "reason" after being trained on nonsense (I don't think the training would even take) Go talk to a flat Earther or other religious zealot. > I think the form of the output and the training sets we've chosen are what're making people believe they're doing stuff significantly like thinking, but it's all the same to an LLM. Yes, but your mistake is asserting that thinking is not…

> We probably agree that LLMs don't have the same understanding of meaning that humans do

I think this is absolutely key, because to the extent LLM output has meaning we care about, I think it's effectively all supplied by us after the fact. This doesn't mean LLMs can't do interesting and useful things, but I consider some current maximalist takes on the state and immediate likely future of AI as doing something a couple steps up the complexity-chain from believing a pocket calculator that solves "1 + 5" must understand what "6" means to a human. That "6" is just some glowing dots until we assign it meaning and context that the pocket calculator doesn't and cannot have, even though it's really good at solving and displaying the results of calculations.

This model explains, I think, the actual experience of using an LLM better than the model I gather some have, which is that they're doing something pretty close to thinking but just get stuff wrong sometimes, as a human gets stuff wrong sometimes (I think they get things wrong differently from how a human does, and it's because they aren't working with the same kind of meaning we are). I think it's the familiar form of the output that's leading people down this way of thinking about what they're doing, and I think it's wrong in ways that matter, both for policy purposes (pleading that they're just doing what humans do to learn and then produce output from what they learned, when it comes to building them with copyright-protected data, falls rather flat with me, for example—I'm not quite to the point of entirely dismissing the argument, but I'm pretty close to sticking that in the "not even wrong" bucket, in part because of my perspective on how they work) and for actually working productively with generative AI. When these programs fail, it's usually not the way a human does, and using heuristics for recognizing places you need to be cautious when dealing with human output will result in mistakes. "Lies" often looks a lot like "truth" with them and can come out of nowhere, because that's not quite what they deal in, not the way humans do. They don't really lie but, crucially, they also don't tell the truth. But they may produce output that contains information that's wrong or correct, and takes a form that's very useful, or not very useful.

> but I'm closer to the position that this is because they haven't been exposed to the same datasets we have, and not necessarily because their fundamental operation is so different.

I'm not super far from agreeing with this, I think, but also think there's probably some approach (or, I'd expect, set of approaches) we need to layer on top of generative AI to make it do something that I'd consider notably close to human-type thinking, in addition to just being able to poke it and make text come out. I think what we've got now are, in human terms, something like severely afflicted schizophrenics with eidetic memories, high levels of suggestibility, and an absence of ego or self-motivation, which turns out to be pretty damn useful things to have but isn't necessarily something we'll get broadly human-level cognition (or better—I mean, they're already better than a lot of people at some tasks, let's face it, zero people who've ever lived could write bad satirical poetry as fast as an LLM can, much as nobody can solve square roots as fast as a pocket calculator) out of if we just do more of it—I doubt that the basic components needed to bridge that gap are present in the current systems at all. I expect we'll see them fail less as we feed them more energy and data, but for their failures to continue to look alien and surprising, always due to that mismatch between the meaning we're assigning to what they're doing, and their internal sense of "meaning", which are correlated (because we've forced them to be) but not dependent on one another in some necessary way. But yes, giving them more sources of "sensory" input and a kind of will, with associated tools, to seek out more input, is likely the direction we'll need to go to make them capable of more things, rather than just somewhat better at what they do now.

[EDIT] As for why I think our ways of discussing how these work matters, aside from the aforementioned reasons, it's that lay-people are taking our lead on this to some extent, and when we come out acting like these are thinking agents in some serious sense (or cynically promoting them as super dangerous and close to becoming real conscious entities on the verge of being insanely "smart" because, gee would you look at that, we're selling the things—ahem, paging Altman) it's a recipe for cooking up harmful misuse of these tools and of laws and policy that may be at-odds with reality.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#97
post #65

Earlier quoted context omitted.

I have wondered that from time to time, why not train an AI system using educational curricula plus some games and play? It might be fascinating to see what comes out using various systems from around the world.

They do train on textbooks.

Textbooks are all you need - https://arxiv.org/pdf/2306.11644

Re: Procedural knowledge in pretraining drives reasoning in large language models

#98
post #15
post #9

Earlier quoted context omitted.

You're close, but there’s an important nuance. The process isn't about "learning how to solve problems in general" in the broad sense. It's more specific: the neural network is trained to mimic the step-by-step process demonstrated by humans solving a specific problem. The distinction is that the software doesn't autonomously derive general problem-solving heuristics from scratch. Instead, it observes examples of how…

Yes, except that I'm not so sure there is a clear distinction between following general instructions and generating new heuristics. It's just a difference in the level of abstraction there, and probably not even that one in any discrete sense, more like a continuum. (Current) models may of course lack sufficient training data to act on a metalevel enough ("be creative problem solvers"), or they may lack deep enough r…

Up to a point, general instructions can be generated from a collection of specific examples by abstracting what is different between them, but it is not clear to me that abstraction is all you need to come up with novel methods.

This seems consistent with the main points in this paper: one to-the-point statement of the answer to a factual question is all you need [1], while, if you don't have an example of a chain of reasoning in which all of the parameters are the same as those in the prompt, more than one example will be needed.

The authors write "We falsify the hypothesis that the correlations are caused by the fact that the reasoning questions are superficially similar to each other, by using a set of control queries that are also superficially similar but do not require any reasoning and repeating the entire experiment. For the control queries we mostly do not observe a correlation." In the examples of control queries that they give, however, this just amounts to embedding the specific answer to the question asked into language that resembles an example of reasoning to a solution (and in the first example, there is very little of the latter.) The result, in such cases, is that there is much less correlation with genuine examples of reasoning to a solution, but it is not yet clear to me how this fact justifies the claim quoted at the start of this paragraph: if the training set contains the answer stated as a fact, is it surprising that the LLM treats it as such?

[1] One caveat: if the answer to a factual question is widely disputed within the training data, there will likely be many to-the-point statements presented as the one correct answer (or - much less likely, I think - a general agreement that no definitive answer can be given.) The examples given in figure 1 are not like this, however, and it would be interesting to know if the significance of individual documents extends to such cases.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#99
post #7

You mean you need humans to step-by-step solve a problem so a neural net can mimic it? It sounds kinda obvious now that I write it out.

No. If I'm understanding correctly it means the software is learning how to solve problems in general by ingesting examples of procedural problem-solving.

I don’t think the paper says anything about general problem-solving? They analyzed 40 reasoning problems and I didn’t find a list of them, but the example they use most often is “find the slope of a line.” Apparently there are many examples in the dataset demonstrating how to find the slope of a line, including computer code, and they all influence the answer.

There are many ways you could ask a question about how to find the slope of a line, and the “generalization” going on seems to be unifying the different ways you could ask and answer this question.

It seems fair to say that the LLM did learn to find the slope of a line? But the question has definitely been solved before, many times.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#100

Earlier quoted context omitted.

But how can you be sure? You talk with confidence as if evidence exists to prove what you say but none of this evidence exists. It’s almost as if you’re an LLM yourself making up a claim with zero evidence. Sure you have examples that correlate with your point but nothing that proves your point. Additionally there exists LLM output that runs counter to your point. Explain LLM output that is correct and novel. There e…

I ignored your other response to me because I didn't see anything in the abstract that contradicted my posts, but maybe there's something deeper in the paper that does. I'll read more of it later. I think, though, the disconnect between us is that I don't see this: > Explain LLM output that is correct and novel. As something I need to do for my position to be strong. It would be if I'd made different claims, but I ha…

It’s ok if you ignore it. I don’t mind.

Your claim is LLMs can’t reason.

And you say you don’t have to explain why LLMs output novel and correct answers. Well I asked for this because the fact that LLMs output correct and novel answers disproves your point. I disproved your point. Think about it. You think I introduced some orthogonal topic but I didn’t. I stated a condition that the LLM meets that disproves your claim.

So if there exists a prompt and answer pair such that the answer is so novel the probability of it being a correlation or random chance is extraordinarily low then your point is trivially wrong right?

Because if the answer wasn’t arrived by some correlative coincidence which you seem to be claiming then the only other possible way is reasoning.

Again such question and answer pairs actually exist for LLMs. They can be trivially generated like my shared link above which it talks about the entire topic of this sub thread without training data to support it.

Humans fail at reasoning all the time yet we don’t say humans can’t reason. So to keep consistency for the criterion we use for humans. If the LLM reasoned it means it can reason even if it clearly gives the wrong answer sometimes.

Additionally you likely claim all humans can reason. What’s your criterion for that? When you look at a human sometimes it outputs correct and novel answers that are not part of its training data (experiences).

It’s literally the same logic but you subconsciously move the goal posts to be much much higher for an LLM. In fact under this higher criterion all mentally retarded humans and babies can’t reason at all.

Post reply on HN