Live data from Hacker News

Procedural knowledge in pretraining drives reasoning in large language models

arxiv.org

61–70 of 104 posts

Re: Procedural knowledge in pretraining drives reasoning in large language models

#61
post #54
post #44

Earlier quoted context omitted.

> Most humans are not that good at reasoning themselves and fall for the same kind of fallacies over and over because of the way they were brought up Disagree that it's easy to pin on "how they were brought up". It seems very likely that we may learn that the flaws are part of what makes our intelligence "work" and be adaptive to changing environments. It may be favourable in terms of cultural evolution for parents t…

> Disagree that it's easy to pin on "how they were brought up". Indeed. That might play a role, but another less politically charged aspect to look at is just: how much effort is the human currently putting in? Humans are often on autopilot, perhaps even most of the time. Autopilot means taking lazy intellectual shortcuts. And to echo your argument: in familiar environments those shortcuts are often a good idea! If y…

> And to echo your argument: in familiar environments those shortcuts are often a good idea!

Only to continue to reaffirm the original post, this was some of the basis for my dissertation. Lower-level practice, or exposure to tons of interactive worked examples, allowed students to train the "mental muscle memory" for coding syntax to learn the more general CS concept (like loops instead of for(int i = 0...). The shortcut in this case is learning what the syntax for a loop looks like so that it can BECOME a shortcut. Once its automatic, then it can be compartmentalized as "loop" instead of getting anxious over where the semicolons go.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#62
post #51

Earlier quoted context omitted.

I very much agree with the perspective that LLMs are not suited for “reasoning” in the sense of creative problem solving or application of logic. I think that the real potential in this domain is having them act as a sort of “compiler” layer that bridges the gap between natural language - which is imprecise - and formal languages (sql, prolog, python, lean, etc) that are more suited for solving these types of problem…

> I think that the real potential in this domain is having them act as a sort of “compiler” layer that bridges the gap between natural language - which is imprecise - and formal languages (sql, prolog, python, lean, etc) that are more suited for solving these types of problems. And then maybe synthesizing the results / outputs of the formal language layer. Basically “agents”. Well, if you do all that, would you say t…

The system as a whole has reasoned twice over, verbally and then logically.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#63
post #35
post #9

Earlier quoted context omitted.

You're close, but there’s an important nuance. The process isn't about "learning how to solve problems in general" in the broad sense. It's more specific: the neural network is trained to mimic the step-by-step process demonstrated by humans solving a specific problem. The distinction is that the software doesn't autonomously derive general problem-solving heuristics from scratch. Instead, it observes examples of how…

Crucially, this is what MacIntyre's narrativity thesis is talking about: If a university professor is giving a lecture on decentralized finance and forks into a recipe for chocolate chip cookies: crack two eggs, add a cup of flour, and fold in brown sugar prior to baking, it would break linearity. A generalizable strategy for synthesizing LLMs differentiated by their training parameters is a tokenization is isolating…

> A generalizable strategy for synthesizing LLMs differentiated by their training parameters is a tokenization is isolating data sets and then establishing a lattice in uniformity within the field of technics.

This comment appears to be incoherent and likely AI-generated text. Let me break down why:

1. While it uses technical-sounding terms related to machine learning (LLMs, tokenization, data sets), the way they're strung together doesn't make logical sense.

2. The grammar is incorrect: - "a tokenization is isolating" is not grammatically valid - The sentence structure breaks down in the middle with two "is" statements - The phrase "establishing a lattice in uniformity within the field of technics" is meaningless jargon

3. If we try to interpret what it might be attempting to say about LLMs (Large Language Models), the ideas don't connect in any meaningful way. "Synthesizing LLMs differentiated by their training parameters" could be trying to discuss creating different LLMs with varying parameters, but the rest doesn't follow logically.

4. The term "field of technics" is particularly suspicious - while "technics" is a real word, it's rarely used in AI/ML discussions and seems thrown in to sound technical.

This text shows common hallmarks of AI-generated content that's trying to sound technical but lacks real meaning - it uses domain-specific vocabulary but combines them in ways that don't make semantic sense, similar to how AI models can sometimes generate plausible-looking but ultimately meaningless technical text.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#64

Earlier quoted context omitted.

This is extremely common in these discussions. Most humans are not that good at reasoning themselves and fall for the same kind of fallacies over and over because of the way they were brought up (their training data so to speak). And yet they somehow think they can argue why or why not LLMs should be able to do the same. If anything, the current limits of these morels show the limits of human cognition which is sprea…

It's because we can put responsibility on humans to be correct but we can't on computers. Humans given the appropriate incentives are very good at their jobs and there is a path for compensation if they screw up. Computers have neither of these things.

I think this is a key argument in how powerful AI can become. We may be able to create incredibly intelligent systems, but at the end of the day you can’t send a computer to jail. That inherently limits the power that will be given over to AI. If an AI accidentally kills a person, the worst that could be done to it is that it is turned off, whereas the owners of the AI would be held liable.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#65

Does this mean LLMs might do better if trained on large amounts of student notes, exams, book reviews and such? That would be incredibly interesting.

I have wondered that from time to time, why not train an AI system using educational curricula plus some games and play? It might be fascinating to see what comes out using various systems from around the world.

They do train on textbooks.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#66
> In the extreme case, a language model answering reasoning questions may rely heavily on retrieval from parametric knowledge influenced by a limited set of documents within its pretraining data. In this scenario, specific documents containing the information to be retrieved (i.e. the reasoning traces) contribute significantly to the model’s output, while many other documents play a minimal role.

> Conversely, at the other end of the spectrum, the model may draw from a broad range of documents that are more abstractly related to the question, with each document influencing many different questions similarly, but contributing a relatively small amount to the final output. We propose generalisable reasoning should look like the latter strategy.

Isn't it much more impressive if a model can generalize from a single example?

Re: Procedural knowledge in pretraining drives reasoning in large language models

#67
post #51

Earlier quoted context omitted.

> I think that the real potential in this domain is having them act as a sort of “compiler” layer that bridges the gap between natural language - which is imprecise - and formal languages (sql, prolog, python, lean, etc) that are more suited for solving these types of problems. And then maybe synthesizing the results / outputs of the formal language layer. Basically “agents”. Well, if you do all that, would you say t…

The system as a whole has reasoned twice over, verbally and then logically.

Well, pfisherman seems to disagree with that use of the word reasoning.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#68

Earlier quoted context omitted.

> We hold AI to a pretty high standard of correctness, as we should, but humans are not that reliable on matters of fact, let alone on rigor of reasoning. I never understood this line of reasoning. 1. Humans can't run faster than 30 mph. 2. Therefore we can't complain if cars/trains/transport always go slower than 30 mph. These comparisons also hide that we are comparing best of AI (massive LLMs) with median/average…

The whole social dynamic of this conversation is amazing. How fast complacency happened. In 2010 if you told me I could get a model to respond approximately as intelligently as a low intelligence human, I would be amazed. As a matter of perspective,I am still amazed. At the same time I see such negative sentiment around the capabilities at their current limits. We are reaching an era of commodified intelligence, whic…

The main issue is the expectations. The companies behind these models marketed them as being intelligent or close to it so it's natural for people to expect that and react as such when the expectations are not met

Re: Procedural knowledge in pretraining drives reasoning in large language models

#69
post #35

Earlier quoted context omitted.

Crucially, this is what MacIntyre's narrativity thesis is talking about: If a university professor is giving a lecture on decentralized finance and forks into a recipe for chocolate chip cookies: crack two eggs, add a cup of flour, and fold in brown sugar prior to baking, it would break linearity. A generalizable strategy for synthesizing LLMs differentiated by their training parameters is a tokenization is isolating…

> A generalizable strategy for synthesizing LLMs differentiated by their training parameters is a tokenization is isolating data sets and then establishing a lattice in uniformity within the field of technics. This comment appears to be incoherent and likely AI-generated text. Let me break down why: 1. While it uses technical-sounding terms related to machine learning (LLMs, tokenization, data sets), the way they're…

And that analysis is also LLM generated. It's turtles all the way down, folks.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#70
post #9

Earlier quoted context omitted.

You're close, but there’s an important nuance. The process isn't about "learning how to solve problems in general" in the broad sense. It's more specific: the neural network is trained to mimic the step-by-step process demonstrated by humans solving a specific problem. The distinction is that the software doesn't autonomously derive general problem-solving heuristics from scratch. Instead, it observes examples of how…

> it observes Observe implies sentience that, without question, a neural net simply does not possess. "It" certainly 'records', or more specifically it 'maps', but there is no observer in sight (npi). > mimic LLM's do not mimic. The magic is mathematical and happening in the high dimensional space. If there is intrinsic underlying pattern and semantic affinities between process X (used in training) and process Y (use…

Define "observation". If it's just sensory and information processing then no, it does not require nor simply sentience.
Post reply on HN