Live data from Hacker News

Procedural knowledge in pretraining drives reasoning in large language models

arxiv.org

71–80 of 104 posts

Re: Procedural knowledge in pretraining drives reasoning in large language models

#71

Going meta a bit: comments so far on this post show diametrically opposing understandings of the paper, which demonstrates just how varied the interpretation of complex text can be. We hold AI to a pretty high standard of correctness, as we should, but humans are not that reliable on matters of fact, let alone on rigor of reasoning.

This is extremely common in these discussions. Most humans are not that good at reasoning themselves and fall for the same kind of fallacies over and over because of the way they were brought up (their training data so to speak). And yet they somehow think they can argue why or why not LLMs should be able to do the same. If anything, the current limits of these morels show the limits of human cognition which is sprea…

Human society, when the individual reaches the limits of their "reasoning", usually produce growths to circumvent these limitations to produce and use artifacts that lurk beyond their limitations. A illiterate can still use Netflix, etc.

The ability to circumvent these limitations, is encoded in company procedures, architecture of hierarchies/gremiums within companies and states. Could AI be "upgraded" beyond human reasoning, by referencing these "meta-organisms" and their reasoning processes that can produce things that are larger then the sum of its parts?

Could AI become smarter by rewarding this meta-reasoning and prompting for it?

"Chat GPT for your next task, you are going to model a company reasoning process internally to produce a better outcome"

This should also allow to circumvent human reasoning bugs - like tribal thinking (which is the reason why we have black and white thinking. You goto agree with the tribes-group-think, else there be civil war risking all members of the tribe. Which is why there always can only be ONE answer, one idea, one plan, one leader - and multiple simultaneous explorations at once as in capitalism cause deep unease)

Re: Procedural knowledge in pretraining drives reasoning in large language models

#72
post #54
post #44

Earlier quoted context omitted.

> Most humans are not that good at reasoning themselves and fall for the same kind of fallacies over and over because of the way they were brought up Disagree that it's easy to pin on "how they were brought up". It seems very likely that we may learn that the flaws are part of what makes our intelligence "work" and be adaptive to changing environments. It may be favourable in terms of cultural evolution for parents t…

> Disagree that it's easy to pin on "how they were brought up". Indeed. That might play a role, but another less politically charged aspect to look at is just: how much effort is the human currently putting in? Humans are often on autopilot, perhaps even most of the time. Autopilot means taking lazy intellectual shortcuts. And to echo your argument: in familiar environments those shortcuts are often a good idea! If y…

> Humans are often on autopilot, perhaps even most of the time

I wonder what responsiveness/results someone would get running an LLM with just ~20 watts for processing and memory, especially if it was getting trained at the same time.

That said, we do have a hardware advantage, what with the enormous swarm of nano-bots using technology and techniques literally beyond our best science. :p

Re: Procedural knowledge in pretraining drives reasoning in large language models

#73
post #72
post #54

Earlier quoted context omitted.

> Disagree that it's easy to pin on "how they were brought up". Indeed. That might play a role, but another less politically charged aspect to look at is just: how much effort is the human currently putting in? Humans are often on autopilot, perhaps even most of the time. Autopilot means taking lazy intellectual shortcuts. And to echo your argument: in familiar environments those shortcuts are often a good idea! If y…

> Humans are often on autopilot, perhaps even most of the time I wonder what responsiveness/results someone would get running an LLM with just ~20 watts for processing and memory, especially if it was getting trained at the same time. That said, we do have a hardware advantage, what with the enormous swarm of nano-bots using technology and techniques literally beyond our best science. :p

There's another advantage:

Humans and human language have co-evolved to be compatible. Language makes no such allowance for the needs and quirks of LLMs. (However to a certain extent we design our LLMs to be able to deal with human language.)

Re: Procedural knowledge in pretraining drives reasoning in large language models

#75

Earlier quoted context omitted.

> it observes Observe implies sentience that, without question, a neural net simply does not possess. "It" certainly 'records', or more specifically it 'maps', but there is no observer in sight (npi). > mimic LLM's do not mimic. The magic is mathematical and happening in the high dimensional space. If there is intrinsic underlying pattern and semantic affinities between process X (used in training) and process Y (use…

Define "observation". If it's just sensory and information processing then no, it does not require nor simply sentience.

There is a word for that: a 'recording'. There is no observer thus no observation.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#76
post #25
post #5

This would explain the unexpected benefits of training on code.

That sounds interesting, but I'm a layman and don't know anything about it. Can you provide a link? I was able to find https://arxiv.org/abs/2408.10914 , but I don't have the context to know whether it's the paper you're talking about.

There was an interview with Zuckerberg about how they initially split training llama chat models on purely normal text and codellama on code, but later realized that if they combine the training set they get a model that is better at both tasks than each specialized one was.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#77

>On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies surprised this gets voted up given the surprising amount of users on HN who think LLMs can't reason at all and that the only way to characterize an LLM is through the lens of a next token predictor. La…

The reality is that both can be true at the same time. Yes they're next token predictors, but sometimes the only way to do that correctly is by actually understanding everything that came before and reasoning logically about it. There's some Sutskever quote that if the input to a model is most of a crime novel, and the next token is the name of the perpetrator, then the model understood the novel.

Transformers are arbitrary function approximators, so there's no hard limitation on what they can or cannot do.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#78
post #48

Earlier quoted context omitted.

Why isn't LLM pre-training based on next token prediction considered "trial and error"? It seems to fit that description pretty well to me.

There is more than one correct answer in reality, LLM pre-training just trains it to respond the same way as the text did. Imagine if school only gave correct if you used exactly the same words as the book, that is not "trial and error".

I can tell you haven't been in a school in while. That is actually a pretty accurate description of what schools are like nowadays.

Re: Procedural knowledge in pretraining drives reasoning in large language models

#79
post #26

It seems obvious to me that LLMs wouldn't be able to find examples of every single problem posed to them in training data. There wouldn't be enough examples for the factual look up needed in an information retrieval style search. I can believe that they're doing some form of extrapolation to create novel solutions to posed problems. It's interesting that this paper doesn't contradict the conclusions of the Apple LLM…

I can believe that they're doing some form of extrapolation to create novel solutions to posed problems You can believe it what sort of evidence are you using for this belief? Edit: Also, the abstract of the Apple paper hardly says "corruption" (implying something tricky), it says that they changed the initial numerical values

Changing numerical values doesn't do anything to impact the performance of state of the art models (4o, o1-mini, preview)

The only thing that does is the benchmark that introduces "seemingly relevant but ultimately irrelevant information"

Re: Procedural knowledge in pretraining drives reasoning in large language models

#80
post #57

Earlier quoted context omitted.

The whole social dynamic of this conversation is amazing. How fast complacency happened. In 2010 if you told me I could get a model to respond approximately as intelligently as a low intelligence human, I would be amazed. As a matter of perspective,I am still amazed. At the same time I see such negative sentiment around the capabilities at their current limits. We are reaching an era of commodified intelligence, whic…

> Even the current limited models change the economics dramatically. Yes, though at the moment they hype is still a lot bigger than the impact. But I am fairly confident that even without any new technical ideas for the networks themselves, we will see a lot more economic impact over the next few years, as people work out how to use these new tools. (Of course, the networks will also evolve still.)

I think the impact is already understated. Every non-technical person I know that's still working has used ChatGPT for work at some point, and quite a few of them are using it regularly. And I'm nowhere near Silicon Valley or any serious tech hub.
Post reply on HN