Live data from Hacker News

From word models to world models

arxiv.org

21–30 of 119 posts

Re: From word models to world models

#21

I doubt that word models can lead to world models. To quote Yann LeCun: "The vast majority of our knowledge, skills, and thoughts are not verbalizable. That's one reason machines will never acquire common sense solely by reading text." https://twitter.com/ylecun/status/1368235803147649028

That just seems like an unfounded hot take. Of course we can explain most of our knowledge, skills, and thoughts in words, that's how we don't lose everything when the next generation comes around lol. It's the core reason we're different from animals.

Now sure you can't describe qualia, but that's basically a subjective artefact of how we sense the world and (to add another unfounded hot take) likely not critical to have an understanding of it on a physical level.

Re: From word models to world models

#22
post #10

It's a surprise to see a paper actually try to solve the problem of modelling thought via language. Nevertheless, it begins with far too many hedges: > By scaling to even larger datasets and neural networks, LLMs appeared to learn not only the structure of language, but capacities for some kinds of thinking There's two hypotheses for how LLMs generate apparently "thought-expressing" outputs: Hyp1 -- it's sampling fro…

If it's "absolutely trivial" to show that LLMs don't have the capacity to form thought, then please publish a paper proving that. So all the "stupid" people studying LLMs that can't come up with such trivial proofs can move on to other stuff.

You may wish to read the paper above. But if you want a quick proof:

1. A thought is a representation of a situation

2. A representation generates entailments of that situation

3. Language is many-to-one translation from these representations to symbols

4. Understanding language is reversing these symbols into thoughts (ie., reprs)

So,

5. If agent A understands sentence X then A forms the relevant representation of X.

6. If agent has a representation it can state entailments of S (eg., counter-facutals).

Now, split X into Xc = "canonical descriptions of S" and trivial permutations Xp.

(st. distribution of Xc,Xp is low, but the tokens of Xp are common)

Form entailments of X, say Y -- sentences that are cannonically implied by the truth of X.

7. If the LLM understood that X entails Y, it would be via constructing the repr S -- which entails S regardless of which sentence in X was used.

8. Train an LLM on Xc and it's accuracy on judging Y entailed by Xp is random.

9. Since using Xp sentences cause it to fail, it does not predict Y via S.

QED.

And we can say,

1. Appearing to judge Y entailed-by X is possible via simple sampling of (X, Y) in historical cases. 2. LLMs are just such a sampling.

so,

3. +Inference to the best explanation:

4. LLMs sample historical cases rather than form representations.

Incidentally, "sampling of historical cases" is already something we knew -- so this entire argument is basically unnecessary. And only necessary because PhDs have been turned into start-up hype men.

Re: From word models to world models

#23

Earlier quoted context omitted.

>It is absolutely trivial to show Hyp2 is false No it's not > Current LLMs can produce impressive results on a set of linguistic inputs and then fail completely on others that make trivial alterations to the same underlying domain. >Indeed: because there're no relevant prior cases to sample from in that case. That's not what that tells us. Humans have weird failure modes that look absurd outside the context of evolut…

The "failure modes" in humans do not show we lack the capacity. Eg., do you have capacity to reason about physics? Well if you're extremely drunk, less so. But not if I permute the name of the object . > I've found more often than not, simply changing names of variables Yes, lol --- why do you think that is? Because in the digitised dataset of "everything ever written" those names correspond to places in that dataset…

> less so. But not if I permute the name of the object.

You need to realize that you wrote it on a forum where the most known joke is "there are two hard things in programming". That would immediately show you how this assumption is exactly false.

Re: From word models to world models

#24
post #10

It's a surprise to see a paper actually try to solve the problem of modelling thought via language. Nevertheless, it begins with far too many hedges: > By scaling to even larger datasets and neural networks, LLMs appeared to learn not only the structure of language, but capacities for some kinds of thinking There's two hypotheses for how LLMs generate apparently "thought-expressing" outputs: Hyp1 -- it's sampling fro…

If it's "absolutely trivial" to show that LLMs don't have the capacity to form thought, then please publish a paper proving that. So all the "stupid" people studying LLMs that can't come up with such trivial proofs can move on to other stuff.

"I have a truly marvelous demonstration that LLMs don't have the capacity to form thought which this margin is too narrow to contain."

Re: From word models to world models

#25

It's a surprise to see a paper actually try to solve the problem of modelling thought via language. Nevertheless, it begins with far too many hedges: > By scaling to even larger datasets and neural networks, LLMs appeared to learn not only the structure of language, but capacities for some kinds of thinking There's two hypotheses for how LLMs generate apparently "thought-expressing" outputs: Hyp1 -- it's sampling fro…

> It is absolutely trivial to show Hyp2 is false

To investigate precisely this question in a clear and unambiguous way, I trained an LLM from scratch to sort lists of numbers. It learned to sort them correctly, and the entropy is such that it's absolutely impossible that it could have done this by Hyp1 (sampling from similar text in the training set).

https://jbconsulting.substack.com/p/its-not-just-statistics-...

Now, there is room to argue that it applies a world-model when given lists of numbers with a hidden logical structure, but not when given lists of words with a hidden logical structure, but I think the ball is in your court to make that argument. (And to a transformer, it only ever sees lists of numbers anyway).

Re: From word models to world models

#26

Earlier quoted context omitted.

> Then they don't in LLMs too LLMs don't get drunk . If a child answers questions from a book of answers then they'll appear to understand the domain insofar as those questions appear. They do not. They will fail to answer questions under, eg., permutations of words (say, a question asks about "norepinephrine" but the book only contains "noradrenaline" etc.). Insofar as a human cannot answer questions under trivial l…

>Insofar as a human cannot answer questions under trivial linguistic permutations then they too do not understand the domain. alright let me humor you for a bit. Lets start with some solid examples of GPT-4 failing this "trivial linguistic permutation" then ?

see, just one reference in the paper: https://arxiv.org/pdf/2302.08399.pdf

Re: From word models to world models

#27
post #11

Earlier quoted context omitted.

Don't try to ham-fist scientific sounding wording into your (very unscientific) argument. This is not a disproof of anything because you failed to define what it means to have the ability to form rational thoughts. With a definition, you would then wanna prove this for humans as a sanity check: Do we never make stupid mistakes? Ok, we make fewer of those than LLMs. Then what is the threshold for accuracy after which…

This entire paper is written as a disproof of the distributional hypothesis. If you want to understand why it's a profoundly unhelpful pseudoscientific idea, this paper is a good start. The test for a capacity C in a system1 has nothing to do with proxy measures of that capacity in system2. The capacity for an oven to cook food may be measured by how much smoke it lets of when burning -- but no amount of "smoke" esta…

>The capacity for an oven to cook food may be measured by how much smoke it lets of when burning -- but no amount of "smoke" establishes that a dry ice machine can cook.

You seem to be talking past me, as nowhere did I claim that LLMs are intelligent. That's the point – Unlike you I do not claim to be able to prove or disprove this. I argue that your comment is the one that is pseudoscientific because you didn't provide (even a semblance of) a rigorous definition of intelligence.

Re: From word models to world models

#28
post #10

Earlier quoted context omitted.

If it's "absolutely trivial" to show that LLMs don't have the capacity to form thought, then please publish a paper proving that. So all the "stupid" people studying LLMs that can't come up with such trivial proofs can move on to other stuff.

You may wish to read the paper above. But if you want a quick proof: 1. A thought is a representation of a situation 2. A representation generates entailments of that situation 3. Language is many-to-one translation from these representations to symbols 4. Understanding language is reversing these symbols into thoughts (ie., reprs) So, 5. If agent A understands sentence X then A forms the relevant representation of X…

> Train an LLM on Xc and it's accuracy on judging Y entailed by Xp is random.

Why? This is obviously wrong in general case. For that to be true Xp and Xc has to have no statistical relationship whatsoever, which statistically is virtually impossible.

Re: From word models to world models

#29
Humans come in all shapes and forms of sensory as well as cognitive abilities. Our true ability to be human comes from objectives (derived from biological and socially bound complex systems) that drive us, feedback loops (ability to morph / affect the goals) and continuous sensory capabilities.

Reasoning is just prediction with memory towards an objective.

Once large models have these perpetual operating sensory loops with objective functions, the ability to distinguish model powered intelligence and human like intelligence tends to drop.

Re: From word models to world models

#30
post #25

It's a surprise to see a paper actually try to solve the problem of modelling thought via language. Nevertheless, it begins with far too many hedges: > By scaling to even larger datasets and neural networks, LLMs appeared to learn not only the structure of language, but capacities for some kinds of thinking There's two hypotheses for how LLMs generate apparently "thought-expressing" outputs: Hyp1 -- it's sampling fro…

> It is absolutely trivial to show Hyp2 is false To investigate precisely this question in a clear and unambiguous way, I trained an LLM from scratch to sort lists of numbers. It learned to sort them correctly, and the entropy is such that it's absolutely impossible that it could have done this by Hyp1 (sampling from similar text in the training set). https://jbconsulting.substack.com/p/its-not-just-statistics-... No…

So this is a really good starting point -- but you havent formulated any hypotheses that can be tested. You've just looked at the graph and "reckoned something".

Formally, what hypotheses are you comparing? What do you think the specific hypothesis of the "AI = stats" person is? It isnt that the NN literally remembers data tokens, right?

In any case:

The issue with forcing NNs to model mathematical features is that the structure of the data itself has those properties. So the distributional hypothesis is true for sorting ordinals.

But it's really obviously false for natural language. The properties of the world are not the properties of word order... being red isnt "red follows words like...".

Post reply on HN