Live data from Hacker News

Language Models Represent Space and Time

arxiv.org

141–150 of 192 posts

Re: Language Models Represent Space and Time

#141
post #41

There was an interesting thought from a Microsoft researcher in an episode of This American Life. He had been given early access to GPT4 and found that it had gained an understanding of gravity and balancing. 3.5 would fail to describe a safe order for stacking 3 eggs, a bottle, a book and a nail. But 4 would give robust answers with logical justification added to each construction step for the tower. The researcher…

> Give the same task to GPT4 and it proves it’s modeled the concept appropriately because it can pull the right words out of the ether. I take your point, but we should be careful when using phrased like “modeled the concept appropriately ” here. That implies correctness. There are many ‘concept models’ that could work well enough. Unless we can inspect the concept model directly (inside the LLM, somehow!) we are lef…

> I take your point, but we should be careful when using phrased like “modeled the concept appropriately” here. That implies correctness.

All models are wrong, but some are useful.

Perhaps we should say "modeled the concept usefully".

> In some ways, my standard here is probably not that different than a rigorous human educational assessment. I’m not talking about standardized tests, I’m talking about adversarial challenges like one would hope to see during a Ph.D. defense.

Usefulness is very contextual obviously. Newton's laws aren't entirely correct but are still useful.

Re: Language Models Represent Space and Time

#142

Earlier quoted context omitted.

If an LLM is trained on just text, then it's internal model can only be based on such text. I think the 'conceit' of LLM as AI advocates is that this model is just as (or at least nearly as) our own internal model - that the 'simplest' model to match the text is indeed our 'real' model that matches reality as a whole. A bit of a stretch surely!

I’m no expert, but the word “just” is tricky. You could say mammalian brain is trained on “just” nerve impulses.

you could, but you’d be wrong. (there are many other signaling modalities!)

Re: Language Models Represent Space and Time

#143

Earlier quoted context omitted.

Since language often describes the world, I think a good language model must include a world model

Well, this has been claimed often, namely by some of the people who have developed statistical language models [1] but it's really not obvious how that should work. Where would a language model find the world model? Where would it store it? And why would it even need it? Obviously I don't know how the human linguistic ability works but it's clear that for us, text, words, language, isn't carrying around with it a rep…

> Where would a language model find the world model? Where would it store it? And why would it even need it?

It is pattern recognition with many layers of abstraction. Obviously it will infer semantic relations at some level. It is a type of machine learning. The entire point of machine learning is to generate a model of the data which can be generalized to new inputs.

It would find the world model in semantic relations in text. It would store it in its vast neural network. It would need a world model in order to perform natural language understanding tasks, which is exactly what it was designed to do. When asked about the world it needs a model of the world in order to generate useful answers.

GPT-4 took months to train on a supercomputer and it generated a neural network of hundreds of gigabytes. What exactly was that supercomputer doing for several months and what exactly would the neural network represent if not a world model? Are you thinking it rote learned answers to every possible question?

> so it seems likely that a world model must be developed first

It is "a" model. It can be completely alien to whatever exists in a human mind but still function and exist as some type of model.

> Why would it ever be possible to derive the representation just from the pointer?

It only needs to have a sufficient representation of the Sea that is useful for the tasks we have trained it to do. It doesn't need to derive "the" one true representation that humans have of the Sea, whatever that is.

Re: Language Models Represent Space and Time

#144
post #38
post #17

Earlier quoted context omitted.

My only experience so far is with chatgpt, and I am not an expert in AI, or even just in LLMs. With those disclaimers, my interactions inform my opinion that the LLM behind chatgpt has no internal world model. It shows no understanding of basic facts and makes very silly mistakes very easily. I have my bias, like anyone else, but in the case of AI in particular, I should say that I don't think there's anything especi…

People are already asking you if you have tried but this pretty much echos my experience and I have used both free and paid versions of ChatGPT (3.5 and 4 respectively) as well as the GPT4 API directly. I had a similar experience between all three of them, although the GPT4 based versions were certainly better in terms of output quality. In my opinion it seems pretty clear that this is a fundamental limit of the arch…

It's so fascinating going through previous GPT-X threads. You can almost everybody making the same kind of extrapolation mistakes you're making. People just don't seem all that interested in revising their model when it's obviously wrong.

Re: Language Models Represent Space and Time

#145

Wait I had the "It's just a stochastic parrot" parroted at me ad nauseum. Was this FUD all along?

The paper it originates from predates ChatGPT (let alone GPT-4). So I think a charitable interpretation is that it was partly wrong (and definitely unimaginative) at the time, and partly it's the case that very few people were expecting "let's just throw more parameters at it" to work and result in emergent general capabilities, which it appears to.

Re: Language Models Represent Space and Time

#146

Earlier quoted context omitted.

If an LLM is trained on just text, then it's internal model can only be based on such text. I think the 'conceit' of LLM as AI advocates is that this model is just as (or at least nearly as) our own internal model - that the 'simplest' model to match the text is indeed our 'real' model that matches reality as a whole. A bit of a stretch surely!

LLMs are a redemonstration of the lesson of sun worship: apes are easily fooled. If the output is indistinguishably 'world modelling', so must its mechanism be. There is no 'world' model in a hashmap from questions to answers about a world . There is only the condition that, given those questions, any old ape would accept those answers. Science is required to determine whether a system has properties, such as a ratio…

Sorry, that's just surface-level tripe. You could make the same dismissive comment about any new technology. You're not saying anything interesting.

Re: Language Models Represent Space and Time

#147

Earlier quoted context omitted.

LLMs are a redemonstration of the lesson of sun worship: apes are easily fooled. If the output is indistinguishably 'world modelling', so must its mechanism be. There is no 'world' model in a hashmap from questions to answers about a world . There is only the condition that, given those questions, any old ape would accept those answers. Science is required to determine whether a system has properties, such as a ratio…

Sorry, that's just surface-level tripe. You could make the same dismissive comment about any new technology. You're not saying anything interesting.

The claim is that you cannot determine whether a system has a 'world model' through inspection of it's abstract token inputs/outputs.

Re: Language Models Represent Space and Time

#148

A lot of HNers were so adamant that the LLMs understand absolutely nothing and that these models are just predicting the next most likely word. I think with this paper it becomes clear that the adamant denial was just human bias talking. LLMs crossed a certain line here. You would think people would be amazed or in awe or react in fear at technological break through a and they often are. The weird part for LLMs was t…

I think LLMs are presenting some uncomfortable philosophical questions for people about how our own brains work and admitting that there is any kind of "intelligence" (even if very basic) in an LLM is an admission that our own brains may work in a similar manner.

This. I think this is the reason why so many people are in denial. Is All of intelligence simply trying to find the best fit curve in an n-dimensional scatter plot of data points?

Re: Language Models Represent Space and Time

#149

Earlier quoted context omitted.

I think LLMs are presenting some uncomfortable philosophical questions for people about how our own brains work and admitting that there is any kind of "intelligence" (even if very basic) in an LLM is an admission that our own brains may work in a similar manner.

For me, one of the most interesting things that have come out of LLMs is the confirmation that humans are very bad at reasoning and, consequently, its' a very bad idea to try and make machines that "think like humans", because that way we'll only make machines with none of the advantages of machines and all the disadvantages of computers. For instance -I'm not trying to be mean and I'm certainly not blaming you in pa…

>-I'm not trying to be mean and I'm certainly not blaming you in particular, because I've seen this very often- but the reasoning that because LLMs can generate language, and humans can generate language not only LLMs are somehow like humans but also humans are like LLMs is not sound.

Nah. Nobody personifies LLMs like this. What you're laying out here is a fundamental mistake that you'd have to be extremely stupid to make. I think barely anyone is making this mistake to even qualify mentioning it.

Seriously who here things that LLMs are anything like humans? That is not the claim. The claim is that LLMs understand you. Intelligence and understanding are clearly orthogonal to "human-like"

Re: Language Models Represent Space and Time

#150

Earlier quoted context omitted.

Yes, I see what you mean. It's frustrating but the word "model" is severely overused in Computer Science and AI and it can cause a lot of confusion. Briefly, a "world model" is a theory possessed by an autonomous agent that describes the entities that exist in the world and how they interact with each other and with the agent, and that the agent can use to make decisions. This is the sense in which "model" is used wh…

Since language often describes the world, I think a good language model must include a world model

A good language model IS a world model. They can be one in the same. Very likely what's going on with something like chatGPT. The world model is simply encoded as text.
Post reply on HN