Live data from Hacker News

Language Models Represent Space and Time

arxiv.org

151–160 of 192 posts

Re: Language Models Represent Space and Time

#151

Earlier quoted context omitted.

Since language often describes the world, I think a good language model must include a world model

Well, this has been claimed often, namely by some of the people who have developed statistical language models [1] but it's really not obvious how that should work. Where would a language model find the world model? Where would it store it? And why would it even need it? Obviously I don't know how the human linguistic ability works but it's clear that for us, text, words, language, isn't carrying around with it a rep…

> Why would it ever be possible to derive the representation just from the pointer?

Of course it cannot derive a representation from just a pointer. Neither humans or machines can do that. The words are not pointers in isolation. They are connected in a network of semantic relations. The model is in the relationships between pointers. Artificial neural network.

Re: Language Models Represent Space and Time

#152

Earlier quoted context omitted.

I dunno. The guy that's freebasing copium sounds more convincing. Sorry.

Well as all you have as a response is an ad hominem you aren't the intended audience anyway. I would love to learn why I am wrong.

You are clearly a smart guy, you are also clearly on the far side of the spectrum. Since it took you a while to notice that I was yanking your chain. That's ok, I'm somewhat of a full spectrum spaz myself.

The problem with smart people is that they have a very hard time to notice when their brain has sidestepped rational thought, and started to go into an emotional latent space. The reason for this is that your emotional life has no problem utilizing complex topics to shield itself from your conscious attention.

Face it. There isn't anything special about what a bio-neuron does. We just choose to ignore a lot of the intricacies of the bio-neuron. Because it is clearly irrelevant in order to sidestep halting. If a task suffers from halts the NN will simply side step the problem. Just like us. We are at a stage of technological development. Where we need to effen stop, and put all our effort into alignment. You realize this is true inside that cholesterol bag you call a brain. Your brain is just doing gymnastics of a preeteen soviet girl level, to stop yourself from getting it. Just get it.

Re: Language Models Represent Space and Time

#153
post #106
post #40

Earlier quoted context omitted.

It was something like Paris - France + Japan = Tokyo I know that's not a source but you could also just google "word2vec", as I recall most of the explainer blog posts had similar examples.

That gets "analogies" rather than "spatial awareness" per-se.

You put quotes around spatial awareness but the commenter being referenced said geography.

Re: Language Models Represent Space and Time

#154
post #17

Earlier quoted context omitted.

My only experience so far is with chatgpt, and I am not an expert in AI, or even just in LLMs. With those disclaimers, my interactions inform my opinion that the LLM behind chatgpt has no internal world model. It shows no understanding of basic facts and makes very silly mistakes very easily. I have my bias, like anyone else, but in the case of AI in particular, I should say that I don't think there's anything especi…

I really don't understand replies like this. This is just the beginning. It is a tiny, tiny view into this new technology. Nobody with an ounce of intelligence looked at the first iPhone and thought "ah well it's just another phone". It's the potential of this new technology that is so impressive. How can you not be impressed by this? I am literally incapable of empathizing with your perspective. All I can think is t…

I felt similarly to you with cryptocurrency but then not much improved in the 10 years since. Maybe LLM's will be different.

Re: Language Models Represent Space and Time

#155

Earlier quoted context omitted.

I'm talking about how many people react to any startling new information or idea. The advent of LLMs is just a relevant example.

No new ideas have been advanced in anything reported here. A state machine change state is all that is going on here. What I'd like to know is why should anyone care about this preprint? Nothing has been stated on how this software will improve the life of the average human being. LK-99 at least would have had lead.somewhere if it worked.

Given that LK99 was a ceramyc it was unlikely to cause any major changes.

Re: Language Models Represent Space and Time

#156
post #154

Earlier quoted context omitted.

I really don't understand replies like this. This is just the beginning. It is a tiny, tiny view into this new technology. Nobody with an ounce of intelligence looked at the first iPhone and thought "ah well it's just another phone". It's the potential of this new technology that is so impressive. How can you not be impressed by this? I am literally incapable of empathizing with your perspective. All I can think is t…

I felt similarly to you with cryptocurrency but then not much improved in the 10 years since. Maybe LLM's will be different.

True but in those 10 Years AI has been moving at a rocket pace. Draw the trend line.

Re: Language Models Represent Space and Time

#157
post #39

Earlier quoted context omitted.

No, I've only used GPT-3 so far. Should I be excited about GPT-4? :) Examples of things I'd file as silly mistakes: - mentioning non-existing settings when asking it about a particular software. - responding to roughly "give me a way to list all rds clusters that are running a version that will reach EOL within the next 12 months" (which I know is almost impossible to do without scraping, as the EOLs are not returned…

> No, I've only used GPT-3 so far. Should I be excited about GPT-4? The difference is extremely dramatic. Any experience with 3 or 3.5 is completely irrelevant for 4.

There is a clear difference, but "extremely dramatic" is obviously hyperbole.

(I use 3.5 and 4, both in the chatgpt interface and via API.)

Re: Language Models Represent Space and Time

#158

Earlier quoted context omitted.

Sorry, that's just surface-level tripe. You could make the same dismissive comment about any new technology. You're not saying anything interesting.

The claim is that you cannot determine whether a system has a 'world model' through inspection of it's abstract token inputs/outputs.

But that's not all this paper is doing. So why bring it up?

Re: Language Models Represent Space and Time

#159
post #39
post #20

Earlier quoted context omitted.

Have you used GPT-4?

No, I've only used GPT-3 so far. Should I be excited about GPT-4? :) Examples of things I'd file as silly mistakes: - mentioning non-existing settings when asking it about a particular software. - responding to roughly "give me a way to list all rds clusters that are running a version that will reach EOL within the next 12 months" (which I know is almost impossible to do without scraping, as the EOLs are not returned…

I have made similar mistakes to all the ones you mentioned. Do I have an internal world model?

Re: Language Models Represent Space and Time

#160
post #143

Earlier quoted context omitted.

Well, this has been claimed often, namely by some of the people who have developed statistical language models [1] but it's really not obvious how that should work. Where would a language model find the world model? Where would it store it? And why would it even need it? Obviously I don't know how the human linguistic ability works but it's clear that for us, text, words, language, isn't carrying around with it a rep…

> Where would a language model find the world model? Where would it store it? And why would it even need it? It is pattern recognition with many layers of abstraction. Obviously it will infer semantic relations at some level. It is a type of machine learning. The entire point of machine learning is to generate a model of the data which can be generalized to new inputs. It would find the world model in semantic relati…

>> GPT-4 took months to train on a supercomputer and it generated a neural network of hundreds of gigabytes. What exactly was that supercomputer doing for several months and what exactly would the neural network represent if not a world model?

I believe GPT-4 was trained on a server farm, not a single computer. In any case what it was doing all that time was going over and over the text in its gigantic training corpus, which was of petabyte size, and optimising the objective P(tₖ|tₖ₋ₙ , ..., tₖ₋₁) i.e. the probability of token k given a "sliding window" of the n preceding (or surrounding) tokens.

There is nothing in this objective that needs a world model, and it is really not obvious why optimising this objective should lead to development of a world model, rather than, or in addition to, a model of the training corpus.

It is easy to see how this stuff works. You can train your own language model easily, although of course it would have to be a smaller language model. For example, you can train a Hidden Markov Model on the text of freely available literary works on project Guttenberg, or on wikipedia pages, and without too much compute (an ordinary laptop will do).

In fact, I recommend that as an exercise and as an experiment to gain a better understanding in how language modelling works, for those who are curious about questions regarding their ability to model something beyond text.

A good textbook to begin with statistical language modelling is "Foundations of Statistical Natural Language Processing" by Manning and Schűtze:

https://nlp.stanford.edu/fsnlp/

Or, just as good, "Speech and Language Processing" by Jurafsky and Martin:

https://web.stanford.edu/~jurafsky/slp3/

Or, if you don't have the time for an entire textbook, "Statistical Language Learning" by Eugene Charniak is an excellent, concise introduction to the subject:

https://archive.org/details/statisticallangu0000char

Post reply on HN