Live data from Hacker News

LLMs aren't world models

yosefk.com

41–50 of 240 posts

Re: LLMs aren't world models

#41

Earlier quoted context omitted.

> With LLMs being unable to count how many Bs are in blueberry, they clearly don't have any world model whatsoever. Is this a real defect, or some historical thing? I just asked GPT-5: How many "B"s in "blueberry"? and it replied: There are 2 — the letter b appears twice in "blueberry". I also asked it how many Rs in Carrot, and how many Ps in Pineapple, amd it answered both questions correctly too.

It is not historical: https://kieranhealy.org/blog/archives/2025/08/07/blueberry-h... Perhaps they have a hot fix that special cases HN complaints?

They clearly RLHF out the embarrassing cases and make cheating on benchmarks into a sport.

Re: LLMs aren't world models

#42
Agree in general with most of the points, except

> but because I know you and I get by with less.

Actually we got far more data and training than any LLM. We've been gathering and processing sensory data every second at least since birth (more processing than gathering when asleep), and are only really considered fully intelligent in our late teens to mid-20s.

Re: LLMs aren't world models

#43
post #37
post #8

Earlier quoted context omitted.

With LLMs being unable to count how many Bs are in blueberry, they clearly don't have any world model whatsoever. That addition (something which only takes a few gates in digital logic) happens to be overfit into a few nodes on multi-billion node networks is hardly a surprise to anyone except the most religious of AI believers.

The core issue there isn't that the LLM isn't building internal models to represent its world, it's that its world is limited to tokens. Anything not represented in tokens, or token relationships, can't be modeled by the LLM, by definition. It's like asking a blind person to count the number of colors on a car. They can give it a go and assume glass, tires, and metal are different colors as there is likely a correlat…

That is just not a solid argument. There are countless examples of LLMs splitting "blueberry" into "b l u e b e r r y", which would contain one token per letter. And then they still manage to get it wrong.

Your argument is based on a flawed assumption, that they can't see letters. If they didn't they wouldn't be able to spell the word out. But they do. And when they do get one token per letter, they still miscount.

Re: LLMs aren't world models

#44
What with this and your previous post about why sometimes incompetent management leads to better outcomes, you are quickly becoming one of my favorite tech bloggers. Perhaps I enjoyed the piece so much because your conclusions basically track mine. (I'm a software developer who has dabbled with LLMs, and has some hand-wavey background on how they work, but otherwise can claim no special knowledge.) Also your writing style really pops. No one would accuse your post of having been generated by an LLM.

Re: LLMs aren't world models

#45
post #44

What with this and your previous post about why sometimes incompetent management leads to better outcomes, you are quickly becoming one of my favorite tech bloggers. Perhaps I enjoyed the piece so much because your conclusions basically track mine. (I'm a software developer who has dabbled with LLMs, and has some hand-wavey background on how they work, but otherwise can claim no special knowledge.) Also your writing…

thank you for your kind words!

Re: LLMs aren't world models

#46

I just tried a few things that are simple and a world model would probably get right. Eg Question to GPT5: I am looking straight on to some objects. Looking parallel to the ground. In front of me I have a milk bottle, to the right of that is a Coca-Cola bottle. To the right of that is a glass of water. And to the right of that there’s a cherry. Behind the cherry there’s a cactus and to the left of that there’s a pean…

this seems like a strange riddle. In my mind I was thinking that regardless of the glass, all of the objects can be seen (due to perspective, and also the fact you mentioned the locations, meaning you're aware of them). It seems to me it would only actually work in an orthographic perspective, which is not how our reality works

You can tell from the response it does understand the riddle just fine, it just gets it wrong.

Re: LLMs aren't world models

#47
post #3
post #2

https://arxiv.org/abs/2501.17186

This is interesting. The "professional level" rating of However: "A significant Elo rating jump occurs when the model’s Legal Move accuracy reaches 99.8%. This increase is due to the reduction in errors after the model learns to generate legal moves, reinforcing that continuous error correction and learning the correct moves significantly improve ELO" You should be able to reach the move legality of around 100% with…

> r4rk1 pp6 8 4p2Q 3n4 4N3 qP5P 2KRB3 w — — 3 27

Can you say 100% you can generate a good next move (example from the paper) without using tools, and will never accidentally make a mistake and give an illegal move?

Re: LLMs aren't world models

#48
post #8

Earlier quoted context omitted.

With LLMs being unable to count how many Bs are in blueberry, they clearly don't have any world model whatsoever. That addition (something which only takes a few gates in digital logic) happens to be overfit into a few nodes on multi-billion node networks is hardly a surprise to anyone except the most religious of AI believers.

> they clearly don't have any world model whatsoever Then how did an LLM get gold on the mathematical Olympiad, where it certainly hadn’t seen the questions before? How on earth is that possible without a decent working model of mathematics? Sure, LLMs might make weird errors sometimes (nobody is denying that), but clearly the story is rather more complicated than you suggest.

> where it certainly hadn’t seen the questions before?

What are you basing this certainty on?

And even if you're right that the specific questions had not come up, it may still be that the questions from the math olympiad were rehashes of similar questions in other texts, or happened to correspond well to a composition of some other problems that were part of the training set, such that the LLM could 'pick up' on the similarity.

It's also possible that the LLM was specifically trained on similar problems, or may even have a dedicated sub-net or tool for it. Still impressive, but possibly not in a way that generalizes even to math like one might think based on the press releases.

Re: LLMs aren't world models

#49
post #37
post #8

Earlier quoted context omitted.

With LLMs being unable to count how many Bs are in blueberry, they clearly don't have any world model whatsoever. That addition (something which only takes a few gates in digital logic) happens to be overfit into a few nodes on multi-billion node networks is hardly a surprise to anyone except the most religious of AI believers.

The core issue there isn't that the LLM isn't building internal models to represent its world, it's that its world is limited to tokens. Anything not represented in tokens, or token relationships, can't be modeled by the LLM, by definition. It's like asking a blind person to count the number of colors on a car. They can give it a go and assume glass, tires, and metal are different colors as there is likely a correlat…

> It's like asking a blind person to count the number of colors on a car.

I presume if I asked a blind person to count the colors on a car, they would reply “sorry, I am blind, so I can’t answer this question”.

Re: LLMs aren't world models

#50

I just tried a few things that are simple and a world model would probably get right. Eg Question to GPT5: I am looking straight on to some objects. Looking parallel to the ground. In front of me I have a milk bottle, to the right of that is a Coca-Cola bottle. To the right of that is a glass of water. And to the right of that there’s a cherry. Behind the cherry there’s a cactus and to the left of that there’s a pean…

Gemini 2.5 Pro gets this correct on the first attempt, and specifically points out the transparency of the glass of water.

https://g.co/gemini/share/362506056ddb

Time to get the ol' goalpost-moving gloves out.

Post reply on HN