Live data from Hacker News

LLMs aren't world models

yosefk.com

231–240 of 240 posts

Re: LLMs aren't world models

#231
post #146

Earlier quoted context omitted.

We're not just stochastic parrots though, we can parrot things stochastically when that has utility, but we can also be original. The first time that work was done, it was sone by a person, autonomously. Current LLMs couldnt have done it the first time

They are much more than stochastic parrots. I have never understood the stochastic parrot interpretation. LLMs (and general deep learning models) are not statistical/stochastic based models. Statistics trivially apply, as they apply to all measurements of judge-able behavior. But the models do not perform statistical operations, nor do their architectures form tunable statistically driven systems. They learn topologi…

> LLMs (and general deep learning models) are not statistical/stochastic based models. Statistics trivially apply, as they apply to all measurements of judge-able behavior. But the models do not perform statistical operations, nor do their architectures form tunable statistically driven systems.

And just like a LLM, confidently wrong.

Re: LLMs aren't world models

#232
post #141

Earlier quoted context omitted.

I think it's hard to take any LLM criticism seriously if they don't even specify which model they used. Saying "an LLM model" is totally useless for deriving any kind of conclusion.

When talking about the capabilities of a class of tools long term, it makes sense to be general. I think deriving conclusions at all is pretty difficult given how fast everything is moving, but there is some realities we do actually know about how LLMs work and we can talk about that. Knowing that ChatGPT output good tokens last tuesday but Sonnet didn't does not help us know much about the future of the tools on gen…

> Knowing that ChatGPT output good tokens last tuesday but Sonnet didn't does not help us know much about the future of the tools on general.

Isnt that exactly what is going to help us understand the value these tools bring to end-users, and how to optimize these tools for better future use? None of these models are copy+pastes, they tend to be doing things slightly differently under the hood. How those differences affect results seems like the exact data we would want here

Re: LLMs aren't world models

#233

Earlier quoted context omitted.

this seems like a strange riddle. In my mind I was thinking that regardless of the glass, all of the objects can be seen (due to perspective, and also the fact you mentioned the locations, meaning you're aware of them). It seems to me it would only actually work in an orthographic perspective, which is not how our reality works

You can tell from the response it does understand the riddle just fine, it just gets it wrong.

I'm not sure it does. I did ask 5 adults this question, with zero context about what we're discussing with AI, just posing it as a riddle. They were split, with lots of uncertainty about the optics of the glass straight on, and where the viewer is vertically with respect the glass' height. The best response, from my wife, brought up the trig aspect to the problem, pointing out that nothing in the question talks about distance to or between the objects. Her assumption was that the peanut could easily be offset behind the glass from her perspective, resulting in the peanut not being visible.

We tried the experiment, and she was right that there are definitely setups of distance and spacing that cause the peanut to not be visible. Try it!

Re: LLMs aren't world models

#234
post #175
post #123

Earlier quoted context omitted.

I don’t know what you mean by “unidirectional relation”. I get that you gave an explanation after the colon, but I still don’t quite get what you mean. Do you just mean that what words I use doesn’t pick out a unique possible experience? That’s true of course, but I don’t know why you call that “unidirectional” I don’t think describing colors to a blind person is nonsense. One can speak of how the different colors re…

IMO, the GP's idea is that you can't explain sounds to a deaf man, or emotions to someone who doesn't feel them. All that needs direct experience and words only point to our shared experience.

Ok, but you can explain properties of sounds to deaf men, and properties of colors to blind men. You can’t give them a full understanding of what it is like to experience these things, but that doesn’t preclude deaf or blind men from having mental models of the world that take into account those senses. A blind man can still reason about what things a sighted person would be able to conclude based on what they see, likewise a deaf man can reason about what a person who can hear could conclude based on what they could hear.

Re: LLMs aren't world models

#235
post #141

Earlier quoted context omitted.

When talking about the capabilities of a class of tools long term, it makes sense to be general. I think deriving conclusions at all is pretty difficult given how fast everything is moving, but there is some realities we do actually know about how LLMs work and we can talk about that. Knowing that ChatGPT output good tokens last tuesday but Sonnet didn't does not help us know much about the future of the tools on gen…

> Knowing that ChatGPT output good tokens last tuesday but Sonnet didn't does not help us know much about the future of the tools on general. Isnt that exactly what is going to help us understand the value these tools bring to end-users, and how to optimize these tools for better future use? None of these models are copy+pastes, they tend to be doing things slightly differently under the hood. How those differences a…

I guess I disagree that the main concern is the differences per each model, rather than the overall technology of LLMs in general. Given how fast it's all changing, I would rather focus on the broader conversation personally. I don't really care if GPT5 is better at benchmarks, I care that LLMs are actually capable of the type of reasoning and productive output that the world currently thinks they are.

Re: LLMs aren't world models

#236
post #235

Earlier quoted context omitted.

> Knowing that ChatGPT output good tokens last tuesday but Sonnet didn't does not help us know much about the future of the tools on general. Isnt that exactly what is going to help us understand the value these tools bring to end-users, and how to optimize these tools for better future use? None of these models are copy+pastes, they tend to be doing things slightly differently under the hood. How those differences a…

I guess I disagree that the main concern is the differences per each model, rather than the overall technology of LLMs in general. Given how fast it's all changing, I would rather focus on the broader conversation personally. I don't really care if GPT5 is better at benchmarks, I care that LLMs are actually capable of the type of reasoning and productive output that the world currently thinks they are.

Sure, but if you're making a point about LLMs in general, you need to use examples from best-in-class models. Otherwise your examples of how these models fail are meaningless. It would be like complaining about how smartphone cameras are inherently terrible, but all your examples of bad photos aren't labeled with what phone was used to capture. How can anyone infer anything meaningful from that?

Re: LLMs aren't world models

#237
I've been looking for a way to play chess/go with a generalist LLM. It wouldn't matter if the moves are bad (I like winning), but being able to chat on unrelated topics while playing the game would take the experience to the next level.

Re: LLMs aren't world models

#238
post #6

Earlier quoted context omitted.

Your being blunt is actually very kind, if you're describing what I'm doing as "being smart and reasoning from first principles"; and I agree that I am not saying something very novel, at most it's slightly contrarian given the current sentiment. My goal is not to cherry-pick failures for its own sake as much as to try to explain why I get pretty bad output from LLMs much of the time, which I do. They are also very u…

Your LLM output seems abnormally bad, like you are using old models, bad models, or intentionally poor prompting. I just copied and pasted your Krita example into ChatGPT, and reasonable answer, nothing like what you paraphrased in your post. https://imgur.com/a/O9CjiJY

The examples are from the latest versions of ChatGPT, Claude, Grok, and Google AI Overview. I did not bother to list the full conversations because (A) LLMs are very verbose and (B) nothing ever reproduces, so in any case any failure is "abnormally bad." I guess dismissing failures and focusing on successes is a natural continuation of our industry's trend to ship software with bugs which allegedly don't matter because they're rare, except with "AI" the MTBF is orders of magnitude shorter

Re: LLMs aren't world models

#239

Earlier quoted context omitted.

> If you want to be pedantic about word definitions, it absolutely is AGI: artificial general intelligence. This isn't being pedantic, it's deliberately misinterpreting a commonly used term by taking every word literally for effect. Terms, like words, can take on a meaning that is distinct from looking at each constituent part and coming up with your interpretation of a literal definition based on those parts.

I didn't invent this interpretation. It's how the word was originally defined, and used for many, many decades, by the founders of the field. See for example: https://www-formal.stanford.edu/jmc/generality.pdf Or look at the old / early AGI conference series: https://agi-conference.org Or read any old, pre-2009 (ImageNet) AI textbook. It will talk about "narrow intelligence" vs "general intelligence," a dichotomy tha…

I'm a week late, but I do appreciate you pointing out this real phenomenon of moving the goalpost. Language is really general, multimodal models even more-so. The idea that AGI should be way more anthropomorphic and omnipotent is really recent. New definitions almost disregard the possibility of stupid general intelligence, despite proof-by-existence living all around us.
Post reply on HN