Live data from Hacker News

LLMs aren't world models

yosefk.com

221–230 of 240 posts

Re: LLMs aren't world models

#221
post #86

Earlier quoted context omitted.

But this is parallel to saying LLMs are not "compelled" by the training algorithms to learn symbolic logic. Which says to me there are two camps on this and the verdict is still out on this and all related questions.

>LLMs are not "compelled" by the training algorithms to learn symbolic logic. I think "compell" is such a unique human trait that machine will never replicate to the T. The article did mention specifically about this very issue: "And of course people can be like that, too - eg much better at the big O notation and complexity analysis in interviews than on the job. But I guarantee you that if you put a gun to their he…

And yet there are two camps on the matter. Experts like Hinton disagree, others agree.

Re: LLMs aren't world models

#222
post #185

The post is based on a misconception. If you read the blog post linked at the end of this message, you'll see how a very small GPT-2 alike transformer (Karpathy nano-gpt trained to a very small size) after seeing just PGN games and nothing more develops an 8x8 internal representation with which chess piece is where. This representation can be extracted by linear probing (and can be even altered by using the probe in…

The post or rather the part you refer to is based on a simple experiment which I encourage you to repeat. (It is way likelier to reproduce in the short to medium run than the others.) From your link: "...The first was gpt-3.5-turbo-instruct's ability to play chess at 1800 Elo" These things don't play at 1800 ELO, though maybe someone measured this ELO without cheating but rather relying on some artifacts of how an en…

>These things don't play at 1800 ELO

Why are you saying 'these things'?. That statement is about a specific model which did play at that level and did not lose track of the pieces. There's no cheating or weirdness.

https://github.com/adamkarvonen/chess_gpt_eval

Re: LLMs aren't world models

#223
post #159

Earlier quoted context omitted.

I'm pretty sure you can do that right now in Claude Code with the right subagent definitions. (For what it's worth, I respect and greatly appreciate your willingness to put out a prediction based on real evidence and your own reasoning. But I think you must be lacking experience with the latest tools & best practices.)

If you're right, there will soon be a flood of software teams with no programmers on them - either across all domains, or in some domains where this works well. We shall see. Indeed I have no experience with Claude Code, but I use Claude via chat, and it fails all the time on things not remotely as hard as orientation in a large code base. Claude Code is the same thing with the ability to run tools. Of course tools h…

I am more skeptical of the rate of AI progress than many here, but Claude Code is a huge step. There were a few "holy shit" moments when I started using it. Since then, after much more experimentation, I see its limits and faults, and use it less now. But I think it's worth giving it a try if you want to be informed about the current state of LLM-assisted programming.

Re: LLMs aren't world models

#224

Earlier quoted context omitted.

Claude Code isn't an LLM. It's a hybrid architecture where an LLM provides the interface and some of the reasoning, embedded inside a broader set of more or less deterministic tools. It's obvious LLMs can't do the job without these external tools, so the claim above - that LLMs can't do this job - is on firm ground. But it's also obvious these hybrid systems will become more and more complex and capable over time, an…

If you want to be pedantic about word definitions, it absolutely is AGI: artificial general intelligence. Whether you draw the system boundary of an LLM to include the tools it calls or not is a rather arbitrary distinction, and not very interesting.

> If you want to be pedantic about word definitions, it absolutely is AGI: artificial general intelligence.

This isn't being pedantic, it's deliberately misinterpreting a commonly used term by taking every word literally for effect. Terms, like words, can take on a meaning that is distinct from looking at each constituent part and coming up with your interpretation of a literal definition based on those parts.

Re: LLMs aren't world models

#225

Earlier quoted context omitted.

If you want to be pedantic about word definitions, it absolutely is AGI: artificial general intelligence. Whether you draw the system boundary of an LLM to include the tools it calls or not is a rather arbitrary distinction, and not very interesting.

> If you want to be pedantic about word definitions, it absolutely is AGI: artificial general intelligence. This isn't being pedantic, it's deliberately misinterpreting a commonly used term by taking every word literally for effect. Terms, like words, can take on a meaning that is distinct from looking at each constituent part and coming up with your interpretation of a literal definition based on those parts.

I didn't invent this interpretation. It's how the word was originally defined, and used for many, many decades, by the founders of the field. See for example:

https://www-formal.stanford.edu/jmc/generality.pdf

Or look at the old / early AGI conference series:

https://agi-conference.org

Or read any old, pre-2009 (ImageNet) AI textbook. It will talk about "narrow intelligence" vs "general intelligence," a dichotomy that exists more in GOFAI than the deep learning approaches.

Maybe I'm a curmudgeon and this is entering get-off-my-lawn territory, but I find it immensely annoying when existing clear terminology (AGI vs ASI, strong vs weak, narrow vs. general) is superseded by a confused mix of popular meanings that lack any clear definition.

Re: LLMs aren't world models

#226

Earlier quoted context omitted.

> If you want to be pedantic about word definitions, it absolutely is AGI: artificial general intelligence. This isn't being pedantic, it's deliberately misinterpreting a commonly used term by taking every word literally for effect. Terms, like words, can take on a meaning that is distinct from looking at each constituent part and coming up with your interpretation of a literal definition based on those parts.

I didn't invent this interpretation. It's how the word was originally defined, and used for many, many decades, by the founders of the field. See for example: https://www-formal.stanford.edu/jmc/generality.pdf Or look at the old / early AGI conference series: https://agi-conference.org Or read any old, pre-2009 (ImageNet) AI textbook. It will talk about "narrow intelligence" vs "general intelligence," a dichotomy tha…

The McCarthy paper doesn't use the term "artificial general intelligence" anywhere. It does use the word "general" a lot in relation to artificial intelligence.

I looked at the AGI conference page for 2009: https://agi-conference.org/2009/

When it uses the term "artificial general intelligence", it hyperlinks to this page: http://www.agiri.org/wiki/index.php?title=Artificial_General...

Which seems unavailable, so here is an archive from 2007: https://web.archive.org/web/20070106033535/http://www.agiri....

And that page says "In Nov. 1997, the term Artificial General Intelligence was first coined by Mark Avrum Gubrud in the abstract for his paper Nanotechnology and International Security". And here is that paper: https://web.archive.org/web/20070205153112/http://www.foresi...

That paper says: "By advanced artificial general intelligence, I mean AI systems that rival or surpass the human brain in complexity and speed, that can acquire, manipulate and reason with general knowledge, and that are usable in essentially any phase of industrial or military operations where a human intelligence would otherwise be needed."

I think that your insisting that AGI means something different than what everyone else means when they say it is not useful, and will only lead to people getting confused and disagreeing with you. I agree that it's not a great term.

Re: LLMs aren't world models

#227
post #200

Earlier quoted context omitted.

> but expecting it to lead to new insights ignores the fact that it's fundamentally an abstraction of the real, not in relationship to it. Where do humans get new insights from?

Generally the experience of insight is prior to any discursive expression. We put our insights in terms of words, they do not arise as such.

Like VLMs then.

Re: LLMs aren't world models

#228
post #203

Earlier quoted context omitted.

> I believe, that human mind does something like that all the time Absolutely not. Human brains have online one-shot training. LLMs weights are fixed and fine-tuning them is a huge multi-year enterprise. Fundamentally it's two completely different architectures.

I really don't like how you rejecting the idea completely. People have online one-shot training, but have you tried to learn how to play on piano? To learn it you need a lot of repetitions. Really a lot. You need a lot of repetitions to learn how to walk, or how to do arithmetic, or how to read English. This is very similar to LLMs, isn't it? So they are not completely different architectures, aren't they? It is more…

> This is very similar to LLMs, isn't it?

No, it isn't at all. The effort humans spend on rote learning is to optimize mechanical precision in performance, not to internalize the concepts.

The concepts of playing the piano you can learn in a couple days. All the rest of the effort is about getting synchronization and timing right.

Re: LLMs aren't world models

#229

One thing I appreciated about this post, unlike a lot of AI-skeptic posts, is that it actually makes a concrete falsifiable prediction; specifically, "LLMs will never manage to deal with large code bases 'autonomously'". So in the future we can look back and see whether it was right. For my part, I'd give 80% confidence that LLMs will be able to do this within two years, without fundamental architectural changes.

> LLMs will never manage to deal with large code bases 'autonomously' Absolutely nothing about that statement is concrete or falsifiable. Hell, you can already deal with large code bases 'autonomously' without LLMs - grep and find and sed goes a long way!

Seems falsifiable to me? If an LLM (+harness) is fully maintaining a project, updating things when dependencies update, handling bug reports, etc., in a way that is considered decent quality by consumers of the project, then that seems like it would falsify it.

Now, that’s a very high bar, and I don’t anticipate it being cleared any time soon.

But I do think if it happened, it would pretty clearly falsify the hypothesis .

Re: LLMs aren't world models

#230
It might be worth noting that humans also struggle with keeping up a coherent world model over time.

Luckily, we don’t have to; we externalize a lot of our representations. When shopping together with a friend we might put our stuff on one side of the shopping cart and our friends’ on the other. There’s a reason we don’t just play chess in our heads but use a chess board. We use notebooks to write things down, etc.

Some reasoning model can do similar things (keep a persistent notebook that gets fed back into the context window on every pass), but I expect that we need a few more dirty representational ist tricks to get there.

In other words, I don’t think it’s an LLMs job to have a world model, but an LLM is just one part of an AI system.

Post reply on HN