Live data from Hacker News

Why Fei-Fei Li and Yann LeCun are both betting on "world models"

entropytown.com

21–30 of 105 posts

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#21
post #14

With all due respect, AI is ultimately a capital game. World models aren’t where real B2B customer revenue comes from—at least compared to today’s LLMs; they’re mainly a better story for raising huge amounts of private capital. Hopefully they figure out how to build the next-gen AI architecture along the way.

The most useful models are image, video, and audio models. It makes sense that we'd make the video models more 4D aware. Text really hogged all the attention. Media is where AI is really going to shine. Some of the most profitable models right now are in music, image, and video generation. A lot of people are having a blast doing things they could legitimately never do before, and real working professionals are able…

That just sounds like text with extra steps.

Fundamentally what AGI is trying to do is to encode ability to logic and reason. Tokens, images, video and audio are all just information of different entropy density that is the output of that logic reasoning process or emulation of logic reasoning process.

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#22

If I was smarter, I would have predicted that not only would everyone else figure out that world models are a critical step, but that as a direct consequence the term "world model" would lose all meaning. Maybe next time. That said, Le Cunn's concept in the blog post is the only one worthy of the title.

[dead]

All abstraction of reality are bound to fail, but some abstractions are more convincing (or indeed more useful) than others.

> Any real AI that veers at control will have to adopt a neurobio path

Maybe. Or maybe it's a useless distraction. Only time will tell what signals are meaningful.

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#23
post #14

With all due respect, AI is ultimately a capital game. World models aren’t where real B2B customer revenue comes from—at least compared to today’s LLMs; they’re mainly a better story for raising huge amounts of private capital. Hopefully they figure out how to build the next-gen AI architecture along the way.

The most useful models are image, video, and audio models. It makes sense that we'd make the video models more 4D aware. Text really hogged all the attention. Media is where AI is really going to shine. Some of the most profitable models right now are in music, image, and video generation. A lot of people are having a blast doing things they could legitimately never do before, and real working professionals are able…

>> The most useful models are image, video, and audio models

This is wrong. The vast majority of revenue is being generated by text models because they are so useful.

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#24

Earlier quoted context omitted.

[dead]

All abstraction of reality are bound to fail, but some abstractions are more convincing (or indeed more useful) than others. > Any real AI that veers at control will have to adopt a neurobio path Maybe. Or maybe it's a useless distraction. Only time will tell what signals are meaningful.

[dead]

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#25
post #14

Earlier quoted context omitted.

The most useful models are image, video, and audio models. It makes sense that we'd make the video models more 4D aware. Text really hogged all the attention. Media is where AI is really going to shine. Some of the most profitable models right now are in music, image, and video generation. A lot of people are having a blast doing things they could legitimately never do before, and real working professionals are able…

That just sounds like text with extra steps. Fundamentally what AGI is trying to do is to encode ability to logic and reason. Tokens, images, video and audio are all just information of different entropy density that is the output of that logic reasoning process or emulation of logic reasoning process.

> Fundamentally what AGI is trying to do is to encode ability to logic and reason.

No? The Wason selection task has shown that logic and reason are not really core nor essential to human cognition.

It's really verging on speculation, but see chapter 2 of Jaynes 1976 - in particular the section on spatialization and the features of consciousness.

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#27

I played with Marble yesterday, Fei-Fei/World Labs' new product. It is the most impressed I've been with an AI experience since the first time I saw a model one-shot material code. Sure, its an early product. The visual output reminds me a lot of early SDXL. But just look at what's happened to video in the last year and image in the last three. The same thing is going to happen here, and fast, and I see the vision fo…

I wasn't actually able to use it because the servers were overloaded. What exactly impressed you (or more generally, what does it actually let you do at the moment?).

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#28

Because they are smart enough to realize current LLM tech is nearing a dead end and cannot serve as a full AGI, even ignoring context and hallucination issues, without actual knowledge of the real world.

Most world models so far are based on transformers, no?

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#29

If I was smarter, I would have predicted that not only would everyone else figure out that world models are a critical step, but that as a direct consequence the term "world model" would lose all meaning. Maybe next time. That said, Le Cunn's concept in the blog post is the only one worthy of the title.

[dead]

> There is no mind

Interesting. What is your response to the cogito?

Re: Why Fei-Fei Li and Yann LeCun are both betting on "world models"

#30
And the pendulum swings back toward representation. It is becoming clear that the LLM approach is not adequate to reach what John McCarthy called human-level intelligence:

Between us and human-level intelligence lie many problems. They can be summarized as that of succeeding in the "common-sense informatic situation". [1]

And the search continues...

[1] https://www-formal.stanford.edu/jmc/human.pdf

Post reply on HN