Live data from Hacker News

The dawn of a world simulator

odyssey.ml

31–40 of 60 posts

Re: The dawn of a world simulator

#31
post #6

For a minute I was like (spoiler alert) « wow the creepy sci-fi theories from the DEVS tv show is taking place »… then I looked up the video and that’s just video generation at this point

That's where this is headed, though. That's the end game.

This should be interesting then: we’ll finally be able to assert whether time is deterministic and the future and past can be modelled/predicted (if you’ve seen the show you know what I mean)

Re: The dawn of a world simulator

#32

As a machine learning researcher, I don't get why these are called world models. Visually, they are stunning. But it's nowhere near physical. I mean look at that video with the girl and lion. The tail teleports between legs and then becomes attached to the girl instead of the tiger. Just because the visuals are high quality doesn't mean it's a world model or has learned physics. I feel like we're conflating these thi…

The tail teleports and reattaches because that is the sort of thing that happens in this special AI world. Even though it looks like a bug, it's actually a physical process being modelled accurately.

Re: The dawn of a world simulator

#33
I guess this might be a chance to plug the fact that Matrix came up with their own Metaverse thing (for lack of a better word) called Third Room, it represented the rooms you joined as spaces/worlds, they built some limited functionality demos before the funding dried up

Re: The dawn of a world simulator

#34

I feel like there's a bit if a disconnect with the cool video demos demonstrated here and say, the type of world models someone like Yann Lecunn is talking about. A proper world model like Jepa should be predicting in latent space where the representation of what is going on is highly abstract. Video generation models by definition are either predicting in noise or pixel space (latent noise if the diffuser is diffusi…

> Video generation models by definition are either predicting in noise or pixel space

I don't see that this follows "by definition" at all.

Just because your output is pixel values doesn't mean your internal world model is in pixel space.

Re: The dawn of a world simulator

#35

Earlier quoted context omitted.

That's where this is headed, though. That's the end game.

This should be interesting then: we’ll finally be able to assert whether time is deterministic and the future and past can be modelled/predicted (if you’ve seen the show you know what I mean)

I think that's actually already provably false if you're bloody-minded enough. I think the proof lies somewhere like Cantor's diagonalization but applied to reality, something like "if you could produce a model sufficiently complex enough to model the future perfectly it wouldn't fit into this current reality because it would require more than this reality's information"

I'm not saying it couldn't be locally violated, but it seems straightforward philosophically that each nesting doll of simulated reality must be imperfect by being less complicated.

Re: The dawn of a world simulator

#36

As a machine learning researcher, I don't get why these are called world models. Visually, they are stunning. But it's nowhere near physical. I mean look at that video with the girl and lion. The tail teleports between legs and then becomes attached to the girl instead of the tiger. Just because the visuals are high quality doesn't mean it's a world model or has learned physics. I feel like we're conflating these thi…

The tail teleports and reattaches because that is the sort of thing that happens in this special AI world. Even though it looks like a bug, it's actually a physical process being modelled accurately.

I'll remind you I am a ML researcher.

So, you need to say more. Or at least give me some reason to believe you rather than state something as an objective truth and "just trust me". In the long response to a sibling I state more precisely why I have never bought this common conjecture. Because that's what it is, conjecture.

So give me at least some reason to believe you. Because you have neither logos nor ethos. Your answer is in the form of ethos, but without the critical requisites.

Re: The dawn of a world simulator

#37

Earlier quoted context omitted.

> Visually, they are stunning. The input images are stunning, model's result is another disappointing trip to uncanny valley. But we feel Ok as long as the sequence doesn't horribly contradict the original image or sound. That is the world model.

> But we feel Ok as long as the sequence doesn't horribly contradict the original image or sound. Is the error I pointed out not "horribly contradicting"? > That is the world model. I would say that if it is non-physical[0] then it's hard to call it a /world/ model. A world is consistent and has a set of rules that must be followed. I've yet to see a claimed world model that actually captures this behavior. Yet it's…

>A world is consistent and has a set of rules that must be followed.

Large language models are mostly consistent, but they have mistakes even in grammar too, from time to time. And it's usually called a "hallucination". Can't we say physics errors are a kind of "hallucination" too, in a world model? I guess the question is, what hallucination rate are we willing to tolerate.

Re: The dawn of a world simulator

#38

Earlier quoted context omitted.

The tail teleports and reattaches because that is the sort of thing that happens in this special AI world. Even though it looks like a bug, it's actually a physical process being modelled accurately.

I'll remind you I am a ML researcher. So, you need to say more. Or at least give me some reason to believe you rather than state something as an objective truth and "just trust me". In the long response to a sibling I state more precisely why I have never bought this common conjecture. Because that's what it is, conjecture. So give me at least some reason to believe you. Because you have neither logos nor ethos. Your…

I think they're joking.

Re: The dawn of a world simulator

#39

Earlier quoted context omitted.

I'll remind you I am a ML researcher. So, you need to say more. Or at least give me some reason to believe you rather than state something as an objective truth and "just trust me". In the long response to a sibling I state more precisely why I have never bought this common conjecture. Because that's what it is, conjecture. So give me at least some reason to believe you. Because you have neither logos nor ethos. Your…

I think they're joking.

If so, I misread and sorry. Their sarcasm is too on point, mimicking claims I've heard made in earnest.

Re: The dawn of a world simulator

#40
post #37

Earlier quoted context omitted.

> But we feel Ok as long as the sequence doesn't horribly contradict the original image or sound. Is the error I pointed out not "horribly contradicting"? > That is the world model. I would say that if it is non-physical[0] then it's hard to call it a /world/ model. A world is consistent and has a set of rules that must be followed. I've yet to see a claimed world model that actually captures this behavior. Yet it's…

>A world is consistent and has a set of rules that must be followed. Large language models are mostly consistent, but they have mistakes even in grammar too, from time to time. And it's usually called a "hallucination". Can't we say physics errors are a kind of "hallucination" too, in a world model? I guess the question is, what hallucination rate are we willing to tolerate.

It's not about making no mistakes, it's about the categorical type of mistakes.

Let's consider language as a world, in some abstract sense. Lies may (or may not) be consistent here. Do they make sense linguistically? But then think about the category of errors where they start mixing languages and sound entirely nonsensical. That's rare with current LLMs in standard usage but you can still get them to have full on meltdowns.

This is the class of mistakes these models are making, not the failing to recite truth class of mistakes.

(Not a perfect translation but I hope this explanation helps)

Post reply on HN