Earlier quoted context omitted.
There are no world models in there, it's trained on arbitrary images/sequences. There are no world models in us, we learn from only specifics in topological space, stitched together in sharp wave ripples. Everything is from detached memories working through optic flow. That's not a world model, it's not even a model. It's an analog. This whole world model thing is another branding phase after language models failed t…
The point is that video generation is not the goal in itself. Just like classifying photos as cat vs dog wasn't the goal in 2013. I know that Sora 2 is not a world model. But what's coming is: Vision-language-action models and planning, spatial AI (SLAM with semantics and 3D reconstruction with interactability and affordance detection). Video diffusion models, photo-to-gaussian-splats, video-to-3D (e.g. from Hunyuan)…
(Unless it's sci-fi and porn that is mainly pushing for human shaped robots.)