Earlier quoted context omitted.
I'm nowhere near an expert but it seemed like he was claiming true understanding of latent space is needed for generating coherent continuations, but Sora demo already has longish videos that are coherent. It's hard for me not to think this is someone just trying to still be right when they are wrong, but I may misunderstand.
The way i understand it is that Sora is mostly just 'moving pictures' with no rhyme or reason. Yann Lecun is interested in videos that tell a 'story', with cause and effect. Like a magician putting his hand into a top hat and pulling out a rabbit, kind of video.
Let me clear a huge misunderstanding
21–30 of 109 posts
Re: Let me clear a huge misunderstanding
#22Earlier quoted context omitted.
I'm nowhere near an expert but it seemed like he was claiming true understanding of latent space is needed for generating coherent continuations, but Sora demo already has longish videos that are coherent. It's hard for me not to think this is someone just trying to still be right when they are wrong, but I may misunderstand.
The way i understand it is that Sora is mostly just 'moving pictures' with no rhyme or reason. Yann Lecun is interested in videos that tell a 'story', with cause and effect. Like a magician putting his hand into a top hat and pulling out a rabbit, kind of video.
Re: Let me clear a huge misunderstanding
#23Earlier quoted context omitted.
The way i understand it is that Sora is mostly just 'moving pictures' with no rhyme or reason. Yann Lecun is interested in videos that tell a 'story', with cause and effect. Like a magician putting his hand into a top hat and pulling out a rabbit, kind of video.
Yeah, the same way ChatGPT is only predicting the next word with no rhyme or reason. However to actually predict the next word so that the entire sentence makes sense and is relevant for the context (e.g. answers a question) you probably must be actually understanding the meaning of the words and the language and have a world model. I can't imagine a NN moving around pixels in the shape of a cat with no understanding…
Re: Let me clear a huge misunderstanding
#24Earlier quoted context omitted.
Yeah, the same way ChatGPT is only predicting the next word with no rhyme or reason. However to actually predict the next word so that the entire sentence makes sense and is relevant for the context (e.g. answers a question) you probably must be actually understanding the meaning of the words and the language and have a world model. I can't imagine a NN moving around pixels in the shape of a cat with no understanding…
Of course. But stable diffusion can do that. It understands what a cat is and can draw cat wearing a hat. Videos are about actions, cause and effect, which is entirely different thing than still pictures.
Re: Let me clear a huge misunderstanding
#25I feel LeCun got roped in debating the likes of Marcus and Yudkowsky. This has made his arguments lose nuance and become rigid. I also can't escape the feeling that if Facebook was tuned into Transformers, they would have shipped earlier, so there must have been some resistance or underestimation that's now repeated "They can't reason", "They can't plan", "They can't understand the world", "They are a distraction / s…
FWIW none of the video models released so far demonstrate any object coherence whatsoever, which suggests they don't have the higher level capabilities you mention yet.
In Sora, as soon as an object is obstructed by an obstacle or goes offscreen, it's likely to disappear or be radically transformed.
Re: Let me clear a huge misunderstanding
#26More like: "Let me clear a huge misunderstanding here by using poorly-defined terms and cramming niche complex ideas elaborated elsewhere into this tweet." Someone correct me, but it seems like he's saying: 1- generative models don't understand the real world 2- generative models that work off of just pixels are more expensive and less useful than a model that represents the contents of the frame with abstract repres…
1. The next frame is easy, but multiple frames is not
2. What works for text doesn't work for video.
Then Sora comes out and shows multiple frames and someone tweets gotcha.
He then tweets without saying he misspoke ..... goes on about the model doesn't understand physics.
And his project, V-JEPA, is the best
He keeps saying stuff about "sucks as a mental model" but doesn't say why that would not apply to text.
https://twitter.com/ylecun/with_replies
Me: If text doesn't need a mental model, I see no reason video needs it. His argument sucks or is badly worded.
Re: Let me clear a huge misunderstanding
#27Re: Let me clear a huge misunderstanding
#28I think one of the things that needs to be said about Sora is the videos that examples that we've seen have been impressive, but this is also how most of OpenAI's examples have been on marketing pages. What matters is when you yourself are actually able to use it. In the past, it was only when I've actually been able to try the tech myself that I've been able to see how successful (or unsuccessful) they really are fo…
HN gets hung up on damning things that aren't perfect _right now_
You also see this with FSD...it isn't perfect today, so HN writes it off forever
Sora is a demo and a teaser of what will be a useful polished tool in three years, that's all
Re: Let me clear a huge misunderstanding
#29I feel LeCun got roped in debating the likes of Marcus and Yudkowsky. This has made his arguments lose nuance and become rigid. I also can't escape the feeling that if Facebook was tuned into Transformers, they would have shipped earlier, so there must have been some resistance or underestimation that's now repeated "They can't reason", "They can't plan", "They can't understand the world", "They are a distraction / s…
> long-term scene coherence, FWIW none of the video models released so far demonstrate any object coherence whatsoever, which suggests they don't have the higher level capabilities you mention yet. In Sora, as soon as an object is obstructed by an obstacle or goes offscreen, it's likely to disappear or be radically transformed.
Re: Let me clear a huge misunderstanding
#30"Furthermore, generating those continuations would be not only expensive but totally pointless." Why would it be pointless? I think there are many creative uses of video continuation model, considering the amount of control it gives you.
If we cut all the videos in the world in half, half of the videos will be continuations. Or another way of saying, all video is continuation, so I'm with you, not sure why it would be pointless. In fact, a continuation based workflow with prompting seems like the easiest way to get a specific effect.