Live data from Hacker News

Let me clear a huge misunderstanding

twitter.com

11–20 of 109 posts

Re: Let me clear a huge misunderstanding

#11
post #2

This is important, even though the Twitter part is useless. Here's the paper, "Revisiting Feature Prediction for Learning Visual Representations from Video"[1] A big problem with machine learning so far has been the lack of some underlying, more abstract model of the subject matter. I don't understand much of this, but apparently something called "representation", a purely mathematical concept, is able to help with t…

https://en.wikipedia.org/wiki/Feature_learning

Re: Let me clear a huge misunderstanding

#14
post #9

More like: "Let me clear a huge misunderstanding here by using poorly-defined terms and cramming niche complex ideas elaborated elsewhere into this tweet." Someone correct me, but it seems like he's saying: 1- generative models don't understand the real world 2- generative models that work off of just pixels are more expensive and less useful than a model that represents the contents of the frame with abstract repres…

I'm nowhere near an expert but it seemed like he was claiming true understanding of latent space is needed for generating coherent continuations, but Sora demo already has longish videos that are coherent. It's hard for me not to think this is someone just trying to still be right when they are wrong, but I may misunderstand.

Re: Let me clear a huge misunderstanding

#15

"Furthermore, generating those continuations would be not only expensive but totally pointless." Why would it be pointless? I think there are many creative uses of video continuation model, considering the amount of control it gives you.

If we cut all the videos in the world in half, half of the videos will be continuations. Or another way of saying, all video is continuation, so I'm with you, not sure why it would be pointless. In fact, a continuation based workflow with prompting seems like the easiest way to get a specific effect.

Re: Let me clear a huge misunderstanding

#16
post #14
post #9

More like: "Let me clear a huge misunderstanding here by using poorly-defined terms and cramming niche complex ideas elaborated elsewhere into this tweet." Someone correct me, but it seems like he's saying: 1- generative models don't understand the real world 2- generative models that work off of just pixels are more expensive and less useful than a model that represents the contents of the frame with abstract repres…

I'm nowhere near an expert but it seemed like he was claiming true understanding of latent space is needed for generating coherent continuations, but Sora demo already has longish videos that are coherent. It's hard for me not to think this is someone just trying to still be right when they are wrong, but I may misunderstand.

The way i understand it is that Sora is mostly just 'moving pictures' with no rhyme or reason. Yann Lecun is interested in videos that tell a 'story', with cause and effect. Like a magician putting his hand into a top hat and pulling out a rabbit, kind of video.

Re: Let me clear a huge misunderstanding

#18

"Furthermore, generating those continuations would be not only expensive but totally pointless." Why would it be pointless? I think there are many creative uses of video continuation model, considering the amount of control it gives you.

It's not pointless. It can potentially generate world simulations. We might be inside one such long video.

Re: Let me clear a huge misunderstanding

#19
The main thing I'm confused about by these comments is that as far as I understand, the Sora model (and many others like it) performs the diffusion process in latent space, and then translates this to pixels.

So it's strange to me to claim that it doesn't have an abstract representation.

But, maybe the latent space of a diffusion-VQVAE pipeline is fundamentally different from that of JEPA, I haven't read the relevant papers for that. Curious if someone could explain if they are different ideas of representation.

Re: Let me clear a huge misunderstanding

#20
I feel LeCun got roped in debating the likes of Marcus and Yudkowsky. This has made his arguments lose nuance and become rigid. I also can't escape the feeling that if Facebook was tuned into Transformers, they would have shipped earlier, so there must have been some resistance or underestimation that's now repeated "They can't reason", "They can't plan", "They can't understand the world", "They are a distraction / side road to AGI".

It is kind of ironic that researchers who claim LLMs lack adaptive intelligence seemingly refuse to adapt their intelligence to LLMs. If even GPT-3 can find logical holes or oversimplification in your arguments about GPTs, at one point this starts becoming embarrassing and unbecoming.

> The generation of mostly realistic-looking videos from prompts does not indicate that a system understands the physical world.

While arguably true, it also does not indicate that a system does not understand the physical world (reflections, collision detection, gravity, object permanence, long-term scene coherence, etc.).

If LeCun wants to argue it does not understand the physical world, he should do so directly. Not attack something that is not directly stated, but rather convincingly and tentatively demo'd (I myself find it hard to argue that a system that generates novel pond reflections has not memorized/stored in weights some generalization program to apply to realistic scene generation).

This demo shows it is not even a wild prediction to guess that soon (consumer tech) we will be able to discuss visual scenes with conversational AIs.

Post reply on HN