This is important, even though the Twitter part is useless. Here's the paper, "Revisiting Feature Prediction for Learning Visual Representations from Video"[1] A big problem with machine learning so far has been the lack of some underlying, more abstract model of the subject matter. I don't understand much of this, but apparently something called "representation", a purely mathematical concept, is able to help with t…
Let me clear a huge misunderstanding
11–20 of 109 posts
Re: Let me clear a huge misunderstanding
#12Why would it be pointless? I think there are many creative uses of video continuation model, considering the amount of control it gives you.
Re: Let me clear a huge misunderstanding
#13[flagged]
Re: Let me clear a huge misunderstanding
#14More like: "Let me clear a huge misunderstanding here by using poorly-defined terms and cramming niche complex ideas elaborated elsewhere into this tweet." Someone correct me, but it seems like he's saying: 1- generative models don't understand the real world 2- generative models that work off of just pixels are more expensive and less useful than a model that represents the contents of the frame with abstract repres…
Re: Let me clear a huge misunderstanding
#15"Furthermore, generating those continuations would be not only expensive but totally pointless." Why would it be pointless? I think there are many creative uses of video continuation model, considering the amount of control it gives you.
Re: Let me clear a huge misunderstanding
#16More like: "Let me clear a huge misunderstanding here by using poorly-defined terms and cramming niche complex ideas elaborated elsewhere into this tweet." Someone correct me, but it seems like he's saying: 1- generative models don't understand the real world 2- generative models that work off of just pixels are more expensive and less useful than a model that represents the contents of the frame with abstract repres…
I'm nowhere near an expert but it seemed like he was claiming true understanding of latent space is needed for generating coherent continuations, but Sora demo already has longish videos that are coherent. It's hard for me not to think this is someone just trying to still be right when they are wrong, but I may misunderstand.
Re: Let me clear a huge misunderstanding
#17Just seeing an error message, and cannot sign in as well. Am I the only one?
Re: Let me clear a huge misunderstanding
#18"Furthermore, generating those continuations would be not only expensive but totally pointless." Why would it be pointless? I think there are many creative uses of video continuation model, considering the amount of control it gives you.
Re: Let me clear a huge misunderstanding
#19So it's strange to me to claim that it doesn't have an abstract representation.
But, maybe the latent space of a diffusion-VQVAE pipeline is fundamentally different from that of JEPA, I haven't read the relevant papers for that. Curious if someone could explain if they are different ideas of representation.
Re: Let me clear a huge misunderstanding
#20It is kind of ironic that researchers who claim LLMs lack adaptive intelligence seemingly refuse to adapt their intelligence to LLMs. If even GPT-3 can find logical holes or oversimplification in your arguments about GPTs, at one point this starts becoming embarrassing and unbecoming.
> The generation of mostly realistic-looking videos from prompts does not indicate that a system understands the physical world.
While arguably true, it also does not indicate that a system does not understand the physical world (reflections, collision detection, gravity, object permanence, long-term scene coherence, etc.).
If LeCun wants to argue it does not understand the physical world, he should do so directly. Not attack something that is not directly stated, but rather convincingly and tentatively demo'd (I myself find it hard to argue that a system that generates novel pond reflections has not memorized/stored in weights some generalization program to apply to realistic scene generation).
This demo shows it is not even a wild prediction to guess that soon (consumer tech) we will be able to discuss visual scenes with conversational AIs.