Okay I think I qualify. I'll bite. LeCun's argument is this: 1) You can't learn an accurate world model just from text. 2) Multimodal learning (vision, language, etc) and interaction with the environment is crucial for true learning. He and people like Hinton and Bengio have been saying for a while that there are tasks that mice can understand that an AI can't. And that even have mouse-level intelligence will be a br…
where LeCun might be prescient should intersect with the nemesis SCHMIDHUBER. They can't both be wrong, I suppose?!
It's only "tangentially" related to energy minimization, technically speaking :) connection to multimodalities is spot-on.
https://www.mdpi.com/1099-4300/26/3/252
To Compress or Not to Compress—Self-Supervised Learning and Information Theory: A Review
With Ravid, double-handedly blue-flag MDPI!
Sunmarized for the layman (propaganda?) https://archive.is/https://nyudatascience.medium.com/how-sho...
>When asked about practical applications and areas where these insights might be immediately used, Shwartz-Ziv highlighted the potential in multi-modalities and tabula
Imho, best take I've seen on this thread (irony: literal energy minimization) https://news.ycombinator.com/item?id=43367126
Of course, this would make Google/OpenAI/DeepSeek wrong by two whole levels (both architecturally and conceptually)