Live data from Hacker News

How New Are Yann LeCun's “New” Ideas?

garymarcus.substack.com

41–50 of 90 posts

Re: How New Are Yann LeCun's “New” Ideas?

#41
post #28

Is Marcus trying to create the impression that somehow he is a more impactful AI contributor than LeCun? It's going to be a tough sell because I know LeCun's name from his technical work whereas I know Marcus' name from him constantly moaning about LeCun on social media. In what _tangible_ ways did Marcus contribute?

The article answers this question in detail.

Gotta love when the question proves they didn't read what they're asking about.

Re: How New Are Yann LeCun's “New” Ideas?

#42
post #2

I expected this to be a smear / petty argument article. In fact, it's a concise, highly specific, quote by quote critique. I don't have enough context to take a side, but this is not just a rant. Beyond their interpersonal disagreements, I do wonder if LeCunn is seeing diminishing marginal returns to deep learning at FB...

Diminishing returns? Have you read the Gato, Palm, Stable Diffusion, etc. papers? Progress is racing ahead. Nothing is stalling... the only thing stopping progress from accelerating even faster is data.

He is talking about Deep learning at FB

Re: How New Are Yann LeCun's “New” Ideas?

#43

Earlier quoted context omitted.

So what? Is he actually right, or is he wrong? A good argument delivered badly is still a good argument.

>LeCun, 2022: Today's AI approaches will never lead to true intelligence (reported in the headline, not a verbatim quote); Marcus, 2018: “deep learning must be supplemented by other techniques if we are to reach artificial general intelligence.” If you think that is substantive evidence for a stolen idea then it's surely not possible for anyone to ever have an original thought.

> If you think that is substantive evidence for a stolen idea

He says in the article that he doesn't think it's a stolen idea:

"I won’t accuse LeCun of plagiarism, because I think he probably reached these conclusions honestly, after recognizing the failures of current architectures."

Re: How New Are Yann LeCun's “New” Ideas?

#44
post #5

Earlier quoted context omitted.

Any specific reason this is relevant to the arguments in his post?

What is he even arguing here? That he has been cheated out of some kind of credit? Credit for what? Afaict he has never actually shown something novel based on his ideas to work in a way that has mattered.

> What is he even arguing here? That he has been cheated out of some kind of credit?

He's (at least to some extent) arguing that if you're going to say someone's paper is "mostly wrong", saying the same things 4 years later should probably warrant a "ok, you were right" at least.

Re: How New Are Yann LeCun's “New” Ideas?

#45

Yann LeCun’s Facebook post from a few days ago now makes more sense to me: https://www.facebook.com/722677142/posts/pfbid035FWSEPuz8Yqe...

From the comments on that post, written by LeCun:

"'[...] Yann LeCun, [...] is on a mission to reposition himself, not just as a deep learning pioneer, but as that guy with new ideas about how to move past deep learning'

First, I'm not 'repositioning myself'. My position paper is in the direct line of things I (and others) have thought about, talked about, and written about for years, if not decades. Gary has merely crashed the party.

My position paper is not at all about 'moving past deep learning'. It's the opposite: using deep learning in new ways, with new DL architectures (JEPAs, latent variable models), and new learning paradigms (energy-based self-supervised learning).

It's not at all about sticking symbol manipulation on top of DL as he suggests in vague terms. It's about seeing reasoning as latent-variable inference based on (hopefully gradient-based) optimization.

Gary claims that my critiques of supervised learning, reinforcement learning, and LLMs (my 'ladders') are critiques of deep learning (his 'ladder'). But they are not. What's missing from SL, RL and LLM are SSL, predictive world models, joint-embedding (non generative) architectures, and latent-variable inference (my rockets). But deep learning is very much the foundation on which everything is built.

In my piece, reasoning is the minimization of an objective with respect to latent variables. If Gary wants to call this 'symbol manipulation' and declare victory, fine. But it's merely a question of vocabulary. It certainly is very much unlike any proposal he has ever made, despite the extreme vagueness of those proposals."

Re: How New Are Yann LeCun's “New” Ideas?

#46

Earlier quoted context omitted.

He is arguing that Yann LeCun is taking ideas from other researchers without citation or credit, and that this is a sign of insecurity and ego.

"Ideas" here being few commonly used words put together barely forming a sentence, not some algorithm, research or deep paper. Ie. the whole "idea" being "hey, for gai we need something different than this gtp3" tweet, not "idea" as in "hey I invented this new thing I call LSTM, check it out [link to paper, results what not]".

> hey I invented this new thing I call LSTM

Well, the LSTM fella does pop up later calling out Lecun for "rehashes but doesn't cite essential work of 1999-2015". Which I guess does mean people with real "ideas" are also fed up with him?

"Deep learning pioneer Jürgen Schmidhuber, author of the commercially ubiquitous LSTM neural network, arguably has even more right to be pissed [...]"

Re: How New Are Yann LeCun's “New” Ideas?

#47

None of Gary’s comments were original either. I don’t know what I’d call this, but I’ve seen similar behavior elsewhere. This weird “flag planting” behavior to try to get credit without doing any actual work, as well as disregarding all prior work. Normally the “predictions” are vague or could be applied to anything. It seems borderline like a mental illness of some sort, but I’m not a mental health professional.

There is a guy at work who paints sometimes presumptuous "# TODO: ..." comments all over the code base without ever actually doing anything about the issues, or discussing with anyone. Similar phenomena. Haha.

Re: How New Are Yann LeCun's “New” Ideas?

#48
This is frustrating:

Consider this:

LeCun, 2022: Today's AI approaches will never lead to true intelligence (reported in the headline, not a verbatim quote); Marcus, 2018: “deep learning must be supplemented by other techniques if we are to reach artificial general intelligence.”

How can that be something that LeCun did not give Marcus credit for? It is borderline self evident, and people have been saying similar things since neural networks were invented. This would only be news if LeCun had said that "neural nets are all you need" (literally, not as a reference to the title of the transformers paper).

And furthermore, if LeCun had said that, there are literally dozens of people who have also said that you need to combine the approaches.

He cites a single line:'LeCun spent part of his career bashing symbols; his collaborator Geoff Hinton even more so, Their jointly written 2015 review of deep learning ends by saying that they “new paradigms are needed to replace rule-based manipulation of symbolic expressions.”'

Well, sure because symbol processing alone is not the answer either. We need to replace it with some hybrid. How is this a contradiction?

To summarize: people have been looking for a productive way to combine symbolic and statistical systems -- there are in fact many such systems proposed with varying degrees of success. LeCun agrees with this approach (no one has anything to lose by endorsing adding things to any model), but Marcus insists he came up with it and he should be cited.

Ugh.

Re: How New Are Yann LeCun's “New” Ideas?

#49
So the idea is that statistical language modelling is not enough. You need a model based on logic too for "real" artificial intelligence. I wonder what the evidence for this claim is? Because the inferences and reasoning GPT3 is already capable of is incredible and beats most expert systems that I know of. And GPT4 is around the corner, Stable Diffusion was published like only a few months ago. I don't see why not more compute, more training data, and better network architectures couldn't lead to leaps and bounds of model improvements. At least for a few more years.
Post reply on HN