Live data from Hacker News

Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

newsweek.com

101–110 of 147 posts

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#101

Earlier quoted context omitted.

The list of great minds who thought that "new fangled thing is nonsense" and later turned out to be horribly wrong is quite long and distinguished

> Heavier-than-air flying machines are impossible. -Lord Kelvin. 1895 > I think there is a world market for maybe five computers. Thomas Watson, IBM. 1943 > On talking films: “They’ll never last.” -Charlie Chaplin. > This ‘telephone’ has too many shortcomings… -William Orton, Western Union. 1876 > Television won’t be able to hold any market -Darryl Zanuck, 20th Century Fox. 1946 > Louis Pasteur’s theory of germs is r…

I'm pretty sure that Lord Kelvin was also in the cohort of fools that bullied Boltzmann to his suicide.

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#102

Earlier quoted context omitted.

The list of great minds who thought that "new fangled thing is nonsense" and later turned out to be horribly wrong is quite long and distinguished

> Heavier-than-air flying machines are impossible. -Lord Kelvin. 1895 > I think there is a world market for maybe five computers. Thomas Watson, IBM. 1943 > On talking films: “They’ll never last.” -Charlie Chaplin. > This ‘telephone’ has too many shortcomings… -William Orton, Western Union. 1876 > Television won’t be able to hold any market -Darryl Zanuck, 20th Century Fox. 1946 > Louis Pasteur’s theory of germs is r…

An important number of those remarks were based on a snapshot of the state of the technology: a fault in not seeing the potential evolution.

Examples of people who could not see non (in some way) dead-ends do not cancel examples of people who correctly saw dead-ends. The lists may even overlap ("if it remains that way it's a dead-end").

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#103
post #93

Earlier quoted context omitted.

> Heavier-than-air flying machines are impossible. -Lord Kelvin. 1895 > I think there is a world market for maybe five computers. Thomas Watson, IBM. 1943 > On talking films: “They’ll never last.” -Charlie Chaplin. > This ‘telephone’ has too many shortcomings… -William Orton, Western Union. 1876 > Television won’t be able to hold any market -Darryl Zanuck, 20th Century Fox. 1946 > Louis Pasteur’s theory of germs is r…

I am just wondering did you have this all somehow saved up or did you pull it out of somewhere? Amazing list of things. Thank You.

Gosh no. I knew most of that list but I'll be honest and tell you that I used ChatGPT to come up with it. it's a collection of quotes to begin with so I think that's okay. I'm not passing off someone else's writing as my own, I'm explicitly quoting them.

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#104
post #89
post #23

As LLMs do things thought to be impossible before, LeCun adjusts his statements about LLMs, but at the same time his credibility goes lower and lower. He started saying that LLMs were just predicting words using a probabilistic model, like a better Markov Chain, basically. It was already pretty clear that this was not the case as even GPT3 could do summarization well enough, and there is no probabilistic link between…

I wanna believe everything you say (because you generally are a credible person) but a few things don't add up: 1. Weakest ever LLM? This one is really making me scratch my head. For a period of time Llama was considered to THE best. Furthermore, it's the third most used on OpenRouter (in the past month): https://openrouter.ai/rankings?view=month 2. Ignoring DeepSeek for a moment, Llama 2 and 3 require a special lice…

Doesn't OpenRouter ranking include pricing?

Not really a good measure of quality or performance but of cost effectiveness

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#105

I remember reading Douglas Hofstadter's Fluid Concepts and Creative Analogies [ https://en.wikipedia.org/wiki/Fluid_Concepts_and_Creative_An... ] He wrote about Copycat, a program for understanding analogies ("abc is to 123 as cba is to ???"). The program worked at the symbolic level, in the sense that it hard-coded a network of relationships between words and characters. I wonder how close he was to "inventing" an L…

What is Dark Matter? How to eradicate cancer? How to have world peace? I don't quite see how pattern-matching, alone, can solve questions like these.

So, how do we solve questions like these? How about collecting a lot of data and looking for patterns in that data? In the process, scientists typically produce some hypotheses, test them by collecting more data and finding more patterns, and try to correlate these patterns with some patterns in existing knowledge. Do you agree?

If yes, it seems to me that LLMs should be much better at that than humans, and I believe the frontier models like o3 might already be better than humans, we are just starting to use them for these tasks. Give it a couple more years before making any conclusions.

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#106
post #77
post #66

Earlier quoted context omitted.

I don't understand. Deterministic and stochastic have very specific meanings. The statement: "To continue my reply I could say this word, more than the others, or maybe that one, a bit less, ..." sounds very much like a probability distribution.

If you really want to think at it as a probability, think at it as "the probability to express correctly the sentence/idea that was modeled in the activations of the model for that token". Which is totally different than "the probability that this sentence continues in a given way", as the latter is like "how in general this sentence continues", but instead the model picks tokens based on what it is modeling in the l…

That's not quite how auto-regressive models are trained (the expression of "ideas" bit). There is no notion of "ideas." Words are not defined like we humans do, they're only related.

And on the latent space bit, it's also true for classical models, and the basic idea behind any pattern recognition or dimensionality reduction. That doesn't mean it's necessarily "getting the right idea."

Again, I don't want to "think of it as a probability." I'm saying what you're describing is a probability distribution. Do you have a citation for "probability to express correctly the sentence/idea" bit? Because just having a latent space is no implication of representing an idea.

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#107
post #91

Earlier quoted context omitted.

The combinatorics on choosing 500 pieces (words) out of a bag of 1.8 million pieces (approx parameters per layer for GPT-3) with replacement, and order matters works out to be something like 10^4600. Maybe you can't call that creativity, but you've got to admit that's a pretty big number.

I said No handwaving with scale. :-)

Right—but why should “new ABS plastic” be the bar for creativity? If the kid builds a structure no one’s ever imagined, from an unimaginably large box of Lego, isn’t that still novel? Sure, it’s made from known parts—but so is language. So are most human ideas.

The demand for outputs that are provably untraceable to training data feels like asking for magic, not creativity. Even Gödel didn’t require “never seen before atoms” to demonstrate emergence.

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#108
post #96
post #42

Earlier quoted context omitted.

Because he has a core belief and based on that core belief he made some statements that turned out to be incorrect. But he kept the core belief and adjusted the statements. So it's not so much about his incorrect predictions, but that these predictions were based on a core belief. And when the predictions turned out to be false, he didn't adjust his core beliefs, but just his predictions. So it's natural to ask, if n…

I have not followed all of LeCun's past statements, but - if the "core belief" is that the LLM architecture cannot be the way to AGI, that is more of an "educated bet", which does not get falsified when LLMs improve but still suggest their initial faults. If seeing that LLMs seem constrained in the "reactive system" as opposed to a sought "deliberative system" (or others would say "intuitive" vs "procedural" etc.) wa…

If you say LLMs are a dead end, and you give a few examples of things they will never be able to do, and a few months later they do it, and you just respond by stating that sure they can do that but they're still a dead end and won't be able to do this.

Rinse and repeat.

After a while you question whether LLMs are actually a dead end

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#109

I remember reading Douglas Hofstadter's Fluid Concepts and Creative Analogies [ https://en.wikipedia.org/wiki/Fluid_Concepts_and_Creative_An... ] He wrote about Copycat, a program for understanding analogies ("abc is to 123 as cba is to ???"). The program worked at the symbolic level, in the sense that it hard-coded a network of relationships between words and characters. I wonder how close he was to "inventing" an L…

What is Dark Matter? How to eradicate cancer? How to have world peace? I don't quite see how pattern-matching, alone, can solve questions like these.

My premise is that pattern-matching unlocks human-level artificial intelligence. Just because LLMs haven't cured cancer yet doesn't mean that LLMs will never be as intelligent as humans. After all, humans haven't cured cancer yet either.

What is intelligence?

Is it reacting to the environment? No, a thermostat can do that.

Is being logical? No, the simplest program can do that.

Is it creating something never seen before? No, a random number generator can do that.

We can even combine all of the above into a program and it still wouldn't be intelligent or creative. So what's the missing piece? The missing piece is pattern-matching.

Pattern-matching is taking a concrete input (a series of numbers or a video stream) and extracting abstract concepts and relationships. We can even nest patterns: we can match a pattern of concepts, each of which is composed of sub-patterns, and so on.

Creativity is just pattern matching the output of a pseudo-random generator against a critique pattern (is this output good?). When an artist creates something, they are constantly pattern matching against their own internal critic and the existing art out there. They are trying to find something that matches the beauty/impact of the art they've seen, while matching their own aesthetic, and not reproducing an existing pattern. It's pattern-matching all the way down!

Science is just a special form of creativity. You are trying to create a model that reproduces experimental outcomes. How do you do that? You absorb the existing models and experiments (which involves pattern-matching to compress into abstract concepts), and then you generate new models that fit the data.

Pattern-matching unlocks AI, which is why LLMs have been so successful. Obviously, you still need logic, inference, etc., but that's the easy part. Pattern-matching was the last missing piece!

Re: Yann LeCun, Pioneer of AI, Thinks Today's LLM's Are Nearly Obsolete

#110
post #23

As LLMs do things thought to be impossible before, LeCun adjusts his statements about LLMs, but at the same time his credibility goes lower and lower. He started saying that LLMs were just predicting words using a probabilistic model, like a better Markov Chain, basically. It was already pretty clear that this was not the case as even GPT3 could do summarization well enough, and there is no probabilistic link between…

LLMs literally are just predicting tokens with a probabilistic model. They’re incredibly complicated and sophisticated models, but they still are just incredibly complicated and sophisticated models for predicting tokens. It’s maybe unexpected that such a thing can do summarization, but it demonstrably can.

The rub is that we don't know if intelligence is anything more than "just predicting next output".
Post reply on HN