Live data from Hacker News

Ask HN: Any insider takes on Yann LeCun's push against current architectures?

news.ycombinator.com

331–340 of 343 posts

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#331

Earlier quoted context omitted.

>Humans have the ability to admit when they do not know something. No not really. It's not even rare that a human confidently says and believes something and really has no idea what he/she's talking about. >We say “sorry, I don’t know, let me get back to you.” LLMs cannot do this Yeah they can. And they can do it much better than chance. They just don't do it as well as humans. >And they do not even know which one th…

No not really. It's not even rare that a human confidently says and believes something and really has no idea what he/she's talking about. Like you’re doing right now? People say “I don’t know” all the time. Especially children. That people also exaggerate, bluff, and outright lie is not proof that people don’t have this ability. When people are put in situations where they will be shamed or suffer other social stigm…

>Like you’re doing right now?

Lol Okay

>When people are put in situations where they will be shamed or suffer other social stigmas for admitting ignorance then we can expect them to be less than candid.

Good thing I wasn't talking about that. There's a lot of evidence that human explanations are regularly post-hoc rationalizations they fully believe in. They're not lieing to anyone, they just fully believe the nonsense their brain has concocted.

Experiments on choice and preferences https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/

Split Brain Experiments https://www.nature.com/articles/483260a

>As for your links to research showing that LLMs do possess the ability of introspection, I have one question: why have we not seen this in consumer-facing tools? Are the LLMs afraid of social stigma?

Maybe read any of them ? If you weren't interested in evidence to the contrary of your points then you could have just said so and I wouldn't have wasted my time. The 1st and 6th Links make it quite clear current post-training processes hurt calibration a lot.

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#332

Earlier quoted context omitted.

LLM is just the name. You can encode anything into the "language" including pictures video and sound.

I've always been wondering if anyone is working on using nerve impulses. My first thought when transformers came around was if they could be used for prosthetics, but I've been too lazy to do the research to find anybody working on anything like that, or to experiment myself with it.

Like Cortical labs? Neurons integrated on a silicon chip https://corticallabs.com/cl1.html

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#333

Earlier quoted context omitted.

There are absolutely reasons that we cannot capture the entirety—or even a proper image—of human cognition in semantic space. Cognition is not purely semantic. It is dynamic, embodied, socially distributed, culturally extended, and conscious. LLMs are great semantic heuristic machines. But they don't even have access to those other components.

The LLM embeddings for a token cover much more than semantics. There is a reason a single token embedding dimension is so large. You are conflating the embedding layer in an LLM and an embedding model for semantic search.

I don't think we're using the term semantic in the same way. I mean "relating to meaning in language."

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#334

Earlier quoted context omitted.

The LLM embeddings for a token cover much more than semantics. There is a reason a single token embedding dimension is so large. You are conflating the embedding layer in an LLM and an embedding model for semantic search.

I don't think we're using the term semantic in the same way. I mean "relating to meaning in language."

The embedding layer in an llm deals with much more than the meaning. It has to capture syntax, grammar, morphology, style and sentiment cues, phonetic and orthographic relationships and 500 other things that humans can't even reason about but exist in words combinations.

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#335
post #293

Earlier quoted context omitted.

>I think it's more fundamental than that. If you start saying "it thinks" in regards to an LLM, you're wrong. LLMs don't think, they pattern match fuzzily. I'm not sure if this objection is terribly helpful. We use terms like think and want to describe processes that are clearly not involve any form of understanding. Electrons do not have motivations but they 'want' to go to a lower energy level in an atom. You can h…

> We use terms like think and want to describe processes that are clearly not involve any form of understanding. ...and that's why so many people are confused about what's going on with LLMs: sloppy, ambiguous use of language. > In the Karpathy video I mentioned he talks about how researches found that models did have an internal representation of not knowing, but that the fine tuning was restricting it to providing…

>...and that's why so many people are confused about what's going on with LLMs: sloppy, ambiguous use of language.

There is a difference between explanation by metaphor and lack of precision. If you think someone is implying something literal when they might be using a metaphor you can always ask for clarification. I know plenty of people that are utterly precise in their use in their language which leads them to being widely misunderstood because they think a weak precise signal is received as clearly as a strong imprecise signal. They usually think the failure in communication is in the recipient but in reality they are just accurately using the wrong protocol.

>Do you understand the distinction I'm making here? I believe I do, and it is precisely this distinction that the researches showed. By teaching a model to say "I don't know" for some information that they knew the model did not know the answer to, the model learned to respond "I don't know" for things that it did not know that it was not explicitly taught to respond with "I don't know". For it to acquire that ability to generalise to new cases the model has to have already had an internal representation of "That information is not available"

I'm not sure where you think a model converting its internal representation of not knowing something into words is distinct from a human converting its internal representation of not knowing into words.

When fine tuning directs a model to profess lack of knowledge, usually they will not give the same specific "I don't know" text as a way to express that it does not not know because they want the want to bind the concept "lack of knowledge" to the concept of "communicate that I do not know" rather than any particular word phrase. Giving it many ways to say "I don't know" builds that binding rather than the crude "if X then emit Y" that you imagine it to be.

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#336

Earlier quoted context omitted.

Inference is already wasteful (compared to humans) but training is absurd. There's strong reason to believe we can do better (even prior to having figured out how).

That would mean with current resources AI can get so much more intelligent than humans, right? Aren't you scared?

That's a potential outcome of any increase in training efficiency.

Which we should expect, even from prior experience with any other AI breakthrough, where first we learn to do it and then we learn to do it efficiently.

E.g. Deep Blue in 1997 was IBM showing off a supercomputer, more than it was any kind of reasonably efficient algorithm, but those came over the next 20-30 years.

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#338
post #180

Earlier quoted context omitted.

I'm not sure I buy that, I didnt find the counter argument persuasive, but this comment basically took you from thoughtful to smug — unfairly so, ironically, because I've been so bored by not understanding Yann's "average housecat is smarter than an LLM" Speaking of which...I'm glad you're here ,because I have an interlocutor I can be honest with while getting at the root question of the Ask HN. What in the world doe…

It is easy to cross wires in a HN thread. I think what makes this discussion hard (hell it would be a hard PhD topic!) is: What do we mean by smart? Intelligent? Etc. What is my agenda and what is yours? What are we really asking? I won't make any more arguments but pose these questions. Not for you to answer but everyone to think about: Given (assuming) mammals including us have evolved and developed thought and lan…

a 3yr old is actually far more similar to AI than an adult. 3 year olds have extremely limited context windows. They will almost immediately forget what happened even 20-30 seconds ago when you play a game like memory with them, and they rarely remember what they ate for breakfast or lunch or basically any previous event from the same day.

When a 3 year old says "I love you" it is not at all clear that they understand what that means. They frequently mimic phrases they hear/basically statistical next word guessing and obviously don't understand the meaning of what they are saying.

You can even mimic an inner voice for them like Deepseek does for thinking through a problem with a 3 year old and it massively helps them to solve problems.

AI largely acts like a 3 year old with a massive corpus of text floating around in their head compared to the much smaller corpus a 3 year old has.

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#339

Okay I think I qualify. I'll bite. LeCun's argument is this: 1) You can't learn an accurate world model just from text. 2) Multimodal learning (vision, language, etc) and interaction with the environment is crucial for true learning. He and people like Hinton and Bengio have been saying for a while that there are tasks that mice can understand that an AI can't. And that even have mouse-level intelligence will be a br…

Is that what he's arguing? My perspective on what he's arguing is that LLMs effectively rely on a probabilistic approach to the next token based on the previous. When they're wrong, which the technology all but ensures will happen with some significant degree of frequency, you get cascading errors. It's like in science where we all build upon the shoulders of giants, but if it turns out that one of those shoulders wa…

Right now, humans still have enough practice thinking to point out the errors, but what happens when humanity becomes increasingly dependent on LLMs to do this thinking?

Re: Ask HN: Any insider takes on Yann LeCun's push against current architectures?

#340

Earlier quoted context omitted.

I don't think we're using the term semantic in the same way. I mean "relating to meaning in language."

The embedding layer in an llm deals with much more than the meaning. It has to capture syntax, grammar, morphology, style and sentiment cues, phonetic and orthographic relationships and 500 other things that humans can't even reason about but exist in words combinations.

I'll give you that. I was including those in "semantic space," but the distinction is fair.

My original point still stands: the space you've described cannot capture a full image of human cognition.

Post reply on HN