Live data from Hacker News

Yann LeCun on GPT-3

facebook.com

201–210 of 253 posts

Re: Yann LeCun on GPT-3

#201

Text reproduced, minus abusive shell of dark patterns: Some people have completely unrealistic expectations about what large-scale language models such as GPT-3 can do. This simple explanatory study by my friends at Nabla debunks some of those expectations for people who think massive language models can be used in healthcare. GPT-3 is a language model, which means that you feed it a text and ask it to predict the co…

this comment reads like it was generated by GPT-3. sum up your points.

Re: Yann LeCun on GPT-3

#202

Reading this is really interesting: > GPT-3 doesn't have any knowledge of how the world actually works. I think this is a philosophical question. There is a view that, basically, there is no such thing as knowledge, just language (or, at least, there is no distinction between knowledge and language). In this view, all there really is is language, which is mostly composed of metaphors and, ultimately, metaphors only r…

If GPT-3 has a consistent position on anything, it's only because the corpus it was trained on was consistent about it. So, for example, it will reliably autocomplete Jabberwocky because there are a lot of copies of this poem in the corpus and they are all the same.

If there were two versions of this poem that started the same way, it would pick between the variations in the corpus randomly. In other cases it might choose based on the style of prose or other stuff like that.

GPT-3 can get some trivia right, but it's only because the editors of Wikipedia already came to consensus about it and Wikipedia was weighted more. It doesn't have a way of coming to a consistent conclusion on its own.

Without consistency, how can it be said to know or believe anything? You might as well ask what a library believes. Sure, the authors may have believed things, but it depends which book you happen to pick up.

Re: Yann LeCun on GPT-3

#203

> Some people have completely unrealistic expectations about what large-scale language models such as GPT-3 can do. Just want to point out that he's saying the people on the upper end of the expectation distribution are wrong, not the people in the middle of it. So if you're takeaway from this is that GPT3 is nothing special, that's probably the wrong message.

His next paragraph claims that Nabla "debunks" the idea that "large language models" can be used in healthcare. That's not just "some people have unrealistic expectations" it's "this tool, when when more advanced and find tuned, will never be appropriate to use in a very broad class of use cases". He also says "GPT-3 has no knowledge of how the world works", which is clearly an overstatement meant to clear up hype, b…

>For example, GPT-3 knows more trivia than I do.

no it doesn't, GPT-3 is a very sophisticated parrot. it doesn't know any trivia, it knows how to put the most likely string of characters next to the one it just saw, it doesn't matter what the text represents. That's the difference between you and the model.

It's basically the Chinese room. You can make an analog GPT-3 by asking a question, recording your answer, handing someone who doesn't understand a word of your language the giant box of tapes, and she tries to match them together until she appears to make sense to listeners

Re: Yann LeCun on GPT-3

#205

Earlier quoted context omitted.

I'm not sure that you can assert that language is tightly coupled to reality, unless you're using the term "reality" to mean something akin to "as one perceives the world" (regardless of whether that perception is correct or not). Most expressions of language that survived from a few thousand years ago are centered around myths, and while those myths may have contained certain moral or ethical lessons (that were and…

a few things: 1. i get where you're coming from 2. yes, language is bottlenecked by human perception, as are all things 3. even the notion of myth and fiction is encoded in language. language is self-descriptive and self-aware and you can separate sense from non-sense. 4. i'm not talking about knowledge or understanding, but of addressing the question of why training on language let's GPT-3 make human-like prediction…

Why does it make human-like predictions as if it knows about reality? Because it's essentially pattern-matching the consensus of the literature it was trained on. Literature as a whole will tend to settle on a consensus sentiment, albeit one that is probably significantly behind the current consensus sentiment (it takes time to accumulate enough mass to move the weights). If your interactions with GPT-3 fall into the rather sizeable consensus that most people either subscribe to or are familiar with then it will certainly prove a decent mimic of understanding. If, however, you attempt to teach it a novel concept or if you dig into its interactions long enough to test for depth of understanding you run into the gaps and GPT-3 either begins to mimic that bullshit artist everyone knows that claims to know things but is only regurgitating platitudes and buzzwords or it begins to mimic behavior that would make you question whether it was sober and/or sane.

There's little doubt that GPT-3 could hold its own quite well in a bout of polite conversation and/or small talk that features in many social situations but that's more a commentary on the limited area of knowledge and behavior that etiquette expects for interactions in such settings.

It could also be trained on the canon of Shakespeare and behave as a prior work of Shakespeare, but if you left one play out of that canon and then attempted to have GPT-3 generate that missing work it wouldn't... at all. It doesn't approximate how Shakespeare thought, it approximates the literature he produced that you used for training.

None of this is to denigrate the achievement that GPT-3 represents, it is merely to point out that it is unfair to GPT-3 to attempt to hold it to the standards of AGI.

Re: Yann LeCun on GPT-3

#206
post #79

Earlier quoted context omitted.

I'm not reducing GPT-3 to the extent that you're suggesting. I'm pointing out (and so does LeCun in his post) that it's a language model designed to continue a sequence of words. It has no understanding of the world and is no particularly suited for knowledge extraction or conversation. > GPT-3 is a really impressive milestone towards AGI We really don't know this. It's a big step for the field of language models, th…

> it's a language model designed to continue a sequence of words. If a language model were able to do this task perfectly, it would be indistinguishable from intelligence, because continuing a sequence of words requires reasoning. You cannot conclude that has no understanding based solely on what it is trained to do when the task it is trained on would be sufficient to demonstrate understanding were it to fully succe…

> you cannot conclude that [a model] has no understanding based solely on what it is trained to do

agreed.

but that's not everything we're basing our conclusions on – we also know that GPT-3 was trained purely on text, and i (and presumably GP) don't think that's a path towards "understanding".

in other words, i think being a language model [trained only using a text corpus] is a valid reason to be skeptical of its potential :)

Re: Yann LeCun on GPT-3

#207

Earlier quoted context omitted.

No. I just typed it on a mobile device. (but maybe that's exactly what a bot would say...) EDIT: Actually - that's no excuse for that awful second sentence. I'm ashamed of myself.

Actually - why would a bot be more likely to make typos and grammatical errors? Surely a slightly careless human is the simpler explanation?

Haha so I was actually thinking about that myself after submitting. Thanks for calling me out -- I guess my only answer to that would be, the bot (by that I meant GPT-3) is using the wrong conjugation of a verb it learned somewhere.

Re: Yann LeCun on GPT-3

#208

Earlier quoted context omitted.

Um, as an outside observer, what is Open about this OpenAI GPT-3 then if they’re selling exclusive rights?

They were forced to give Microsoft exclusive access, because it was one of the terms of Microsoft's billion-dollar cloud credit investment. But you can't pay employees with cloud credits, so time will tell whether it was a correct decision. (It probably was. And I exaggerate slightly; the investment included a substantial sum of real dollars too. But most people see that billion dollar investment and think it's all d…

Is that only for GPT-3? Or for everything they produce?

Re: Yann LeCun on GPT-3

#209
post #2

It's nice to hear from someone who knows what they're talking about that GPT-3 is just a fancy and expensive autocomplete. The hype in some circles about it went as far as comparing it to AGI at some point which is just ridiculous.

You are just a fancy and efficient autocomplete too. When you speak or write, some words have a higher probability than others. You pick alternatives, but they are limited. Of course there are more layers in the human mind, but GPT-3 is a really impressive milestone towards AGI. It's so easy to downplay every advanced tech, it's actually fun. Planes? Just a flying metal tube. Self landing rockets? Just applied physic…

>You are just a fancy and efficient autocomplete too. When you speak or write, some words have a higher probability than others

No you're not, and that is very easy to disprove. Look at the sentence "John took the water bottle out of the backpack so that it would be lighter". What does it refer to in the sentence, the bottle or the backpack?

Did it statistically come to you or did you need to consult Google? No, you know the right answer, it's the backpack. Why? Because you have a physical understanding of the world. The bottle doesn't get lighter, when you take it out of the backpack, the backpack does, because the bottle is not in there any more. This is not statistics, it's not manipulating strings, it's having a fundamental physical model of the world in your head, and an idea about how entities operate in it.

When you talk you don't do random statistical inference, you match language to the semantics you want to express, which is not statistical.

Re: Yann LeCun on GPT-3

#210

Reading this is really interesting: > GPT-3 doesn't have any knowledge of how the world actually works. I think this is a philosophical question. There is a view that, basically, there is no such thing as knowledge, just language (or, at least, there is no distinction between knowledge and language). In this view, all there really is is language, which is mostly composed of metaphors and, ultimately, metaphors only r…

> I think this is a philosophical question.

If it's only philosophical, then me saying that Hacker News Website itself has 'knowledge' of everything we discuss about is also philosophical. Same can be applied to plain paper books.

How about any web application? A for loop? anything which can generate something for you?

Post reply on HN