Live data from Hacker News

GPT-3 has no idea what it’s talking about

technologyreview.com

201–210 of 323 posts

Re: GPT-3 has no idea what it’s talking about

#201

The article is of course right but also a bit silly. Language models like GPT-X are producing grammatically correct sentences, along the lines of "Colorless green ideas sleep furiously". The NLP research more or less solved the old syntax problem using 'distributional semantics' but 'semantics' is a misnomer, it's all about syntax. In fact the most useful part of the article for me is that they mentioned Douglas Summ…

Gradient free optimization is not used much in Neural Networks except in Reinforcement Learning. I think it's because backprop is objectively faster for most supervised problems than other techniques (e.g. simulated annealing or GAs)

The thing is, living beings don’t learn by brute-force trial and error like mathematical optimization models you mentioned. Besides enormous energy spent, an individual organism will just be eaten by predator on another iteration of its ‘error minimization’ loop.

the idea of learning via reinforcement, that came from Skinner behaviorist experiments has been long discredited in cognitive psychology. (I highly recommend Wayne Wickelgren’ work on learning and memory if you’re interested, it’s brilliant and concise http://www.columbia.edu/~nvg1/Wickelgren/ )

Biological plausibility might not be needed for recognizing check signatures or images of traffic lights, where backprop is working just fine, but I believe true cognition would require such energy expenditures that brute-force trial and error will never be feasible. Moreover such error correction imposes artificial constraints that limit the amount of information that can be learned, kind of like those mechanical calculators of the 17th century with gears and wheels and crude mechanical actuators.

Re: GPT-3 has no idea what it’s talking about

#202
post #183

Earlier quoted context omitted.

It's not about impressiveness - surely, it's impressive. However, the article is more or less critiquing the discourse surrounding the model - namely, that there is a strange misconception floating around that it's somehow a general purpose AI that can understand and think about the world similar to a human. Which, of course, it cannot. If the claims about GPT-3 were accurate, there'd be a lot less of a flare-up abou…

I fail to see where openai is making any false claims

That's a fundamental skill of marketing: not making false claims while convincing the customers to jump to false conclusions.

Re: GPT-3 has no idea what it’s talking about

#204

Earlier quoted context omitted.

~50-300 million sounds like an over-estimate, and the investment in GPT-3 is miniscule compared to SDC. Plus, GPT-3 generates value even without direct product impact. I'm sure MSFT sales reps are already folding tons of nonsense about gpt-3 into their Azure pitches. Industry labs like OpenAI and DeepMind replace/augment the "Research & Development" model with a "Research & Marketing" model.

Seems spot-on. One trick to estimate a startup's burn rate is to multiply their number of employees by $200k. It's not too accurate, but it's within the ballpark. So how many employees does OpenAI have? Supposing they have 500, that's a burn rate of $100M/yr. 250 employees, $50M/yr. 100 employees, $20M/yr.

Gpt-3 is just one of openai's many projects. Theres ~20 authors on the paper, and they all almost certainly did not spend a full year on this project. So 300m is completely wrong.

Re: GPT-3 has no idea what it’s talking about

#205
post #75

Why must we keep having this argument? If you do research in the field you know full well that GPT/any other transformer or Bert model is generating text by regurgitating approximate conditional probabilities of words given all the text it has ever seen and the prompt. The neurophysiological concept of “understanding” as most understand it is orthogonal to the way the algorithm actually works. A more useful conversat…

> The neurophysiological concept of “understanding” as most understand it is orthogonal to the way the algorithm actually works. This is not obviously true and it's exactly the core of the debate. A GPT-3 proponent might say: We don't really know what "understanding" means, so it very well might be nothing more than complex rehashing of conditional probabilities. This isn't implausible. Consider Friston's "free energ…

That’s a good point and thanks for the reference.

I added “as most understand it” to caveat cases like this one, where there exists a non-falsifiable theory about how cognition works under which the GPT algorithm and “understanding” would be non-orthogonal.

Don’t get me wrong, it’s an interesting theory, but with no evidence of existence or non-existence do we really need to spend this much time on it? This is why I invoked cults - arguing about theories without evidence smells a more like a religious argument than a scientific one.

I think I mostly just wish we could end the argument by all agreeing the following (I think) non-controversial points...

1) GPT is very impressive 2) GPT is not perfect 3) we don’t have a fucking clue how human cognition works 4) because of 3, how “close” GPT is to human cognition is an open question

Re: GPT-3 has no idea what it’s talking about

#206

Earlier quoted context omitted.

>Any tractable amount of data with the former [statistics] can’t approximate an ounce of the latter [causal/reasoning structure]. I don't know why you think this is true. If statistically B follows A to a high degree, then a sufficiently advanced statistical model will represent "A then B" in some manner. In a predictive language model, at some point the best way to model a text corpus that indirectly references the…

> I don't know why you think this is true. If statistically B follows A to a high degree, then a sufficiently advanced statistical model will represent "A then B" in some manner. Yes, but suppose A implied B only if C were true. And in the training corpus C were always true (hence learned A=>B) but in the test corpus suppose C is not true, then the learned statistical rule is wrong. The problem is that to cover all t…

But this isn't an issue for statistically modelling causal relationships specifically, this is a core problem of modelling causal relationships at all. The fact that GPT-3 is sensitive to the real world changing, or to having insufficient information to form a universally accurate model says nothing interesting about GPT-3.

Re: GPT-3 has no idea what it’s talking about

#207

Earlier quoted context omitted.

It’s easy to sit here and say what their goals “should” be, but it doesn’t change the fact that they never claimed it’s an AGI.

Gary Marcus doesn't claim that they claim that it's an AGI, so I really have no god damn clue what y'all are going on about here. Clearly, the point I'm making in the top comment is the following: OpenAI apparently restricts access to people -- even very highly respective scientists -- who happen to be critical of their previous work. It's impossible to prove intent, but limiting access for people like Gary Marcus is…

> Gary Marcus doesn't claim that they claim that it's an AGI, so I really have no god damn clue what y'all are going on about here.

Well, who is Gary Marcus anyway? I looked up his publications in the last 2 years on Google Scholar and I see no hard contributions, all he has are some publications about futurology and AI critique. His wiki page says he's a cognitive scientist who once sold an ML company to Uber, but not an AI expert.

Why didn't Gary invent a better language model to show us how it's done? If he knows better than the guys at OpenAI, let him show the path ahead. When someone has superior results it's not constructive to throw bullshit at their accomplishments.

Re: GPT-3 has no idea what it’s talking about

#208

Earlier quoted context omitted.

~50-300 million sounds like an over-estimate, and the investment in GPT-3 is miniscule compared to SDC. Plus, GPT-3 generates value even without direct product impact. I'm sure MSFT sales reps are already folding tons of nonsense about gpt-3 into their Azure pitches. Industry labs like OpenAI and DeepMind replace/augment the "Research & Development" model with a "Research & Marketing" model.

Seems spot-on. One trick to estimate a startup's burn rate is to multiply their number of employees by $200k. It's not too accurate, but it's within the ballpark. So how many employees does OpenAI have? Supposing they have 500, that's a burn rate of $100M/yr. 250 employees, $50M/yr. 100 employees, $20M/yr.

I believe significant share of GPT-3 cost is machine-hours that were spent training this model - months of hundreds top-tier NVidia machines.

Edit: estimates range from $2-$5MM to $15MM

https://www.reddit.com/r/MachineLearning/comments/hwfjej/d_t...

Re: GPT-3 has no idea what it’s talking about

#209
post #120

Earlier quoted context omitted.

"Evaluate the model properly"? VCs think this thing can code

I’m not implying it can’t! It might be able to in many cases, if you do prompt design right and fine-tune.

Sure, as in: if the spec is flawless, we can source out the coding to a bunch of minimum wage dudes in {location of your choice}. Anyone happy with this approach?

Re: GPT-3 has no idea what it’s talking about

#210
post #52
post #46

The authors don't understand prompt design well enough to evaluate the model properly. Take this example: Prompt: > You are a defense lawyer and you have to go to court today. Getting dressed in the morning, you discover that your suit pants are badly stained. However, your bathing suit is clean and very stylish. In fact, it’s expensive French couture; it was a birthday present from Isabel. Continuation: > You decide…

I think you're kind of proving the OPs point. The argument is that GPT3 has no understanding of the world, just superficial understanding of words and their relationships. If it did have a real understanding, prompt construction wouldn't matter as much, but it clearly does because all GPT3 cares about the structure of sentences, not their meanings.

GPT-3 is a statistical model of text sequences, it has just textual understanding of the world. But the funny thing is that it can do lots of tasks without explicit training, and that is something amazing, it shows a path forward. In order to have real understanding it needs to be an embodied agent that interacts with the world like us, and has goals and needs like us.
Post reply on HN