Live data from Hacker News

Gopher – A 280B parameter language model

deepmind.com

21–30 of 127 posts

Re: Gopher – A 280B parameter language model

#22
post #15

The closer we get to artificial intelligence, the more we raise the bar for what qualifies as AI (as we should). Gopher/GPT-3 are already much more accurate than the average human at technical information retrieval (trivial to see from the dialogue transcripts: how many Americans know what a Schwarzschild metric is?). The focus on ethics and equity for these algorithms is interesting too, as the average human holds m…

The problem seems to be that these models provide fairly accurate information at many occasions and occasionally complete blunders. Humans provide less accurate information most of the time but with a certain amount of self-reflection/meta-cognition, and they will usually recognize total blunders or display reasonable uncertainty about them.

There are only very few applications where it would make sense to take the risk and use an AI that occasionally makes gigantic mistakes without any understanding why. Even seemingly harmless applications like automated customer support could go horribly wrong.

Re: Gopher – A 280B parameter language model

#23
post #15

The closer we get to artificial intelligence, the more we raise the bar for what qualifies as AI (as we should). Gopher/GPT-3 are already much more accurate than the average human at technical information retrieval (trivial to see from the dialogue transcripts: how many Americans know what a Schwarzschild metric is?). The focus on ethics and equity for these algorithms is interesting too, as the average human holds m…

When those language models are wrong or biased, the user will have a worse experience in all three of those scenarios. At least when we look at search results now, we can prune for the facts. Those language models are ingesting that same data to give a monolithic answer your a query. Less transparent, less safe.

Re: Gopher – A 280B parameter language model

#24

It should have some uncertainty when it says there are no French-speaking countries in South America. French Guiana is there, but it's not clear it counts as a "country in South America" since it's part of France. Technically you could say France is (partially) a country in South America, and France definitely is French-speaking. The way the question is phrased is unclear as to whether French Guiana should count, and…

I think you're missing the point. That section was to show that the model is sometimes wrong and lacks the self-awareness to be uncertain about that wrong answer.

They're transparently providing an example where their product doesn't work well. Find me another product, even an OSS project that does the same on their landing page.

Re: Gopher – A 280B parameter language model

#26
post #15

The closer we get to artificial intelligence, the more we raise the bar for what qualifies as AI (as we should). Gopher/GPT-3 are already much more accurate than the average human at technical information retrieval (trivial to see from the dialogue transcripts: how many Americans know what a Schwarzschild metric is?). The focus on ethics and equity for these algorithms is interesting too, as the average human holds m…

When those language models are wrong or biased, the user will have a worse experience in all three of those scenarios. At least when we look at search results now, we can prune for the facts. Those language models are ingesting that same data to give a monolithic answer your a query. Less transparent, less safe.

I don't see a difference. Large language models can also return their sources, as in the example on the Gopher blog post. This will lead to a quicker answer and equal transparency.

Re: Gopher – A 280B parameter language model

#27
post #18

Earlier quoted context omitted.

One has to wonder if the final response is the first glimmer of an artificial sense of humor. Failing at simple arithmetic after nailing some advanced physics answers has the air of playful bathos.

Nothing like a little anthropomorphism to completely distort otherwise good faith interpretations of bot behavior.

How is the impression of playfulness not a good faith interpretation?

You of course know that the model is not capable of thought or reasoning - only the appearance of them as needed to match its training corpus. A training corpus of completely human generated data. As such, how could anything it does, be anything but anthropomorphic?

Now, if this model were trained exclusively on a corpus of mathematical proofs stripped of natural language commentary, the expectation that you seem to have would be more appropriate.

Re: Gopher – A 280B parameter language model

#28
post #19

Earlier quoted context omitted.

One has to wonder if the final response is the first glimmer of an artificial sense of humor. Failing at simple arithmetic after nailing some advanced physics answers has the air of playful bathos.

I think it's more likely that 5 came out because if it ever saw the answer, 105, before, it was split into the tokens [10][5] of which it only 'remembered' one. Or the numbers were masked when training (something that was done with BERT-like models) so it just knew enough to put a random one in

That seems likely and fair.

What moved me to post is that that kind of silly answer is the exact sort of shenanigans that I would pull if I were cast as the control group in a Turing test.

I already do such things winkingly when talking with my preschooler to send him epistemic tracer rounds and see if he's listening critically

Re: Gopher – A 280B parameter language model

#29
post #22
post #15

The closer we get to artificial intelligence, the more we raise the bar for what qualifies as AI (as we should). Gopher/GPT-3 are already much more accurate than the average human at technical information retrieval (trivial to see from the dialogue transcripts: how many Americans know what a Schwarzschild metric is?). The focus on ethics and equity for these algorithms is interesting too, as the average human holds m…

The problem seems to be that these models provide fairly accurate information at many occasions and occasionally complete blunders. Humans provide less accurate information most of the time but with a certain amount of self-reflection/meta-cognition, and they will usually recognize total blunders or display reasonable uncertainty about them. There are only very few applications where it would make sense to take the r…

Accuracy is improving rapidly though. I agree that the current accuracy levels are not high enough to be relied upon.

> Humans ... they will usually recognize total blunder

I question this assumption. I don't believe this is true, even for subject matter experts. I've worked with radiology data where experts with 10+ years of experience make blunders that disagree with a consensus panel of radiologists.

Re: Gopher – A 280B parameter language model

#30

Next to "Human Expert", I'd like to see it compared to "Average American" or "Average College Grad". That might be more of a realistic notion of how close this model is to everyday US citizenry rather than experts. Sure I'd love to see a radiology assistant, too.

It might be fun for a laugh.

What actual value would an AI that produces answers similar to the average person have, though? Non-expert answers for interesting questions are pretty much meaningless -- the whole point of an advanced society is that we can avoid knowing anything about most things and focus on narrow expertise.

Post reply on HN