Live data from Hacker News

GPT-3: Language Models Are Few-Shot Learners

arxiv.org

121–130 of 212 posts

Re: GPT-3: Language Models Are Few-Shot Learners

#121

Earlier quoted context omitted.

Dude, I’m sorry, but the average person will not know the difference between that and a regular buzzfeed article or YouTube comment. We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles. It’s over, we lost. Morpheus: What if I told you that, throughout your whole life,…

Hello. Gwern and I trained the GPT-2 1.5B model that powers /r/SubSimulatorGPT2. https://www.reddit.com/r/SubSimulatorGPT2/ I've been basically living and breathing GPT-2 for ... gosh, it's been 6 months or so. The past few months have been a lot of StyleGAN2 and a lot of BigGAN, but before that, it was very "make GPT-2 sing and dance in unexpectedly interesting ways" type work. I don't claim to know a lot. But occas…

Thank you for this comment. As someone who played a bit with GPT, it was very poignant for me. I still think it's incredible that GPT can put up such convincing facades, that it can generate genuinely novel and interesting text... but it's bittersweet, too, that it can't go any further with them. The ideas are lost in the context window.

I play AI dungeon on occasion, which uses GPT2 to generate freeform adventures. And I find over time that it's not really GPT2 that's writing stories, it's me. GPT2 is putting out plausible strings of words, but I'm the one giving them meaning, culling the parts that go off track, and guiding it in a direction I want to go.

And it is a bit melancholy. You see possibilities, nuances, subtexts, and meanings. The neural net sees words.

Re: GPT-3: Language Models Are Few-Shot Learners

#122

Earlier quoted context omitted.

But as a VERY publicly watched lab, they have a serious duty I was nodding right along with you, and then... OpenAI has no duty. It doesn't matter if they're publicly watched. What matters is whether the field of AI can be advanced, for some definition of "advanced" equal to "the world cares about it." It's important to let startups keep their spirit. Yeah, OpenAI is one of the big ones. DeepMind, Facebook AI, OpenAI…

I think all researchers and science communicators have a duty to present science in a way which educates and edifies, and doesn't mislead. It's not just that they're successful, but that their publicity gives them a prominent role as science communicators. Science is all about and questioning your assumptions, and acknowledging limitations. They claim the public interest in their charter. I think it's reasonable to d…

It's easy to say: they 'have a duty to present science in a way which educates and edifies, and doesn't mislead'. But sometimes it takes years even for scientists to really understand what they have created or discovered. It's cutting edge, not well known, hard to communicate. How could lay people keep up where not even scientists have grasped it fully?

Of course, if the same scientists were asked about something where the topic has settled, they could be more effective communicators.

Re: GPT-3: Language Models Are Few-Shot Learners

#123

Earlier quoted context omitted.

Dude, I’m sorry, but the average person will not know the difference between that and a regular buzzfeed article or YouTube comment. We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles. It’s over, we lost. Morpheus: What if I told you that, throughout your whole life,…

Hello. Gwern and I trained the GPT-2 1.5B model that powers /r/SubSimulatorGPT2. https://www.reddit.com/r/SubSimulatorGPT2/ I've been basically living and breathing GPT-2 for ... gosh, it's been 6 months or so. The past few months have been a lot of StyleGAN2 and a lot of BigGAN, but before that, it was very "make GPT-2 sing and dance in unexpectedly interesting ways" type work. I don't claim to know a lot. But occas…

> is not the same thing as "promote any idea/product."

GPT-3 seems to have quite a few paragraphs worth of context. A simple way to promote your product online with it is to give it a prefix of:

---

Comment1: Superbrush is amazing - I literally couldn't live without it. No other brush is as good.

Comment2: This brush is really good for tangled hair, and I love the soft smooth surface.

Comment3:

---

Then let it write a comment. Of all the comments it writes, manually filter a few thousand good ones, and use those as seeds to generate more, which you post all over the web. There's no need to do any training - the generic model should be fine given the right prefix.

Re: GPT-3: Language Models Are Few-Shot Learners

#124

Even though this was the GPT-3-generated text that humans most easily identified as machine-written, I still like it a lot: Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm Subtitle: Joaquin Phoenix pledged to not change for each awards event Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am…

I don't know if it says something about text generation or human text processing, but whenever I read an example of computer generated text, all through I think "I can't tell this is machine generated, it seems completely natural," and the only giveaway is that at the end I have no idea what it said. It's a pretty eerie feeling. It's as though both the AI and my short-term processing only pay attention to a context o…

This happens to me when I'm reading in a language I'm not very good at (German). Each sentence may make sense, but overall I feel I didn't get the point. I guess it's a cummulative error situation, where you reach a threshold after which the point is lost.

Re: GPT-3: Language Models Are Few-Shot Learners

#126
post #109

Earlier quoted context omitted.

Hello. Gwern and I trained the GPT-2 1.5B model that powers /r/SubSimulatorGPT2. https://www.reddit.com/r/SubSimulatorGPT2/ I've been basically living and breathing GPT-2 for ... gosh, it's been 6 months or so. The past few months have been a lot of StyleGAN2 and a lot of BigGAN, but before that, it was very "make GPT-2 sing and dance in unexpectedly interesting ways" type work. I don't claim to know a lot. But occas…

Great comment! Was it generated with GPT2 or GPT3? I understood all sentences but as a whole I will need to revisit.

I think it was written by a human, but the human had spent so much time with GPT-2 that they'd begun to emulate its writing style.

Re: GPT-3: Language Models Are Few-Shot Learners

#127

Even though this was the GPT-3-generated text that humans most easily identified as machine-written, I still like it a lot: Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm Subtitle: Joaquin Phoenix pledged to not change for each awards event Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am…

I don't know if it says something about text generation or human text processing, but whenever I read an example of computer generated text, all through I think "I can't tell this is machine generated, it seems completely natural," and the only giveaway is that at the end I have no idea what it said. It's a pretty eerie feeling. It's as though both the AI and my short-term processing only pay attention to a context o…

This has some parallels to generic random corporate PR or marketing speak. Some communications are already so automated and dehumanized that we are used to random content signed by a pseudo real person that we gloss over, and are more easily fall for GPT like generated content that has similar form. Edit: I mean I guess someone of a previous generation used to only read the newspaper and letters would more instantly spot something is wrong.

Re: GPT-3: Language Models Are Few-Shot Learners

#128

Earlier quoted context omitted.

I don't know if it says something about text generation or human text processing, but whenever I read an example of computer generated text, all through I think "I can't tell this is machine generated, it seems completely natural," and the only giveaway is that at the end I have no idea what it said. It's a pretty eerie feeling. It's as though both the AI and my short-term processing only pay attention to a context o…

To me it reads like a child telling a story, but that this child has an adult's ability to use language. When children tell a story they aren't going anywhere with it but don't know how to cover it up.

Maybe someone drunk then?

Re: GPT-3: Language Models Are Few-Shot Learners

#129
post #26

Earlier quoted context omitted.

How much does that cost?

21.3 USD/h in vast.ai (8X Tesla V100, 118.8 TFLOPS)

that's only 256GB, which isn't enough. I'm not sure it's even possible to nvlink 16 v100s. I'd love to try it out for $40/hr if it were possible though.

Re: GPT-3: Language Models Are Few-Shot Learners

#130
This looks like a big deal to me:

1. First of all, the authors successfully trained a model with 173 BILLION PARAMETERS. The previous largest model in the literature, Google’s T5, had "only" 11 billion. With Float32 representations, GPT-3-173B's weights alone occupy ~700GB of memory (173 billion params × 4 bytes/param). A figure in the 100's of billions is still 3 orders of magnitude smaller than the 100’s of trillions of synapses in the human brain [a], but consider this: Models with trillions of weights are suddenly looking... achievable.

2. The model achieves competitive results on many NLP tasks and benchmarks WITHOUT FINETUNING. Let me repeat that: there is no finetuning. There is only unsupervised (i.e., autoregressive) pretraining. For each downstream NLP task or benchmark, the pretrained model is given text instructions, and possibly sample text with questions and answers. The NLP tasks on which the model was tested include translation, question-answering, cloze tasks, unscrambling words, using novel words in sentences, and performing 3-digit arithmetic.

3. The model is tested only in a ZERO-SHOT or FEW-SHOT setting. In other words, for each NLP task, the pretrained model is given text instructions with zero examples, or text instructions with a small number of examples (typically 10 to 100). As with human beings, GPT-3-173B doesn't need lots of examples to perform competitively in novel NLP tasks.

4. The results reported by this paper on all NLP tasks and benchmarks should be seen as a BASELINE. These results likely could be meaningfully improved with conventional finetuning.

5. The model’s text generation FOOLS HUMAN BEINGS, without having to cherry-pick examples.

--

[a] https://www.google.com/search?q=number+of+synapses+in+human+...

Post reply on HN