Live data from Hacker News

GPT-3: Language Models Are Few-Shot Learners

arxiv.org

81–90 of 212 posts

Re: GPT-3: Language Models Are Few-Shot Learners

#81
post #29
post #21

How do you go about running a model this large?

Realistically you'd be able to train up to 760M param models. For that you'd need 32gb+ VRAM GPUs which I think AWS might have. You can try iwth 16gb VRAM GPUs, but you would need to figure out FP16. https://github.com/shawwn is doing some work in the GPT-2 space including using TPUs instead -- which has given him pretty good results.

32GB RAM is nowhere near what's necessary. It can barely fit the GPT-2 models on there with a batch size in the single digits. We'll need extensive model parallelism libraries (Zero2 from Microsoft) to run this at all.

Re: GPT-3: Language Models Are Few-Shot Learners

#82

Earlier quoted context omitted.

They can, they can crowd out good comments with an absolutely crushing volume of crap. While humans can put out a lot of crap, bots can do orders of magnitude more. The issue is that it is an effort multiplier for when a small number of people want to target a particular forum.

I imagine the problem isn’t so much volume of crap comments as much as the tailoring of crap comments. Imagine if every tweet-into-the-void from a human with 50 followers reliably got engaging replies. Bots taunting your grammar mistakes, bots selectively quoting your prior tweets to point out contradictions, bots cleverly insinuating that your tweets reveal problematic sympathies. So much of our noise-filtering is i…

> What happens when every spam comment seems to understand the OP, even when the OP’s true audience is negligible?

My hope is the next step will be filtering by insightfulness/usability of a comment and then those best bots bought and used by next stack overflow: https://xkcd.com/810/

Re: GPT-3: Language Models Are Few-Shot Learners

#83

I am not a fan of this trend of "Language Models Are X" in recent work particularly out of OpenAI. I think it's a rhetorical sleight of hand which hurts the discourse. Like, the exact same paper could have instead been titled "Few-Shot Learning with a Large-Scale Language Model" or similar. But instead there seems to be this extremely strong desire to see certain ineffable qualities in neural networks. Like, it's a l…

Yeah its kind of odd why OpenAI makes these weird titles.

"Few-Shot Learning with a Large-Scale Language Model" makes more sense.

Even with their robot hand paper, they titled it along the lines of "we solved a rubrix cube" not "a robot hand manipulated the cube and solved it"

Re: GPT-3: Language Models Are Few-Shot Learners

#84

I am not a fan of this trend of "Language Models Are X" in recent work particularly out of OpenAI. I think it's a rhetorical sleight of hand which hurts the discourse. Like, the exact same paper could have instead been titled "Few-Shot Learning with a Large-Scale Language Model" or similar. But instead there seems to be this extremely strong desire to see certain ineffable qualities in neural networks. Like, it's a l…

But as a VERY publicly watched lab, they have a serious duty

I was nodding right along with you, and then...

OpenAI has no duty. It doesn't matter if they're publicly watched. What matters is whether the field of AI can be advanced, for some definition of "advanced" equal to "the world cares about it."

It's important to let startups keep their spirit. Yeah, OpenAI is one of the big ones. DeepMind, Facebook AI, OpenAI. But it feels crucial not to reason from the standpoint of "they have achieved success, so due to this success, we need to carefully keep an eye on them."

Such mindsets are quite effective in causing teams to slow down and second-guess themselves. Maybe it's not professional enough, they reason. Or perhaps we're not clear enough. Maybe our results aren't up to "OpenAI standards."

As to your specific point, yes, I agree in general that it's probably good to be precise. And perhaps "Language Models Are Few-Shot Learners" is less precise than "Maybe Language Models Are Few-Shot Learners."

But let's be real for a moment: this is GPT-3. GPT-2 is world-famous. It's ~zero percent surprising that GPT-3 is "something big." So, sure, they're few-shot learners.

In time, we'll either discover that language models are in fact few shot learners, or we'll discover that they're not. And that'll be the end of it. In the meantime, we can read and decide for ourselves what to think.

Re: GPT-3: Language Models Are Few-Shot Learners

#85

I am not a fan of this trend of "Language Models Are X" in recent work particularly out of OpenAI. I think it's a rhetorical sleight of hand which hurts the discourse. Like, the exact same paper could have instead been titled "Few-Shot Learning with a Large-Scale Language Model" or similar. But instead there seems to be this extremely strong desire to see certain ineffable qualities in neural networks. Like, it's a l…

For a specific example of how I think their framing is unhelpful: in the LAMBADA evaluation (sec. 3.1), they suggest that one-shot performance is low "perhaps...because all models still require several examples to recognize the pattern." This may be the first thing you'd think of for a few-shot learner, but then why is zero-shot performance higher than one-shot? If you remember that you're working with a language model, there's another possible explanation: the model probably models the last paragraph is a narrative continuation of the previous ones, and gets confused by the incongruity or distractors. (The biggest model is able to catch on to the incongruity, but only when it's seen it before, i.e., with >1 example.) Of course, this is just one possible explanation, and it's arguable, but the point is I think it's more useful to think of this as a language model being used for few-shot learning, not a few-shot learner where language modeling is an implementation detail.

Re: GPT-3: Language Models Are Few-Shot Learners

#86
post #28

This part really freaked me out... GPT-2 couldn't do math: Context → Passage: Saint Jean de Br´ebeuf was a French Jesuit missionary who travelled to New France in 1625. There he worked primarily with the Huron for the rest of his life, except for a few years in France from 1629 to 1633. He learned their language and culture, writing extensively about each to aid other missionaries. In 1649, Br´ebeuf and another missi…

Passage: Saint Jean de Br´ebeuf was a French Jesuit missionary who travelled to New France in 1625. There he worked primarily with the Huron for the rest of his life, except for a few years in France from 1629 to 1633. He learned their language and culture, writing extensively about each to aid other missionaries. In 1649, Br´ebeuf and another missionary were captured when an Iroquois raid took over a Huron village . Together with Huron captives, the missionaries were ritually tortured and killed on March 16, 1649. Br´ebeuf was beatified in 1925 and among eight Jesuit missionaries canonized as saints in the Roman Catholic Church in 1930.

Question: How many years did Saint Jean de Br´ebeuf stay in New France before he went back to France for a few years?

Answer: 4

Explanation: The model used the arithmetic expression - 1629 + 1633 = 4.

NAQANet (trained on DROP) - came out in 2019 is able to do reasoning, you have to click result twice. First once it thinks it got it from passage, second attempt it tries to do arithmetic.

https://demo.allennlp.org/reading-comprehension/MjEzMjE1Ng==

Re: GPT-3: Language Models Are Few-Shot Learners

#87

Even though this was the GPT-3-generated text that humans most easily identified as machine-written, I still like it a lot: Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm Subtitle: Joaquin Phoenix pledged to not change for each awards event Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am…

Dude, I’m sorry, but the average person will not know the difference between that and a regular buzzfeed article or YouTube comment. We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles. It’s over, we lost. Morpheus: What if I told you that, throughout your whole life,…

[deleted]

Re: GPT-3: Language Models Are Few-Shot Learners

#89
With things like this, we will need to change how the media work, and how we read news. Every sentence, every factual statement will need to be verified by some kind of chain of trust involving entities with reputation, or be labeled as "fiction/opinion".

Re: GPT-3: Language Models Are Few-Shot Learners

#90
post #63

Even though this was the GPT-3-generated text that humans most easily identified as machine-written, I still like it a lot: Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm Subtitle: Joaquin Phoenix pledged to not change for each awards event Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am…

What is it about ai generated texts that on skimming through it it makes sense, but if you try to slow down and understand it feels absurd and surreal.

There are some very prominent politicians whose primary mode of speech is the same
Post reply on HN