How do you go about running a model this large?
Realistically you'd be able to train up to 760M param models. For that you'd need 32gb+ VRAM GPUs which I think AWS might have. You can try iwth 16gb VRAM GPUs, but you would need to figure out FP16. https://github.com/shawwn is doing some work in the GPT-2 space including using TPUs instead -- which has given him pretty good results.
GPT-3: Language Models Are Few-Shot Learners
81–90 of 212 posts
Re: GPT-3: Language Models Are Few-Shot Learners
#82Earlier quoted context omitted.
They can, they can crowd out good comments with an absolutely crushing volume of crap. While humans can put out a lot of crap, bots can do orders of magnitude more. The issue is that it is an effort multiplier for when a small number of people want to target a particular forum.
I imagine the problem isn’t so much volume of crap comments as much as the tailoring of crap comments. Imagine if every tweet-into-the-void from a human with 50 followers reliably got engaging replies. Bots taunting your grammar mistakes, bots selectively quoting your prior tweets to point out contradictions, bots cleverly insinuating that your tweets reveal problematic sympathies. So much of our noise-filtering is i…
My hope is the next step will be filtering by insightfulness/usability of a comment and then those best bots bought and used by next stack overflow: https://xkcd.com/810/
Re: GPT-3: Language Models Are Few-Shot Learners
#83I am not a fan of this trend of "Language Models Are X" in recent work particularly out of OpenAI. I think it's a rhetorical sleight of hand which hurts the discourse. Like, the exact same paper could have instead been titled "Few-Shot Learning with a Large-Scale Language Model" or similar. But instead there seems to be this extremely strong desire to see certain ineffable qualities in neural networks. Like, it's a l…
"Few-Shot Learning with a Large-Scale Language Model" makes more sense.
Even with their robot hand paper, they titled it along the lines of "we solved a rubrix cube" not "a robot hand manipulated the cube and solved it"
Re: GPT-3: Language Models Are Few-Shot Learners
#84I am not a fan of this trend of "Language Models Are X" in recent work particularly out of OpenAI. I think it's a rhetorical sleight of hand which hurts the discourse. Like, the exact same paper could have instead been titled "Few-Shot Learning with a Large-Scale Language Model" or similar. But instead there seems to be this extremely strong desire to see certain ineffable qualities in neural networks. Like, it's a l…
I was nodding right along with you, and then...
OpenAI has no duty. It doesn't matter if they're publicly watched. What matters is whether the field of AI can be advanced, for some definition of "advanced" equal to "the world cares about it."
It's important to let startups keep their spirit. Yeah, OpenAI is one of the big ones. DeepMind, Facebook AI, OpenAI. But it feels crucial not to reason from the standpoint of "they have achieved success, so due to this success, we need to carefully keep an eye on them."
Such mindsets are quite effective in causing teams to slow down and second-guess themselves. Maybe it's not professional enough, they reason. Or perhaps we're not clear enough. Maybe our results aren't up to "OpenAI standards."
As to your specific point, yes, I agree in general that it's probably good to be precise. And perhaps "Language Models Are Few-Shot Learners" is less precise than "Maybe Language Models Are Few-Shot Learners."
But let's be real for a moment: this is GPT-3. GPT-2 is world-famous. It's ~zero percent surprising that GPT-3 is "something big." So, sure, they're few-shot learners.
In time, we'll either discover that language models are in fact few shot learners, or we'll discover that they're not. And that'll be the end of it. In the meantime, we can read and decide for ourselves what to think.
Re: GPT-3: Language Models Are Few-Shot Learners
#85I am not a fan of this trend of "Language Models Are X" in recent work particularly out of OpenAI. I think it's a rhetorical sleight of hand which hurts the discourse. Like, the exact same paper could have instead been titled "Few-Shot Learning with a Large-Scale Language Model" or similar. But instead there seems to be this extremely strong desire to see certain ineffable qualities in neural networks. Like, it's a l…
Re: GPT-3: Language Models Are Few-Shot Learners
#86This part really freaked me out... GPT-2 couldn't do math: Context → Passage: Saint Jean de Br´ebeuf was a French Jesuit missionary who travelled to New France in 1625. There he worked primarily with the Huron for the rest of his life, except for a few years in France from 1629 to 1633. He learned their language and culture, writing extensively about each to aid other missionaries. In 1649, Br´ebeuf and another missi…
Question: How many years did Saint Jean de Br´ebeuf stay in New France before he went back to France for a few years?
Answer: 4
Explanation: The model used the arithmetic expression - 1629 + 1633 = 4.
NAQANet (trained on DROP) - came out in 2019 is able to do reasoning, you have to click result twice. First once it thinks it got it from passage, second attempt it tries to do arithmetic.
https://demo.allennlp.org/reading-comprehension/MjEzMjE1Ng==
Re: GPT-3: Language Models Are Few-Shot Learners
#87Even though this was the GPT-3-generated text that humans most easily identified as machine-written, I still like it a lot: Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm Subtitle: Joaquin Phoenix pledged to not change for each awards event Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am…
Dude, I’m sorry, but the average person will not know the difference between that and a regular buzzfeed article or YouTube comment. We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles. It’s over, we lost. Morpheus: What if I told you that, throughout your whole life,…
Re: GPT-3: Language Models Are Few-Shot Learners
#88Re: GPT-3: Language Models Are Few-Shot Learners
#89Re: GPT-3: Language Models Are Few-Shot Learners
#90Even though this was the GPT-3-generated text that humans most easily identified as machine-written, I still like it a lot: Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm Subtitle: Joaquin Phoenix pledged to not change for each awards event Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am…
What is it about ai generated texts that on skimming through it it makes sense, but if you try to slow down and understand it feels absurd and surreal.