Live data from Hacker News

GPT-3: Language Models Are Few-Shot Learners

arxiv.org

191–200 of 212 posts

Re: GPT-3: Language Models Are Few-Shot Learners

#191
post #174

They say there are no stupid questions, so here is mine: If there are Billions of parameters in the SOTA models, how do we argue that they are not over fitting?

That's section 4 of OP.

Thank you. It is quite a labor to even skim through the 50+ page paper. Your poignant reply was quite helpful to draw my attention to the issue of contamination. After reading the section carefully, I think my understanding of over fitting is very much improved at least in so far as models like GPT-3 are concerned.

Clearly, the authors have given careful considerations to the issue of contamination and have provided reasonable analysis and a careful argument regarding over fitting the existing benchmarks.

On the other I was wondering if the authors would like to consider purposefully creating a type of "out of sample data" for "creative evaluation"? Of course, GPT is no stranger to creativity, so it would be a fascinating challenge to come up with methods to create such datasets that are truly creative and challenge GPT-{N} to prove its mettle.

For example, would it be possible to engage a really good creative writer* along with a highly experienced school teacher to take on the Reading Comprehension task and create few "tricky" evaluation samples that not only go above and beyond the contamination objections but also challenge the human intelligence to be careful not to fall into common traps?

This way lies a different evaluation metric - a subjective one perhaps, but it's a start. Just a thought experiment - that's all.

* so that they can come up with new ways to trick GPT/humans a teacher knows the common mistakes the average student makes

Edit: Duh, my head immediately screamed GANs the moment I pressed submit, lol. But I am not sure if GANs make sense for NLP tasks. Like do they make sense if humans/domain experts try to solve them?

Re: GPT-3: Language Models Are Few-Shot Learners

#193
post #191

Earlier quoted context omitted.

That's section 4 of OP.

Thank you. It is quite a labor to even skim through the 50+ page paper. Your poignant reply was quite helpful to draw my attention to the issue of contamination. After reading the section carefully, I think my understanding of over fitting is very much improved at least in so far as models like GPT-3 are concerned. Clearly, the authors have given careful considerations to the issue of contamination and have provided…

You might be interested in the ELECTRA model. It's the solid first success I've seen of a GAN-like framework in NLP. It also has links to why GANs still don't do so great in NLP in its references.

Re: GPT-3: Language Models Are Few-Shot Learners

#194

Earlier quoted context omitted.

Hello. Gwern and I trained the GPT-2 1.5B model that powers /r/SubSimulatorGPT2. https://www.reddit.com/r/SubSimulatorGPT2/ I've been basically living and breathing GPT-2 for ... gosh, it's been 6 months or so. The past few months have been a lot of StyleGAN2 and a lot of BigGAN, but before that, it was very "make GPT-2 sing and dance in unexpectedly interesting ways" type work. I don't claim to know a lot. But occas…

I really like your observation about memory. Because you seem open minded to wild ass guesses and going meta: I have a hunch that general intelligence will be the ability to learn from mistakes. Not just optimization. I mean applying the scientific method. Hypothesis, prediction, run experiment, compare expected vs actual. And having a notion, any notion, to explain the delta between expected and actual. Am total noo…

It's deeper than that.

Currently, there's no research into torturing AI. Why not?

A pain response is universal across most life forms with a nervous system. We seek to replicate a nervous system. Pain would seem to be far easier to replicate than the scientific method.

My wife sat me down and told me a story that horrified me. She had to get it off her chest, and I was sad it happened to her than to me. She was sitting around on the porch and felt something on her leg, and brushed it off. When she got up and looked down, apparently she had stepped on a poor snail. His shell was... And he was...

He wasn't dead. So she frantically looked up what to do. But there was nothing to do. Snails in that situation can't be helped, and the most humane thing is to put it out of its writing anguish, its full-body torture.

She put on some boots, took it out to the sidewalk, and stomped it as hard as she could. And that was the story of that snail.

You probably felt more for that snail than you've ever felt for any AI bot. Why?

It's worth considering.

Re: GPT-3: Language Models Are Few-Shot Learners

#195
post #191

Earlier quoted context omitted.

Thank you. It is quite a labor to even skim through the 50+ page paper. Your poignant reply was quite helpful to draw my attention to the issue of contamination. After reading the section carefully, I think my understanding of over fitting is very much improved at least in so far as models like GPT-3 are concerned. Clearly, the authors have given careful considerations to the issue of contamination and have provided…

You might be interested in the ELECTRA model. It's the solid first success I've seen of a GAN-like framework in NLP. It also has links to why GANs still don't do so great in NLP in its references.

Thanks a lot.

If I may ask one more question, would you happen to know if the authors or other researchers who are entertaining any theoretical work on the experimental design and training methodologies of GPT/BERT? As in why does it work? What is the significance of training via the "fill-in-the-blanks" method?

Don't get me wrong - the work is great and the SOTAs are amazing, I would be just happy to have a chat to discuss and bounce some ideas what all this means and why do these methods seem to be working so well. Papers/articles/blog-posts are always a pleasure to read!

Re: GPT-3: Language Models Are Few-Shot Learners

#196
post #165
post #94

Earlier quoted context omitted.

Do you have a source? I am genuinely curious as I can't find it and would like to see the context

https://www.youtube.com/watch?v=0TRtSk-ufu0 Around 19:10~. Though I messed up, he didn't say 'genuinely'. He said "full stop, truly, legitimately, we have an algorithm that can learn".

Thanks!

Re: GPT-3: Language Models Are Few-Shot Learners

#197

Earlier quoted context omitted.

Hello. Gwern and I trained the GPT-2 1.5B model that powers /r/SubSimulatorGPT2. https://www.reddit.com/r/SubSimulatorGPT2/ I've been basically living and breathing GPT-2 for ... gosh, it's been 6 months or so. The past few months have been a lot of StyleGAN2 and a lot of BigGAN, but before that, it was very "make GPT-2 sing and dance in unexpectedly interesting ways" type work. I don't claim to know a lot. But occas…

> is not the same thing as "promote any idea/product." GPT-3 seems to have quite a few paragraphs worth of context. A simple way to promote your product online with it is to give it a prefix of: --- Comment1: Superbrush is amazing - I literally couldn't live without it. No other brush is as good. Comment2: This brush is really good for tangled hair, and I love the soft smooth surface. Comment3: --- Then let it write…

To be a bit less wordy: try it. You stand to earn lots of money.

Narrator: it didn't work

(Going into the reasons it doesn't actually work in practice is... lengthy. It's human dynamics. Would you buy a product from a sales guy that can't remember your name? That's sales 101. And loading up the context window only gets you so far. That "working memory" is tiny, tinytinytiny. Even at 1024 tokens, it means you have to boil down the entire history of an interaction to a few pages at most. Which is a lot, sure, but it's this balancing act where you'll need to retrain the model to support your custom context format for your specific "slots" – a "slot" being a piece of knowledge, like the client's name. Or you can try encoding all of that in natural language, AI dungeon style. But I recently played AI dungeon and pretended to be buying a router from the store. The cashier stripped down and started jacking off onto his desk. I don't have high hopes for our ability to control these models in a business context.)

Re: GPT-3: Language Models Are Few-Shot Learners

#198

Earlier quoted context omitted.

To be fair, human poem authors generate and reject tons of bad poems too!

The difference is that the best authors know how to reject their own bad poems. I think the next step is a GAN that learns to weed out bad writing.

To be honest, I'm not sure that even the best authors reject their own poor work. "The complete works of X" for any poet X is frequently full of, well, duds. It's other people who extract the best of X's output and popularize it.

Re: GPT-3: Language Models Are Few-Shot Learners

#199
post #55
post #50

10^4 petaflop-second/days. They missed an opportunity to be the first paper to measure their computation in mole flops.

Rather, chemists constantly miss opportunities to use actual numbers instead of their lazy legacy mole nonsense. Nobody really seems to use SI prefixes beyond peta or occasionally exa. But they could have called this 900 zetta-flop. (10^4 peta-flop/s-days)

But "mole flop" is one of the best units ever. It's better than furlongs per fortnight.

Re: GPT-3: Language Models Are Few-Shot Learners

#200

Earlier quoted context omitted.

They can, they can crowd out good comments with an absolutely crushing volume of crap. While humans can put out a lot of crap, bots can do orders of magnitude more. The issue is that it is an effort multiplier for when a small number of people want to target a particular forum.

I imagine the problem isn’t so much volume of crap comments as much as the tailoring of crap comments. Imagine if every tweet-into-the-void from a human with 50 followers reliably got engaging replies. Bots taunting your grammar mistakes, bots selectively quoting your prior tweets to point out contradictions, bots cleverly insinuating that your tweets reveal problematic sympathies. So much of our noise-filtering is i…

Would you be fooled by that? No because once it became well known, people would adjust their heuristics for judging who's real. If that's too hard, platforms would help, such as by verifying identity more thoroughly. I've never heard any realistic description of how AI bots could somehow undermine society even close to the amount that humans already do.
Post reply on HN