Live data from Hacker News

GPT-2: 6-Month Follow-Up

openai.com

71–80 of 98 posts

Re: GPT-2: 6-Month Follow-Up

#71
post #28

"Cornell University is studying human susceptibility to digital disinformation generated by language models." "The Middlebury Institute of International Studies Center on Terrorism, Extremism, and Counterterrorism (CTEC) is exploring how GPT-2 could be misused by terrorists and extremists online." "The University of Oregon is developing a series of “bias probes” to analyze bias within GPT-2." But apparently no univer…

It seems like this is missing the point of public data? When you make an edit to Wikipedia, anyone in the world can read it. You don't benefit when they read an article, but it doesn't cost you anything either.

"Anyone" includes researchers. That's part of the deal. Yes, they benefit, but you aren't harmed. That's zero-sum thinking.

Re: GPT-2: 6-Month Follow-Up

#72
post #66

The OpenAI approach to managing the release of the larger dataset strikes me as totally flawed and upside down. The biggest concern the team seem to have is that the fully trained GPT2 model will be used to spread propaganda and misinformation. They also imply that the biggest hurdle to training a similar model is money needed to pay for the training resources. The problem with this approach is that the users most li…

Uh, you're forgetting about spammers and malware authors.

Re: GPT-2: 6-Month Follow-Up

#73
post #28

"Cornell University is studying human susceptibility to digital disinformation generated by language models." "The Middlebury Institute of International Studies Center on Terrorism, Extremism, and Counterterrorism (CTEC) is exploring how GPT-2 could be misused by terrorists and extremists online." "The University of Oregon is developing a series of “bias probes” to analyze bias within GPT-2." But apparently no univer…

Whoa, it’s remarkable how well that paper holds up for being nearly 60 years old. I wonder if it’s just a timeless problem.

Re: GPT-2: 6-Month Follow-Up

#74
post #66

The OpenAI approach to managing the release of the larger dataset strikes me as totally flawed and upside down. The biggest concern the team seem to have is that the fully trained GPT2 model will be used to spread propaganda and misinformation. They also imply that the biggest hurdle to training a similar model is money needed to pay for the training resources. The problem with this approach is that the users most li…

I’d recommend re-reading the original GPT2 announcement, particularly this section regarding their release policy:

This decision, as well as our discussion of it, is an experiment: while we are not sure that it is the right decision today, we believe that the AI community will eventually need to tackle the issue of publication norms in a thoughtful way in certain research areas.

This release approach is an experiment used to force the conversation around a release strategy before we actually and unambiguously need it.

Re: GPT-2: 6-Month Follow-Up

#75
post #28

"Cornell University is studying human susceptibility to digital disinformation generated by language models." "The Middlebury Institute of International Studies Center on Terrorism, Extremism, and Counterterrorism (CTEC) is exploring how GPT-2 could be misused by terrorists and extremists online." "The University of Oregon is developing a series of “bias probes” to analyze bias within GPT-2." But apparently no univer…

I'm not sure if this captures this full set of tensions at play. I'm a huge fan of Doug Engelbart, and he was my inspiration for a long time...but it turns out augmentation of power without checks and balances can end up very messy (I'm also a fan of James Madison, though Engelbart is closer to a first love!).

One could argue that I actually started work on misinformation though Engelbart. I was working on various projects specifically to achieve a vision like his, and it is still my guiding light. But it turns out that some technologies (initially focusing on Facebook and YouTube's engagement dynamics...), without appropriate checks and balances, are sort of anti-augmentation of intellect. They give an asymmetric advantage to those who are trying to weaken our intellects. So I ended up dropping those projects to attempt to address urgent misinformation issues in May of 2016.

Going back to GPT-2 and research release, I and my co-author recently went deep into understanding the types of risks and tradeoffs in a recent paper. You can see the summary here: https://medium.com/@aviv/reducing-malicious-use-of-synthetic... or go directly to arXiv https://arxiv.org/pdf/1907.11274.pdf . The goal of our paper is specifically to go past the angry invective of the "here is the most important problem" (in your case, "AI inequality") and actually dive into the weeds of threat models and tradeoffs.

Re: GPT-2: 6-Month Follow-Up

#76
post #28

"Cornell University is studying human susceptibility to digital disinformation generated by language models." "The Middlebury Institute of International Studies Center on Terrorism, Extremism, and Counterterrorism (CTEC) is exploring how GPT-2 could be misused by terrorists and extremists online." "The University of Oregon is developing a series of “bias probes” to analyze bias within GPT-2." But apparently no univer…

We (the Middlebury Institute's CTEC) are an extremism and terrorism research lab, and so we're tracking the ways that tech is used by terrorists and extremists. For a lot of nonstate orgs with sophisticated propaganda arms, an ideologically cohesive text generation capability would be a huge advantage in scaling up info ops. We are looking to measure whether or not GPT-2 or other neural text generators are useful for…

I think their point isn't that terrorists leveraging this tech not a problem. It is certainly a problem. But the greater problem being a few large entities being the only ones who have access to or control over it.

I think it's pretty clear that terrorists or any other bad actor will find great value & utility in this tech. The article from OpenAI says 'Humans can be convinced by synthetic text.' & research at Cornell found people find it almost as convincing as New York Times articles. I would be interested in learning about the methods you guys are using to determine of it? I wonder how that could be measured?

So let's assume the answer is 'YES! This technology is dangerous". The Middlebury program, Cornell, and more and more universities and research groups find the same thing. Then what will the recommendations be? Certainly not to release it into the wild. I think they will be to keep it locked up. To keep it in the hands of a few large and powerful companies, with the resources to 'manage' such a thing.

This seems to be what the original comment is trying to illustrate, and I think it's an interesting point to consider the implications of long term. The tech exists now. There is no going back. So is it worse to let it out of the box, or to let a but a few have control over it?

Re: GPT-2: 6-Month Follow-Up

#77
post #28

"Cornell University is studying human susceptibility to digital disinformation generated by language models." "The Middlebury Institute of International Studies Center on Terrorism, Extremism, and Counterterrorism (CTEC) is exploring how GPT-2 could be misused by terrorists and extremists online." "The University of Oregon is developing a series of “bias probes” to analyze bias within GPT-2." But apparently no univer…

It seems like this is missing the point of public data? When you make an edit to Wikipedia, anyone in the world can read it. You don't benefit when they read an article, but it doesn't cost you anything either. "Anyone" includes researchers. That's part of the deal. Yes, they benefit, but you aren't harmed. That's zero-sum thinking.

I think you are correct that nothing is taken away from an author when someone reads their Wikipedia article.

Perhaps what the poster above you is saying is that there is a continuum of information, some more personal and sensitive, like your current location, and some less personal, like the Wikipedia article on Elephants.

Taking data about specific humans, (or humans in general), and turning it into code that has predictive power seems like a different type of power than the knowledge given by an encyclopedia.

Re: GPT-2: 6-Month Follow-Up

#78
Even if GPT-2 were released, very very few would have the hardware to run it because of gpu ram running out (and doing some sort of load-unload system would make training times unfeasibly long). And those who have the hardware to run it, has probably already made a version of their own or reasons not to. So I'm wondering if this GPT-2 hype is a genuine concern of openai, or if it's mostly a PR flex to say 'Look at us, we made a good model!'.

As an example, look here by Nvidia https://devblogs.nvidia.com/training-bert-with-gpus/ who made GPT-2 8B, which is ~5 times as large as GPT-2.

Re: GPT-2: 6-Month Follow-Up

#79
post #57

Earlier quoted context omitted.

Ouch! So 11GB is nowhere close to being enough, then. I wonder if even switching to FP16 will be adequate?

Might be able to get 745M down to work on a single GPU. I'm definitely not using all 24GB, so fp16 might be able to get it down enough.

How would you use fp16 to get it to work on a single GPU? And if you did, what GPU should you use?

Re: GPT-2: 6-Month Follow-Up

#80
post #12

Earlier quoted context omitted.

No. We weren't sure if that'd be a good idea since it wasn't trained with low-precision, and 345M thankfully didn't require going that far. 744M might, though. (Another option is model parallelism since I have 2 GPUs and that might be enough, perhaps freezing more layers and training incrementally, or reducing the 1024 token window to smaller ones like 700.)

fp16 saves a lot of memory and is worth doing. I've not had trouble fine tuning all these models with fp16.

Have you fine tuned 774 successfully using a single GPU?
Post reply on HN