Live data from Hacker News

GPT-2: 6-Month Follow-Up

openai.com

21–30 of 98 posts

Re: GPT-2: 6-Month Follow-Up

#21
post #14

Earlier quoted context omitted.

Many important NLP tasks have almost nobody publicly working on them Well, then perhaps you should go work on them, instead of ranting here.

Why the ad hominem? I am pointing a problem of allocation of ressources on the AI research field. It's not to me to fixe that, but yes I am actively working on a logical fallacies detector which is the first of human history and works for the 256 possible forms of syllogisms, I'm expanding it to other logical forms such as modus ponens/tollens.

It's not to me to fixe that

There's nothing to fix. People work on what they want to work on. Things that seem important to you are not important to me, and the opposite. I'm OK with that.

Re: GPT-2: 6-Month Follow-Up

#22
Hopefully someone will make a working demo of it, like Adam King did for 345M. People should be able to experiment with this stuff without relying on the hype of press releases:

https://medium.com/@VictorBanev/interrogating-gpt-2-345m-aaf...

Not sure why open AI doesn't do this themselves. That fully aligns with their stated mission.

Re: GPT-2: 6-Month Follow-Up

#23
post #12
post #9

Earlier quoted context omitted.

Are you using FP16?

No. We weren't sure if that'd be a good idea since it wasn't trained with low-precision, and 345M thankfully didn't require going that far. 744M might, though. (Another option is model parallelism since I have 2 GPUs and that might be enough, perhaps freezing more layers and training incrementally, or reducing the 1024 token window to smaller ones like 700.)

You maybe should try tensorflow automatic mixed precision! https://github.com/zihangdai/xlnet/pull/200

Re: GPT-2: 6-Month Follow-Up

#24
post #12
post #9

Earlier quoted context omitted.

Are you using FP16?

No. We weren't sure if that'd be a good idea since it wasn't trained with low-precision, and 345M thankfully didn't require going that far. 744M might, though. (Another option is model parallelism since I have 2 GPUs and that might be enough, perhaps freezing more layers and training incrementally, or reducing the 1024 token window to smaller ones like 700.)

On the NVIDIA GPT- 2 implementation:

>What would be the largest model one could train across 2x 2080Ti?

>~800M gpt2. this is largely due to the memory required to House parameters + optimizer states. If one uses a smaller optimizer than Adam training something larger should be possible. Make sure to turn on activation checkpointing with —checkpoint-activations

https://twitter.com/TheRealRPuri/status/1161322580126470145

Re: GPT-2: 6-Month Follow-Up

#25

Are there any real use case for GPT-2? Does it solve any problem? I've read almost all state of the art leaderboards of all Nlp tasks of paperswithcode.com and truth is except text generation, openAI has not one state of the art, they are not even visible in leaderboards. OpenAI is maybe the AI research center with the biggest funding and comparatively to other well known (Microsoft, Facebook, Google or even zalando.…

> are there any real use case?

Writing Bloomberg's "market wrap" articles.

Re: GPT-2: 6-Month Follow-Up

#26
post #22

Hopefully someone will make a working demo of it, like Adam King did for 345M. People should be able to experiment with this stuff without relying on the hype of press releases: https://medium.com/@VictorBanev/interrogating-gpt-2-345m-aaf... Not sure why open AI doesn't do this themselves. That fully aligns with their stated mission.

https://talktotransformer.com/

Re: GPT-2: 6-Month Follow-Up

#27
post #12

Earlier quoted context omitted.

No. We weren't sure if that'd be a good idea since it wasn't trained with low-precision, and 345M thankfully didn't require going that far. 744M might, though. (Another option is model parallelism since I have 2 GPUs and that might be enough, perhaps freezing more layers and training incrementally, or reducing the 1024 token window to smaller ones like 700.)

On the NVIDIA GPT- 2 implementation: >What would be the largest model one could train across 2x 2080Ti? >~800M gpt2. this is largely due to the memory required to House parameters + optimizer states. If one uses a smaller optimizer than Adam training something larger should be possible. Make sure to turn on activation checkpointing with —checkpoint-activations https://twitter.com/TheRealRPuri/status/11613225801264701…

They haven't released such models, though, and I don't know if it would be drop-in compatible with the OA GPT-2-774M checkpoint (they're training their own GPT-2s using their own webtext corpus).

Re: GPT-2: 6-Month Follow-Up

#28
"Cornell University is studying human susceptibility to digital disinformation generated by language models."

"The Middlebury Institute of International Studies Center on Terrorism, Extremism, and Counterterrorism (CTEC) is exploring how GPT-2 could be misused by terrorists and extremists online."

"The University of Oregon is developing a series of “bias probes” to analyze bias within GPT-2."

But apparently no university studies the social and economic impact of using terabytes of public data to train algorithms that for all practical reasons end up being inaccessible to an average person.

If things go on the way they're going right now, in 20 years millions of people will be "mechanical turked". Most of information processing tools will be mediated exclusively through companies like Google and Amazon. They will be less like normal tools (e.g. word processors) and more like systems you have to be a part of. Can you imagine the levels of inequality involved? The hyper-centralization of power? This is the foremost challenge presented by AI, not some hypothetical nonsense involving terrorists using a text generator.

And it's not like there aren't any solutions. Douglas Engelbart, for example, pointed out a great way of introducing technology into society without screwing most of the society over:

http://dougengelbart.org/content/view/138

We kind of followed his vision for a while, with good results, but AI seems to be going in an entirely different direction.

Re: GPT-2: 6-Month Follow-Up

#29
I was able to take all of Donald Trumps tweets and using GPT2 to make a program that would mimic his tweets.

I found that it might be very effective. I have the test at

https://docs.google.com/forms/d/1p7tlobl5y5plBCu_enK4KawR7B8...

I got the information from trumptwitterarchive.com

I also explored creating a system that could recognize fake tweets from real ones and I believe I got 94% accuracy. It was a Bayes classifier but I think I have to double check my work.

Re: GPT-2: 6-Month Follow-Up

#30
>As part of our staged release strategy, our current plan is to release the 1558M parameter model in a few months, but it’s plausible that findings from a partner, or malicious usage of our 774M model, could change this.

This seems naive but I think it's a misdirection. Of course the model will have malicious users. Propaganda teams started testing its integration as soon as it was released. It's likely that OpenAI is counting on this for insights into HOW the model can be used maliciously. It's also possible that the model results have inherent trackable markers and OpenAI can later say that X% of social media posts were made using this model.

So what are the positive applications, aside from prettifying data like sports and weather reports?

Even with Skyrim's 800+ books, you frequently ran into the same book. Imagine libraries filled with plausible text that hides nuggets of lore seeded by developers. Along with more realistic text-to-speech this can allow games to support a large diversity of NPCs that have true radiant dialogue and sound more realistic than "I saw a mudcrab the other day".

With some modifications, I think models like this can outweigh even their nefarious applications:

Defense against text decomposition analysis. The model can be used to obfuscate writing patterns that can reveal a person's identity, either by randomizing form or standardizing it. Take your post and run it through the formatter to get the same idea and intent, but in a style that can't be traced to your other writing. Or you reform it into style of Ernest Hemmingway, like thousands of others.

Realtime plausible deniability encryption. Messages in a monitored chat can look like mundane conversation but contain encrypted messages. This would require the model accept seeds and work partially in reverse to diff two sets of text to reveal the hidden message.

In it's current form it doesn't look like it can do any of those things, but there's the potential.

Post reply on HN