Live data from Hacker News

Better Language Models and Their Implications

blog.openai.com

111–120 of 138 posts

Re: Better Language Models and Their Implications

#111
post #100

Sure, not releasing the full trained model probably delays it, but sooner or later a bad actor will do their own scraping and train their own model and share it around and the genie will be out of the bottle. Then what? I think we need to be conducting AI research (and building software generally) under the assumption that all of it will eventually be repurposed by bad actors. How would our practices be different if…

wasn't that the point of this whole openai thing? they didn't like the idea of there being a club with just google in it that had access to resources and funding to collect and train on massive datasets so they were going to be the "bad actors" who would do their own scraping, train their own models and share them around? isn't it supposed to be called OPENai? they don't want to share the data because they don't want…

>faint praise for soccer champion who apparently keeps winning games through poor performance.

I can’t disagree.

Re: Better Language Models and Their Implications

#112

This is so crazy good, someone needs to do a Turing test by sending it to some unsuspecting publishers. I get the feeling that debatepocolypse is not far away. Every forum can now be spammed with reasonable sounding gibberish that humans will have to slog through.

Few MIT students did that with their simple text generator. It is one way to expose conferences with low submission standards.

http://news.mit.edu/2015/how-three-mit-students-fooled-scien...

Re: Better Language Models and Their Implications

#113
post #100

Sure, not releasing the full trained model probably delays it, but sooner or later a bad actor will do their own scraping and train their own model and share it around and the genie will be out of the bottle. Then what? I think we need to be conducting AI research (and building software generally) under the assumption that all of it will eventually be repurposed by bad actors. How would our practices be different if…

wasn't that the point of this whole openai thing? they didn't like the idea of there being a club with just google in it that had access to resources and funding to collect and train on massive datasets so they were going to be the "bad actors" who would do their own scraping, train their own models and share them around? isn't it supposed to be called OPENai? they don't want to share the data because they don't want…

I have had the same impression regarding their work on dota. They got a lot of publicity with it but their work is not open at all. They have released neither their code which runs the bots on dota2 nor their training code nor the final model. All we have is video recordings of a few games against humans.

Re: Better Language Models and Their Implications

#114
post #54

Earlier quoted context omitted.

Agreed that "simply" scaling up with more compute will result in progress and useful systems, and work in that direction is interesting and valuable. But, while we may not need new architectures or training objectives to make progress, we do need them to approach human level sample complexity. Humans don't need to read through 40 GB of text multiple times to learn to write.

> Agreed that "simply" scaling up with more compute will result in progress and useful systems, and work in that direction is interesting and valuable. But, while we may not need new architectures or training objectives to make progress, we do need them to approach human level sample complexity. Yes, agreed. Nothing I said above contradicts that! :-) > Humans don't need to read through 40 GB of text multiple times to…

After 40 GB of text the model doesn't know anything about how the world works, and it shows many times in the examples. Nobody would do some of those mistakes, not even young kids. Other mistakes are more subtle but still show a total lack of understanding.

Then yes, it's enough to write text that nobody really cares about and that could cover a lot of what we read.

Re: Better Language Models and Their Implications

#115

Sure, not releasing the full trained model probably delays it, but sooner or later a bad actor will do their own scraping and train their own model and share it around and the genie will be out of the bottle. Then what? I think we need to be conducting AI research (and building software generally) under the assumption that all of it will eventually be repurposed by bad actors. How would our practices be different if…

CommonCrawl already has open dataset in petabyte size ready on AWS. Even if it didn’t exist, scrapping 80GB of data in AWS is trivial. I am surprised authors considered this as such a big deal. Also notice that performance is not anywhere close to humans. It sort of works and it’s astonishing that it does but long way to go before we have to fear weaponizing text generation.

Re: Better Language Models and Their Implications

#116
post #25

This tech can easily be used to flood humanity’s shared brain with auto-generated propaganda. Schizophrenia of the internet in a way. There is plenty of incentive with Google algorithms favoring number of words and relevant keywords in content for rankings - you could have NLP bots lifting junk sites to top results. To step ahead in that chess game, a detection tool for fake would be just training grounds for better…

Web of trust. Your results and how much you trust given text is based on who gave it to you. You assign trust to people you know, then it's a small world effect and recursive trust calculation based on who they trust.

Centralization won't work. Whether something is good or bad or fake must be subjective and based on your personal network / your beliefs. Otherwise, long term, I don't see how you could avoid dystopia.

Re: Better Language Models and Their Implications

#117
post #96
post #63

Earlier quoted context omitted.

40 GB is surprisingly close. I estimate I already read at least 4 GB of text so far. That's just 10 times more samples. I probably write better than GPT-2, but certainly not faster.

4GB of English? 500 words per single-spaced page, 5 letters on average per English word, so 2500 bytes/page. 4 gigs then would be 1,600,000 pages. That's 219 single-spaced pages per day every single day for 20 years straight. High, but I guess not outside the realm of possibility depending on the complexity of the text.

For perspective, 219 pages amounts to about a railway novel in a day, or about two ~800-page fantasy doorstoppers in a week. At a rule of thumb of about a page per minute for light reading, it's just under four hours: a hefty investment, but well within the realm of possibility. High compared to the general population, but table stakes for a book club or otherwise avid reader.

I'd've easily doubled or tripled that over the summer as a teenager.

Re: Better Language Models and Their Implications

#118

Is anyone else troubled by them not releasing the source model/dataset/parameters here? Yes, the technology can be used for malicious means - but would argue that "DeepFaking" language is FAR less of a problem than "DeepFaking" video/photo/audio... which already occurs. Seems like they went back on their charter to share AI developments broadly ("not concentrate power") under the excuse of "safety." (These results lo…

Quote from their charter [0]: Today this includes publishing most of our AI research, but we expect that safety and security concerns will reduce our traditional publishing in the future, while increasing the importance of sharing safety, policy, and standards research.

They referenced that from the article you commented on in the section 'release strategy' [1]: [...] we see this current work as potentially representing the early beginnings of such concerns, which we expect may grow over time. This decision, as well as our discussion of it, is an experiment: while we are not sure that it is the right decision today, we believe that the AI community will eventually need to tackle the issue of publication norms in a thoughtful way in certain research areas. Other disciplines such as biotechnology and cybersecurity have long had active debates about responsible publication in cases with clear misuse potential, and we hope that our experiment will serve as a case study for more nuanced discussions of model and code release decisions in the AI community.

[0] https://blog.openai.com/openai-charter/ [1] https://blog.openai.com/better-language-models/#releasestrat...

Re: Better Language Models and Their Implications

#119
post #13

Just as many pesticides mimic the hormonal and chemical signals of pests to drive certain behaviors that lead to eradication, this work mimics the linguistic signals of humans. I think viewing it metaphorically as the most sophisticated humanicide discovered to date is probably appropriate. Consider that conventional munitions make an effective pesticide but are not used due to their side effects. Instead, chemicals…

[deleted]

Re: Better Language Models and Their Implications

#120

Earlier quoted context omitted.

There’s a natural way to parallelize these models so that using 128 GPUs is the same as a 128x batch size. You can similarly simulate 128x batch size by accumulated gradients before backpropping. So you can test on just one or a few GPUs before you run the full thing. By that point you know it’s going to work, it’s just a matter of how well and whether you could’ve done nominally better with different tuning. There’s…

Thanks. >By that point you know it’s going to work, it’s just a matter of how well and whether you could’ve done nominally better with different tuning. This can't be true in all cases, right? I'm assuming that for many initially promising results on less-compute when they scale it, the results aren't impressive. I'm very curious to know what is the trials-to-success rate of publishable results when big-compute is th…

It’s indeed a very high trials to success ratio. Again though, there’s enough papers preceding this one that you could have good confidence in the effort. Another thing that helps is orgs like OpenAI have their own servers, rather than renting ec2 instances.

You also don’t just launch that many things and them ignore it. You monitor it to make sure nothing is going terribly wrong.

But yeah there’s also the fact that if you’re Google, throwing $2m worth of compute at something becomes worth it for some reason (eg Starcraft)

Post reply on HN