Live data from Hacker News

OpenAI Trains Language Model, Mass Hysteria Ensues

approximatelycorrect.com

41–50 of 119 posts

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#41
post #35
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

I've just read i.e https://twitter.com/gdb/status/1096098366545522688 and even though it's "best of 25" (I guess cherry-picked by a human) - this is mind-blowing. I am actually having a very hard time believing this is legit generated text.

Definitely impressive work, but the fact that this is hard to distinguish from human text, if true, is pretty sad for humans. Even sadder if anyone reading this could be swayed by such an argument.

Heck, maybe having to compete with this will raise human discourse (Joking).

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#42
post #38
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

> - The _whole point_ is that our model is not special and that other people can reproduce and improve upon what we did. We hope that when they do so, they too will reflect about the consequences of releasing their very powerful text generation models. If this is your whole point, then I think you are missing something fundamental. Implementing these models doesn't require reflection, or introspection, or any sort of…

Exactly. This is like holding up spam samples or how spammers operate from the spam detecting work. That side (and the cultural discussions) needs all the headstart it can get, not be complacent that some arbitrary "experts" will patronizingly "protect" them.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#43
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

> The _whole point_ is that our model is not special and that other people can reproduce and improve Only people with a large amount of money and a lot of expertise. What you are doing is the opposite of democratizing AI.

Actually this shows why OpenAI matters. Google have been training and refining Transformer architectures for years; how unlikely is it nobody tried training a language model at this scale or larger with similar results?

Yet from Google we heard nothing. Which is the optimal decision for them - they only lose by blowing the whistle.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#44
post #34
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

Deciding if feeding the media with fear was worth the attention you will get wasn't easy, ha. Let me tell you, you are the shame of the profession.

Yep, put out handpicked samples to stoke fear, then release nothing of the internals to stoke more fear and act like self-appointed gods.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#45
post #21

Earlier quoted context omitted.

>Secondly, have you seen the results? I was dumbfounded and fascinated. I spent hours reading the samples. Yes, I've seen the result. They're nice but, as the article points out, not extraordinary compared to state of the art, open NLP research. OpenAI's behaviour here smells of Gibsonesque 'anti-marketing', using the misunderstanding of AI and its capabilities in the general population as a means to stir up publicit…

> not extraordinary compared to state of the art, open NLP research > misrepresents progress in the field Can you point me to some examples of unsupervised learning with similar results? Not asking for rhetorical purposes; I just genuinely was shocked by how compelling their results were, especially given this was unsupervised. > OpenAI's behaviour here smells of Gibsonesque 'anti-marketing' I don't disagree that the…

>Can you point me to some examples of unsupervised learning with similar results? Not asking for rhetorical purposes; I just genuinely was shocked by how compelling their results were, especially given this was unsupervised.

Model wise this is just openAI's GPT with some very slight modification (laid out in the paper).

Ilya has now commented in the thread and essentially made the same point, this is state of the art performance, but reproducible by everyone because it uses a known architecture.

The secrecy and controversy makes no sense if the model is open, even the methodology of data collection is laid out. There is no safety here assuming that anybody who wants to rebuild the model can do so simply by putting enough effort into rebuilding the dataset, which is not an issue for a seriously malicious actor.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#46
post #29
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

Ok but isn’t this the opposite of OpenAI’s “nukes are safer when multiple actors have them” strategy wrt AI? I’m also confused by the threat models earnestly put forth in your blog post. Are we really concerned about deep faking someone’s writing? The plain word already demands attribution by default: we look for an avatar, a handle, a domain name to prove the person actually said this.

> Ok but isn’t this the opposite of OpenAI’s “nukes are safer when multiple actors have them” strategy wrt AI?

It seems more like the "nukes are safer when multiple rational state level actors have them", rather than anyone able to pull a git repo.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#47
post #31
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

So a small number of individuals decided what's best for everybody? How is that open? How is that not centralization of power?

If they did release it, there would be an equivalent outcry about how OpenAI was contributing to fake news, etc.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#48
post #37

It seems disingenuous that this article fails to quote examples of GPT-2’s stunning results, or give any contrasting results from BERT to support the claim that this is all normal and expected progress. Like many, I was viscerally shocked that the results were possible, the potential to further wreck the Internet seemed obvious, and an extra six months for security actors to prepare a response seemed like normal good…

What's so shocking about this? Why do we trust this in the hands of a few self-appointed experts than anyone else? Are they supposed to be more moral than any others? What will security experts do in six months that wouldn't benefit from more security experts looking at it? Why do you care that garbage text is machine generated, from a spammer or influencer, or a mechanical turk? If it's volume you're concerned about, should we complain when search/recommendation engines already aggregate and reweight a tiny opinion into a continuous out-of-proportion stream that can last you a lifetime to consume? What is the practical difference to have more volume existing "out there"?

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#50
post #6
post #4

Earlier quoted context omitted.

Cheaper cost to put out a bigger volume of content.

Cheaper than paying "influencers". Paid blogging was huge during the .com era. I wonder if this could be adapted, with suitably good speech synth, to produce podcasts en masse.

It could. Adobe has a model to generate speech from arbitrary text. All it needs is typical samples from a speaker and transcripts for training. You could easily make it sound like Obama, for example. It will match intonation, timing, etc. and maybe insert the occasional "uhm" or "uh" when appropriate.
Post reply on HN