Earlier quoted context omitted.
Cheaper cost to put out a bigger volume of content.
Is volume really what dictates whether or not you can impersonate someone? It's never seemed that way to me.
OpenAI Trains Language Model, Mass Hysteria Ensues
61–70 of 119 posts
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#62Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#63It seems disingenuous that this article fails to quote examples of GPT-2’s stunning results, or give any contrasting results from BERT to support the claim that this is all normal and expected progress. Like many, I was viscerally shocked that the results were possible, the potential to further wreck the Internet seemed obvious, and an extra six months for security actors to prepare a response seemed like normal good…
Why? There were news about bots writing news ~5 years ago. Given a few simple facts the AI generated the regular info-scarce but fluffy news-piece.
Now OpenAI added better everything (better language models, more data, better "long-term memory" for overall text coherence), and we got better fluff.
It seems like a GAN and a simple Markov chain generator. (Even if it's not that simple of course.)
And maybe it's the equivalent of the "modern art meme" style transferred to AI/ML research. ( https://i.pinimg.com/236x/71/e1/21/71e12151f4b59d8433d32c126... )
What I'm trying to convey is that wrecking the net with auto-trolls was already possible, but for some reason Mechanical Turk was cheaper.
> OpenAI warned everyone of an “exploit” in which text humans can trust to be human-generated
Sokal already did that, and so did http://thatsmathematics.com/mathgen/ ... but of course this might be qualitatively different, because it can be targeted. (Weaponized, if you will.) But the defense/antidote is the same, but it takes a lot more than 6 months to make people better at critical thinking, but maybe you already heard about the difficulties of that :)
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#64Earlier quoted context omitted.
I have seen the results and I don't get why people think this is any more dangerous than journalists who selectively report to fit a predetermined agenda or make shit up on the spot. Which, today, is a lot of them.
It's Bulk. Same reason why spam is a problem.
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#65Earlier quoted context omitted.
> Ok but isn’t this the opposite of OpenAI’s “nukes are safer when multiple actors have them” strategy wrt AI? It seems more like the "nukes are safer when multiple rational state level actors have them", rather than anyone able to pull a git repo.
Yep. Maybe I misunderstood the subtler points of OpenAI’s “democratize AI” strategy, and this has been the plan all along. But AFAIK they haven’t put an “among a few rational state actors” asterisk on anything up until now. Regardless, I agree with TFA that this is a silly and arbitrary time to yell “fire.” It’s PR.
True. On the PR side though, it'd be incredibly hard to say "we want to make replication moderately difficult, but not too difficult." Everyone would end up arguing exactly how much should be released, how it would prevent X,Y,Z folks from contributing to AI, etc.
> Regardless, I agree with TFA that this is a silly and arbitrary time to yell “fire.” It’s PR.
Alternatively, it does provide good insight into the reactions in the community as a whole, and continues the conversation on exactly how much should be released. Maybe I'm not far enough into the ML community, but the decision not to put the "keys to the kingdom" on github for every script kiddie to weaponize seems reasonable to me, especially as a precedent.
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#66Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…
Have you done a plagiarism search on that text to see how similar it is to the input corpus? I'm by no means an ML expert, but I've played around with models for random name generation and one thing I've noticed is that as the models become more accurate, they also become much more likely to just regurgitate existing names verbatim. So if you search the list of names and notice something that seems particularly realistic, it could be because it's literally taken in whole or in part from the training data set!
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#67Earlier quoted context omitted.
I've just read i.e https://twitter.com/gdb/status/1096098366545522688 and even though it's "best of 25" (I guess cherry-picked by a human) - this is mind-blowing. I am actually having a very hard time believing this is legit generated text.
Definitely impressive work, but the fact that this is hard to distinguish from human text, if true, is pretty sad for humans. Even sadder if anyone reading this could be swayed by such an argument. Heck, maybe having to compete with this will raise human discourse (Joking).
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#68The strongest counterargument I've seen to OpenAI's decision is that the decision won't end up mattering, because someone else will eventually replicate the work and publish a similar model. But it still seems like a reasonable choice on OpenAI's part–they're warning us that some language model will soon be good enough for malicious use (e.g. large-scale astroturfing/spam), but they're deciding it won't be theirs (and giving the public a chance to prepare).
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#69Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…
Here are a few that comes to mind.
-Secrecy? but how will you continue to exist on the PR scene if you don't release anything?
-Are you willing to pay every developer who is able to replicate your paper, more than what the black market would pay?
-How are you working on incentive alignment to make sure that all people who can replicate your results have more incentive to do good than bad, specially in the current environment where users and valuable data are silo-ed by a few companies?
-Misdirection to keep an edge, i.e. planting bugs/ Not fixing bugs for public ; spreading false results; only working on problems that need high resources to limit the number of actor who will be able to replicate ?
-Tracking the people who have the competence to replicate and take preemptive measures.
-Restrictions on GPU/CPU/silicone wafer.
Who can regulate? How can we regulate? What are the negative consequence of regulation? What happens if we don't, at what odds and time horizon?
Re: OpenAI Trains Language Model, Mass Hysteria Ensues
#70Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…
> - I suggest going over some of the samples generated by the model. Many people react quite strongly, e.g., https://twitter.com/justkelly_ok/status/1096111155469180928 . Have you done a plagiarism search on that text to see how similar it is to the input corpus? I'm by no means an ML expert, but I've played around with models for random name generation and one thing I've noticed is that as the models become more acc…
(The talking unicorn example on their page is also meant to demonstrate that, no, it's not just memorizing, but I think it's a bit more compelling to check from the raw samples)