Live data from Hacker News

OpenAI Trains Language Model, Mass Hysteria Ensues

approximatelycorrect.com

91–100 of 119 posts

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#91
post #35
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

I've just read i.e https://twitter.com/gdb/status/1096098366545522688 and even though it's "best of 25" (I guess cherry-picked by a human) - this is mind-blowing. I am actually having a very hard time believing this is legit generated text.

I couldn't be more disappointed with this bullshit honestly. The texts have almost zero coherence and keep repeating the same patterns (which they presumably learned from the data set) over and over again. If this is their best out of 25 samples then they aren't going to fool anyone.

>Recycling is NOT good for the world.

>It is bad for the environment,

>it is bad for our health,

>and it is bad for our economy.

>Recycling is not good for the environment.

>Recycling is not good for our health.

>Recycling is bad for our economy.

>Recycling is not good for our nation.

The first paragraph keeps repeating the is for the pattern 8 times.

>And THAT is why we need to |get back to basics| and |get back to basics| in our recycling efforts.

"get back to the basics" is repeated twice in the same sentence.

>Everything from the raw materials (wood, cardboard, paper, etc.),

>to the reagents (dyes, solvents, etc.)

>to the printing equipment (chemicals, glue, paper, ink, etc.),

>to the packaging,

>to the packaging materials (mercury, chemicals, etc.)

>to the processing equipment (heating, cooling, etc.),

>to the packaging materials,

>to the packaging materials that are shipped overseas and

>to the packaging materials that are used in the United States.

It literally repeated packaging 5 times in the same sentence and the overall structure was repeated 9 times. Also what type of packaging is based on mercury?

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#92
post #54

Earlier quoted context omitted.

Exactly. This is like holding up spam samples or how spammers operate from the spam detecting work. That side (and the cultural discussions) needs all the headstart it can get, not be complacent that some arbitrary "experts" will patronizingly "protect" them.

If you look at it as a PR stunt, it is almost certainly a good idea. If a bad actor can auto-generate text that is not really distinguishable from something written by a human, how does a community with open membership (eg, HN) protect itself? I imagine this technology will enable interesting new attacks against online communities; we havn't seen that for a while. OpenAI are extremely sensible to draw attention to th…

No, a more effective PR stunt would be to release the model, and better ones, and make it so easy any idiot could use them. THAT would catch the attention of Congress, and THAT would result in funds and lesiglation to combat it. This won’t even register on a sub committees staffers wet dream. It is not human nature to pay attention to far off hypothetical abstract threats, only concrete and immediate ones. You could release a thousand papers like this and it wouldn’t do anything even approaching the effect of congressmen and their staff getting assloads of fake but convincing email/docs/etc, the press being indicated with thousands of fake but convincing tips, of tens of thousands of people calling the police because some asshats are spamming them with convincing letters from their dead grandma or whatever, of convincing communication to banks or brokers, letters to agencies claiming widespread danger (ie there is salmonella in half the food at xyz), kids sending forged letters to their school from their supposed parents to let them leave campus, and so on. I’m sure you can think of better examples.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#93
post #35

Earlier quoted context omitted.

I've just read i.e https://twitter.com/gdb/status/1096098366545522688 and even though it's "best of 25" (I guess cherry-picked by a human) - this is mind-blowing. I am actually having a very hard time believing this is legit generated text.

I couldn't be more disappointed with this bullshit honestly. The texts have almost zero coherence and keep repeating the same patterns (which they presumably learned from the data set) over and over again. If this is their best out of 25 samples then they aren't going to fool anyone. >Recycling is NOT good for the world. >It is bad for the environment, >it is bad for our health, >and it is bad for our economy. >Recyc…

The parts you criticise are the parts I was most impressed with. These sorts of repetitions can be persuasive in writing/arguments, and it's impressive to me that a model learned this type of writing.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#94
post #54

Earlier quoted context omitted.

If you look at it as a PR stunt, it is almost certainly a good idea. If a bad actor can auto-generate text that is not really distinguishable from something written by a human, how does a community with open membership (eg, HN) protect itself? I imagine this technology will enable interesting new attacks against online communities; we havn't seen that for a while. OpenAI are extremely sensible to draw attention to th…

I’m not entirely sure that that bad actor would get any more scalablity form it than from a Mechanical Turk farm, at least as far as impact goes. It seem that as far as information warfare goes “less is more” works quite well and they rely on targeted people to spread the news for them. When you want to drive an agenda you don’t need unique 100,000 comments you need a good copy pasta. Overall I’m sick of this dramati…

But a Mechanical Turk is traceable and definitely not anonymous. Using a self contained model somewhere on a server/cluster/workstation could be.

Regarding an agenda, sure, good pasta is fine and all, and regular ol people are fine, but it is not cost effective. This is a million times cheaper, which means you can use it everywhere, not just the obvious places, you can be everywhere, and you can do more than just push a couple big items, you could push tens of thousands of them, micro targeted all the way down to the individual. Don’t dismiss it so easy—the potential scale is far, far larger than anything existing to date.

And I would note that the reason 100,000 comments aren’t effective now is precisely because they are too formulaic, too obviously fake when used on such a large scale. This has the potential to create real, live, seemingly active and believable online communities of millions of people, all at fractions and fractions and fractions of a penny compared to current methods. People read news, then comments (or reviews or whatever), because they use them to determine the validity of the content they just read; if it’s no longer possible to tell from the comments what’s a scam and what isn’t... well, you could do a lot of things with that.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#95
post #72

Earlier quoted context omitted.

But ... it's not novel. We could already generate convincing gibberish years ago. Now the novelty is that this can be better targeted. But even simple Markov-chain based text generators were good enough to fool people for a bit. And there was always people that had too much free time to write. A lot. (See for example the crackpots and conspiracy theorists that bombard physics forums. See the 9/11, Zeitgeists, etc. mo…

Markov-chain generators are extremely lacking in long-term coherency. They rarely even make complete sentences, much less stay on topic! They were not convincing at all-- and many of the GPT-2 samples are as "human-like" as average internet comments. Conjecture: GPT-2 trained on reddit comments could pass a "comment turing test", where the average person couldn't distinguish whether a comment is bot or human with bet…

That's an indictment of reddit comments more than AI. Remember that conditioned on the human-provided seed prompt, there is no statistical surprise (the definition of information) in the generated text. If all reddit comments are are riffs on the OP based on second-hand information, well then they may as well be bot-generated already.

At this stage, these AI's can only help. Imagine we are given this tool that can generate samples from the "uninformative but realistic looking text" distribution, we can then put it in a discriminator to filter out blabbering bots and humans together, or invert it to summarize the small kernel of information, and that would be a great thing. The better these models learn about typical human behavior the better off we are at identifying the truly exceptional. It's when AI starts to sense and incorporate novel information from the non-human environment that you really have to worry.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#96

Earlier quoted context omitted.

Markov-chain generators are extremely lacking in long-term coherency. They rarely even make complete sentences, much less stay on topic! They were not convincing at all-- and many of the GPT-2 samples are as "human-like" as average internet comments. Conjecture: GPT-2 trained on reddit comments could pass a "comment turing test", where the average person couldn't distinguish whether a comment is bot or human with bet…

That's an indictment of reddit comments more than AI. Remember that conditioned on the human-provided seed prompt, there is no statistical surprise (the definition of information) in the generated text. If all reddit comments are are riffs on the OP based on second-hand information, well then they may as well be bot-generated already. At this stage, these AI's can only help. Imagine we are given this tool that can ge…

>That's an indictment of reddit comments more than AI.

Perhaps, but that's the world we live in. I suspect the average reddit commenter is already more articulate than the average person (citation needed, I know. But reddit skews highly educated young male in a first-world country. There's no way they do worse than a worldwide average).

Other than that, I agree with your comment.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#97
post #29

Earlier quoted context omitted.

Ok but isn’t this the opposite of OpenAI’s “nukes are safer when multiple actors have them” strategy wrt AI? I’m also confused by the threat models earnestly put forth in your blog post. Are we really concerned about deep faking someone’s writing? The plain word already demands attribution by default: we look for an avatar, a handle, a domain name to prove the person actually said this.

> Ok but isn’t this the opposite of OpenAI’s “nukes are safer when multiple actors have them” strategy wrt AI? It seems more like the "nukes are safer when multiple rational state level actors have them", rather than anyone able to pull a git repo.

More like “nukes are safer when we control them and the rest of you cite them”

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#98

To what extent is this not just finding text samples written in its training sample and regurgitating it near verbatim?? -Non ml guy

You bring up a good point. Without seeing their code and training metrics, how do we know that this isn’t some extremely overfitted model?

From the paper:

"All models still underfit WebText and held-out perplexity has as of yet improved given more training time."

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#99
post #72
post #54

Earlier quoted context omitted.

If you look at it as a PR stunt, it is almost certainly a good idea. If a bad actor can auto-generate text that is not really distinguishable from something written by a human, how does a community with open membership (eg, HN) protect itself? I imagine this technology will enable interesting new attacks against online communities; we havn't seen that for a while. OpenAI are extremely sensible to draw attention to th…

But ... it's not novel. We could already generate convincing gibberish years ago. Now the novelty is that this can be better targeted. But even simple Markov-chain based text generators were good enough to fool people for a bit. And there was always people that had too much free time to write. A lot. (See for example the crackpots and conspiracy theorists that bombard physics forums. See the 9/11, Zeitgeists, etc. mo…

But ... it's not novel.

I work in this field, and yes, this is very novel (at least in terms of the quality).

It's the biggest improvement in quality I've ever seen. The long term coherence is so much better than anything else that has ever been built.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#100
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

I think you should at least release a small portion of the training data (e.g. anything recycling related) so people can measure to what extent the model is generating new sentences and to what extent it's just regurgitating training data.
Post reply on HN