Live data from Hacker News

OpenAI Trains Language Model, Mass Hysteria Ensues

approximatelycorrect.com

111–119 of 119 posts

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#111
post #71

Earlier quoted context omitted.

A lot of people have results similar to this - but most people generating a paragraph of slightly_weird_but_plausible_if_you_read_quickly text using a primped version of BERT one time out of 25 regarded it as more or less pointless. But journalists don't. This would be ok if this is the first time that anyone had a media go wild over AI story. But actually this has happened 10000 times this year already.

Seems like the way it worked is that the blog post was discussed here and on Twitter and many people thought it was interesting. Then some journalists picked it up and wrote about it. That much is nothing out of the ordinary. It is interesting (at least to those of us who aren't natural language researchers) so why shouldn't we talk about it? Why shouldn't journalists write about it? Inevitably their mildly controver…

I'd shrug and move on, but the problem is that I believe that these flaps about AI are distracting attention from the real concerns and forces that are having a serious impact on people now.

The distortion of public debate caused by community exclusiveness on social platforms, by the curation and manipulation of social feeds and by the dynamics of online debate where the loudest and angriest voices dominate is one place that we could do with some focus.

Another place is the management of simple models - plain Jane stuff like a learned classifier - people are making these with Python and R and releasing them into infrastructures and apps and we don't know what they are and where they are and how they are interacting.

Instead we have wizard of oz style stories to distract us from who's actually hiding behind the curtain. If we fall for this then we may find ourselves living in a vicious totalitarian society with no obvious way out of it.

Journalists should write about it in an informed and professional way, that's fine. But they need to write about stories that are impactful and important, and if they were to write about this one in this way ("text scrambler makes a pretty good paragraph one out of 30 tries, has no idea of what is going on") they would get no clicks (there will now be a second wave of follow-ups like that to ride on the coattails of the story). Instead they have to make it sound like robots are going to take children from schools and experiment on them live on TV, and this makes them famous and rich.

There is no real revision of the story because the follow on stories disappear from view while search engines and other journalists use the original hysteria. Look at what happened with the two negotiating bots at facebook (the game was to negotiate to get books and balls, the bots tended to use a short hand to negotiate rather than the english they were trained on) This was "Facebook researchers have to pull the plug on AI that they no longer understand", and that is the narrative that we will have on that story more or less forever.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#112
post #108
post #93

Earlier quoted context omitted.

The parts you criticise are the parts I was most impressed with. These sorts of repetitions can be persuasive in writing/arguments, and it's impressive to me that a model learned this type of writing.

> These sorts of repetitions can be persuasive in writing/arguments That is the saddest part. It's not because AI is good, it's because we count saying "X is good/bad" 3 times as a persuasive argument. It won't be hard to learn this kind of "arguing", it's just sad that's what we're teaching our AIs to do and get excited when they do it.

> saying "X is good/bad" 3 times as a persuasive argument

I didn't say that it's a persuasive argument, I said that it can be persuasive IN arguments. There's nothing sad about an AI learning it, or people being happy with it, it's very impressive.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#113
post #87
post #80

Earlier quoted context omitted.

I'm not sure it has much in the way of implications. There is no real profit to be made by generating realistic looking text. Spammers don't work that way, spammers haven't cared about realistic looking text for years. Nor have spam filters cared much about text for a long time, exactly because it's so easy to randomise. Anti-spam is not a good reason to hold back on language generation models, in my view. As for HN,…

You’re fooling yourself if you think there are no significant uses of text generation. Fake news, propaganda, advertising, fake reviews, fake everything. Fabricated email from friends family and colleagues. Whole online communities fabricated out of whole cloth. It is a weapon, and a powerful one.

No, it's useless and I speak from experience of dealing with spammers who forged mail from friends family and colleagues in the past.

People are not trivial automatons who can have their opinions rewritten on the fly by auto-generated text. If auto-generated text reaches into its giant grab-bag of learned expressions and produces something actually interesting or insightful, people might be interested in that new line of thinking, but if - like many of these examples - it's essentially rambling if coherent nonsense then it won't have any impact at all.

So I rather think it's you fooling yourself. You've been reading comments online for years without knowing who or what produced them. If you discovered half of them were artificial tomorrow, what difference would it make? The people around you are already judging arguments based on the content, not their volume or who wrote them.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#114
post #23

Ilya from OpenAI here. Here's our thinking: - ML is getting more powerful and will continue to do so as time goes by. While this point of view is not unanimously held by the AI community, it is also not particularly controversial. - If you accept the above, then the current AI norm of "publish everything always" will have to change - The _whole point_ is that our model is not special and that other people can reprodu…

> - If you accept the above, then the current AI norm of "publish everything always" will have to change

Ok, accepting that premise, what people/organisations would you share the research with and based on what criteria?

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#115
post #52

Earlier quoted context omitted.

>Can you point me to some examples of unsupervised learning with similar results? Not asking for rhetorical purposes; I just genuinely was shocked by how compelling their results were, especially given this was unsupervised. Model wise this is just openAI's GPT with some very slight modification (laid out in the paper). Ilya has now commented in the thread and essentially made the same point, this is state of the art…

> Model wise this is just openAI's GPT with some very slight modification (laid out in the paper). > The secrecy and controversy makes no sense if the model is open, even the methodology of data collection is laid out. This is exactly why I found the results so compelling: It suggests that this technology is already accessible to some big players: The odds that a Big Corp. or govt agency has already begun using the t…

Interesting to think about whether state actors already have such technology.

If they did, I bet it would be used for automated "troll farms".

Like weaponized malicious ELIZA, it would have fake user profiles reacting to keywords, spinning suitable counter-argumentation and/or lies for as long as it takes to change opinions and perceptions, relentlessly, day and night.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#116
post #72

Earlier quoted context omitted.

But ... it's not novel. We could already generate convincing gibberish years ago. Now the novelty is that this can be better targeted. But even simple Markov-chain based text generators were good enough to fool people for a bit. And there was always people that had too much free time to write. A lot. (See for example the crackpots and conspiracy theorists that bombard physics forums. See the 9/11, Zeitgeists, etc. mo…

Markov-chain generators are extremely lacking in long-term coherency. They rarely even make complete sentences, much less stay on topic! They were not convincing at all-- and many of the GPT-2 samples are as "human-like" as average internet comments. Conjecture: GPT-2 trained on reddit comments could pass a "comment turing test", where the average person couldn't distinguish whether a comment is bot or human with bet…

I know they are extremely lacking, but compared to that a hyper-fancy NN with layers and layers of the darkest of black magic, trained at the zenith of the night for thousands of man years in the crypts of the terror itself, the TPU ... yeah, so it's not surprising it's better.

But it's no symbolic reasoning. It's not constructing a counter-argument from your argument. It simply lives off previous epic rap battles of internet flamewar history about .. well, about anything, since it's the Internet, and people like to chat, talk, write essays on every topic there is. Satire too. So there is always something to build that lang model on.

Though that will come too. Eventually.

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#117
post #99
post #72

Earlier quoted context omitted.

But ... it's not novel. We could already generate convincing gibberish years ago. Now the novelty is that this can be better targeted. But even simple Markov-chain based text generators were good enough to fool people for a bit. And there was always people that had too much free time to write. A lot. (See for example the crackpots and conspiracy theorists that bombard physics forums. See the 9/11, Zeitgeists, etc. mo…

But ... it's not novel. I work in this field, and yes, this is very novel (at least in terms of the quality). It's the biggest improvement in quality I've ever seen. The long term coherence is so much better than anything else that has ever been built.

No, I'm sorry, I wasn't precise enough. Yes, it's an amazing feat of engineering, and a truly great peak of text generation. But it's that. Text generation.

Yes, it can serve as great customized propaganda generator, and yes, people can be spin 'round and 'round with it. But they can be already with pretty much anything, from the simplest of phrases from "make X great again" to the elaborate scams of new age bullshit.

I simply disagree on the "virulence" or weaponization factor of this with others. (Especially when it comes to the possible "defenses", none can be "deployed" in 6 months. You can't teach critical thinking to billions of people overnight.)

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#118
post #117
post #99

Earlier quoted context omitted.

But ... it's not novel. I work in this field, and yes, this is very novel (at least in terms of the quality). It's the biggest improvement in quality I've ever seen. The long term coherence is so much better than anything else that has ever been built.

No, I'm sorry, I wasn't precise enough. Yes, it's an amazing feat of engineering, and a truly great peak of text generation. But it's that. Text generation. Yes, it can serve as great customized propaganda generator, and yes, people can be spin 'round and 'round with it. But they can be already with pretty much anything, from the simplest of phrases from "make X great again" to the elaborate scams of new age bullshit…

I've worked in the computational propaganda field, and I tend to agree that there is no real known defense yet.

I don't have a strong opinion about if they should have released this model or not.

I do know it would make a great commercial spam generator though. Want a million product reviews which seem legitimate quickly? This is the thing..

Re: OpenAI Trains Language Model, Mass Hysteria Ensues

#119
post #109
post #105

Earlier quoted context omitted.

> some of the samples generated by the model Mostly it's scary not because it's good - as writing goes, it's quite bad. It forms coherent sentences, but otherwise it's nonsense. I've seen similar nonsense producers in early 90s on basis of Markov chains and what not. No, the scary part is how much it reminds me of what I am reading in the media all the time. My current pet concern is that AIs will start passing the T…

> I've seen similar nonsense producers in early 90s on basis of Markov chains and what not. Exactly. When it comes to generating a large volume of apparently-good sentences, non-AI (or classical) approaches are still better than good. Those will be equally disruptive, since the defending side is yet to develop a proper countermeasure based on the "sensible"-ness of content. Plus, they will be much easier to customize…

> Exactly. When it comes to generating a large volume of apparently-good sentences, non-AI (or classical) approaches are still better than good.

Can you cite your source? I find this hard to believe.

Post reply on HN