Live data from Hacker News

Facebook uses 1.5B Reddit posts to create chatbot

bbc.com

141–150 of 235 posts

Re: Facebook uses 1.5B Reddit posts to create chatbot

#141
post #84

Earlier quoted context omitted.

I feel like there's something really interesting behind that choice that probably isn't that flattering to Facebook.

I would imagine Reddit, being a forum of threaded posts, has far, far, far more conversational interactions than Facebook where everything is basically one-shot, no threading. You want to train a convo bot on conversations. I doubt there's much more to it.

Facebook could have used all the 1 on 1 chats but maybe they didn't to avoid making it obvious that they have access?

Re: Facebook uses 1.5B Reddit posts to create chatbot

#142
post #78

Earlier quoted context omitted.

Good food is good food. Some good food happens to be vegan. It isn't hugely "special" especially these days, when Indian food is reasonably popular; this, incidentally, debunks the notion that vegans all eat weird concoctions of soy meant to resemble meat. I'm sure some do, but a curry which happens to contain no animal products is much more appealing.

Perhaps I'm being pedantic, to the point, most Indian food is vegetarian, not vegan. They love their milk, cheese, and honey.

[deleted]

Re: Facebook uses 1.5B Reddit posts to create chatbot

#143
post #135
post #124

Earlier quoted context omitted.

So what's the defense?

that's the rub. every AI sound/word/picture editor i've ran into says something along the lines of "we're releasing this data set to help stay secure in this day and age of easy counterfeiting of X.", but they never really mention how you apply the data in an adversarial way against itself -- they just sort of hand-wave that part. Same with fake AI generated Obama video and sound, and earlier data-set generated chatb…

If its secret or not publicly available people will argue using Occam’s razor or that only “State actors” could use this. With the subtext being your not important enough.

With the data public its more akin to driveby ssh login attempts. Not being important doesn’t mean your not under attack and people can take the necessary precautions.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#145
post #135
post #124

Earlier quoted context omitted.

So what's the defense?

that's the rub. every AI sound/word/picture editor i've ran into says something along the lines of "we're releasing this data set to help stay secure in this day and age of easy counterfeiting of X.", but they never really mention how you apply the data in an adversarial way against itself -- they just sort of hand-wave that part. Same with fake AI generated Obama video and sound, and earlier data-set generated chatb…

If you're curious to learn more about what's actually being done: https://arxiv.org/abs/1905.12616

Re: Facebook uses 1.5B Reddit posts to create chatbot

#146
post #137
post #126

Earlier quoted context omitted.

Yes and no. You're right that there's no button you can push to delete an entire account history, but wrong that there's no way to remove HN comments. We take care of deletion requests for people every day. We don't want anyone to get in trouble from anything they posted to HN, there's nearly always something we can do, and we don't send people away empty-handed. I can only think of one or two cases where we weren't…

I get the reasoning, but I don't see this applied to some platforms. Reddit and Discord allow you to both delete and edit older comments, and there's no limits on how far back you can go (so you can, if you wanted, edit or delete your entire history). Under the GDPR a subject is allowed full erasure rights. If I say I want you to delete my content from x date to y date, or a particular post, or everything entirely th…

I noticed a few days back you didn't like it when a user made a new account

I think you must have misunderstood whatever the moderation comment was, there's no prohibition on throwaway or multiple accounts. Just against using them to violate the site guidelines which is a different thing.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#147

Earlier quoted context omitted.

I would not encourage using the model for anything other than AI research -- we're still in the early days of dialogue, and there are a lot of unexplored avenues. There are still nuances around safety, controlling generation, consistency, and knowledge involvement. For instance, the bot cannot remember what you said even a few turns ago, due to limitations in memory size. In the paper, we did explore what happens whe…

But toxicity and quality is subjective. The technical achievement is undeniably brilliant, but the quality of the personality is subject to opinion - as I mentioned, I did not personally enjoy the agreeability of the bot. What's toxic today may not be toxic tomorrow and vice versa. It's just a matter of time before a model of this size can be run on commodity hardware and somebody will take the brakes off and/or atte…

For a more toxic version of a similar kind of bot, check out SubSimulatorGPT2: https://www.reddit.com/r/SubSimulatorGPT2/top/?sort=top&t=al...

Unfortunately you can't talk to it. (I've wanted to retrain a version that you can interact with dynamically, someday.)

Re: Facebook uses 1.5B Reddit posts to create chatbot

#150
post #71

Blog post: https://ai.facebook.com/blog/state-of-the-art-open-source-ch... Paper: https://arxiv.org/pdf/2004.13637.pdf Open Source: https://parl.ai/projects/recipes/ Ask us anything, the Facebook team behind it is happy to answer questions here.

For most the evaluations you reported at engagingness (expect Figure 16). Did you also look humanness? I would be especially interested how human your chatbot is compared to real humans (Figure 17). This would be similar to a turning test.
Post reply on HN