Live data from Hacker News

Facebook uses 1.5B Reddit posts to create chatbot

bbc.com

131–140 of 235 posts

Re: Facebook uses 1.5B Reddit posts to create chatbot

#134
post #3

> Numerous issues arose during longer conversations. Blender would sometimes respond with offensive language, and at other times it would make up facts altogether. I mean, to be fair, I've had many conversations like that...

What did they expect training it on Reddit posts? Seems to me the bot is working fine.

Risky chat.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#135
post #124
post #86

Earlier quoted context omitted.

That's definitely a fair concern. We believe that open science and transparency are the right approach here. By releasing it, we ensure that everyone is on the same page with respect to abilities and defense.

So what's the defense?

that's the rub.

every AI sound/word/picture editor i've ran into says something along the lines of "we're releasing this data set to help stay secure in this day and age of easy counterfeiting of X.", but they never really mention how you apply the data in an adversarial way against itself -- they just sort of hand-wave that part.

Same with fake AI generated Obama video and sound, and earlier data-set generated chatbots; it's plastered all over the projects things like "Since these methods are available we think that it's important that this data is disseminated so that other's can use it to validate real world data sources", but again -- how?

We have the real data, we have the fake data -- how is this diff done, exactly?

I'm willing to bet it isn't as easy as all the AI researchers who release this stuff claim it may be.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#137
post #126

Earlier quoted context omitted.

Gentle reminder to those who may not know - you can remove your Reddit comments but you're not able to remove your HN comments. Food for thought!

Yes and no. You're right that there's no button you can push to delete an entire account history, but wrong that there's no way to remove HN comments. We take care of deletion requests for people every day. We don't want anyone to get in trouble from anything they posted to HN, there's nearly always something we can do, and we don't send people away empty-handed. I can only think of one or two cases where we weren't…

I get the reasoning, but I don't see this applied to some platforms. Reddit and Discord allow you to both delete and edit older comments, and there's no limits on how far back you can go (so you can, if you wanted, edit or delete your entire history).

Under the GDPR a subject is allowed full erasure rights. If I say I want you to delete my content from x date to y date, or a particular post, or everything entirely then that shouldn't be an issue. A request may be bothersome, but that's what happens when you don't offer that functionality natively.

I noticed a few days back you didn't like it when a user made a new account, except with the internet these days and how everything is archived for all time, throwaway's are the only option. Building a comment history is extremely dangerous, especially when you might forget what details you may have posted or how meta-data can leak through (such as what subs you post in, any details you posted that could identify you etc).

You can't have it both ways: no to multiple accounts and also no to control over your data. I might have 50 accounts, dislike it? Give me proper control over my comments. (to be honest, it may just be worth making a new account for every comment for maximum privacy, it's extreme, but it's a viable option).

If I want to delete them, that's my choice to freely make. Your thoughts or concerns are not relevant to me, thankfully, the GDPR agrees.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#138

Earlier quoted context omitted.

> By releasing it, we ensure that everyone is on the same page with respect to abilities and defense. This seems like wishful thinking. Having knowledge and having the resources to do something with it are two very different things.

Defending against such an "attack" is much easier if the technology is widely available and many people can play around with it and explore the limits.

So is crafting such an attack. Given that the attacks are obvious, and I haven't seen much word on defenses, the result seems inevitable.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#139
post #76
post #2

How can facebook turn a dumpster fire like reddit in a bot that response with more empathy than a human? Didn't Facebook just merge all fb messenger and whatsapp data and trained a NN on the new chat db?

Hi, paper author here. The model was fine-tuned on (non-Reddit) data that was specifically designed to give it positive conversational skills, like empathy, knowledge, and personality (see: https://arxiv.org/abs/2004.13637 ). No FB data was used to train these models, which is what allowed us to open source it.

I don't understand that last bit. Can you expand a bit please? Earlier comments in this thread say the copyright on a model is a grey area and can be classified as fair use

https://news.ycombinator.com/item?id=23094601

Post reply on HN