Live data from Hacker News

Facebook uses 1.5B Reddit posts to create chatbot

bbc.com

111–120 of 235 posts

Re: Facebook uses 1.5B Reddit posts to create chatbot

#111
post #92

Earlier quoted context omitted.

I was really blown away by the results you achieved. Amazing work! My jaw hit the floor when I saw the witty farewell "fun guy" quip, and I was in stitches when I read the song about baking. I look forward to the day I can take the model for a spin - unfortunately I don't have the requisite $18,000 hardware ;) I have a few questions: Could this be used as a tool to get a feel for public sentiment? For example, could…

I would not encourage using the model for anything other than AI research -- we're still in the early days of dialogue, and there are a lot of unexplored avenues. There are still nuances around safety, controlling generation, consistency, and knowledge involvement. For instance, the bot cannot remember what you said even a few turns ago, due to limitations in memory size. In the paper, we did explore what happens whe…

Was the bot nonsensical without the fine tuning, or just subjectively a worse conversational partner?

Re: Facebook uses 1.5B Reddit posts to create chatbot

#113
post #12

Earlier quoted context omitted.

Thank you. I can't really fathom why the BBC would not think to link to the _actual_ source of this news.

The BBC, like a lot of news orgs, wants to maximize advertising impressions, and the way to do that is to keep almost all links pointed to itself. A link to the more substantive source is a reader lost.

>The BBC, like a lot of news orgs, wants to maximize advertising impressions

the BBC is to a large degree publicly funded and a public service broadcaster, and advertisements are only shown outside of the United Kingdom. IIRC over 75% of their funding comes from British license fees, most of the rest from licensing their content to third parties outside of the UK under a separate commercial branch.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#114

Earlier quoted context omitted.

That's interesting! Did Facebook ask permission to create derivative works (the bot) from Reddit posts, I wonder, or does this fall under web-scraping law? If I recall Reddit users still retain rights to their posts unless Reddit the company provides some sort off broad grants? If they did not, this is an interesting example a company potentially making a great deal of money (if the bot is sold as something) from con…

Making or not making money is such a weird way for people to see things. That's part of why I love the Free Software movement so much and abhor the CC-*-NC licences. Fortunately, Reddit has the exception where they can give out access to anyone they want. But I still think StackOverflow is the gold standard: CC-BY-SA. No restriction on making money. Maybe a platinum standard would be CC-BY.

The point is not about the money - the point is using data contributed by users without the proper license to create something that might yield revenue which will then not be shared or payed forward in any way to the contributors. We have all worked hard to create the data used by companies to sell ads to us and make massive amounts of money. I guess I got a couple gigs of free email? Cool...

I also understand that most apps make us sign our lives away, but if I don't (as in the Reddit case) and I actually have rights to the data I sure as heck don't want that data used ANYWAY to power more of this stuff.

Probably a gross overreaction, but it seems like an externality that we've kinda just accepted as society that I'd like to see change a bit.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#115
post #71

Blog post: https://ai.facebook.com/blog/state-of-the-art-open-source-ch... Paper: https://arxiv.org/pdf/2004.13637.pdf Open Source: https://parl.ai/projects/recipes/ Ask us anything, the Facebook team behind it is happy to answer questions here.

Would FB possibly put some instances up for people to chat with?

It's way too heavy/expensive to host as an individual.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#116

Earlier quoted context omitted.

To even have a snowball's chance at success, they would have had to make use of reddit's voting system. Tons of toxicity and disinformation still makes it up into highly upvoted comments, but I'd expect throwing away heavily downvoted comments to exclude a good fraction of the utter crap.

downvotes are not necessarily indicative of bad ideas or comment but more about alignment to each specific sub-reddit groupthink.

I was thinking of a possible way to improve the downvote issue. Make users either comment or upvote an existing child comment to downvote.

I'm sure you'll get tons of "u suk" comments but there's just as many who won't even bother since they need to do two things now.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#117
post #10

How many of these reddit posts were themselves bot posts? Is it bots all the way down now?

the bots teaching the bots. (teacher bots for kid bots :O). the literal definition of machine learning

Sounds like a Matrix reference.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#118
post #104

A reminder that you can obtain the majority of Reddit posts/comments via BigQuery (via Pushshift). No need to write your own scraper. https://console.cloud.google.com/bigquery?p=fh-bigquery&d=re... https://console.cloud.google.com/bigquery?p=fh-bigquery&d=re... It appears to be roughly up to August 2019 for posts, October 2019 for comments.

There's one for HN too, or at least used to be: https://bigquery.cloud.google.com/dataset/bigquery-public-da...

Gentle reminder to those who may not know - you can remove your Reddit comments but you're not able to remove your HN comments.

Food for thought!

Re: Facebook uses 1.5B Reddit posts to create chatbot

#119

Earlier quoted context omitted.

Making or not making money is such a weird way for people to see things. That's part of why I love the Free Software movement so much and abhor the CC-*-NC licences. Fortunately, Reddit has the exception where they can give out access to anyone they want. But I still think StackOverflow is the gold standard: CC-BY-SA. No restriction on making money. Maybe a platinum standard would be CC-BY.

The point is not about the money - the point is using data contributed by users without the proper license to create something that might yield revenue which will then not be shared or payed forward in any way to the contributors. We have all worked hard to create the data used by companies to sell ads to us and make massive amounts of money. I guess I got a couple gigs of free email? Cool... I also understand that m…

In Reddit's case, that's the deal. You get a website to share things on with other people, and the value exchange involves you giving full licence to Reddit and giving relicense rights to Reddit.

Personally, I find that a very fair deal and clearly other people do as well. I think it actually yields positive externalities because we get things that wouldn't exist otherwise because the transaction costs outweigh the value, but the transaction costs are an inherent cost and I don't want to levy them. Fortunately, Reddit gives me the ability to not levy them and to guarantee that I won't levy them.

In fact, this is part of the magic of Free Software: true freedom to use. Yes, Google can use so much work which was done and it doesn't have to pay any of it back to Torvalds or Greg Kroah-Hartman or even me for the minor changes I made to libraries. This is freedom. I prefer it. And fortunately the world is aligned in this direction.

Re: Facebook uses 1.5B Reddit posts to create chatbot

#120
post #20

I fed the script to M-x doctor, and had a nice chat. --- I am the psychotherapist. Please, describe your problems. Each time you are finished talking, type RET twice. Hi how are you today? How do you do? What brings you to see me? Doing well. My favorite food is cake. I just bought one because I got promoted at work! Is it because you got promoted at work that you came to me? Thanks so much, I just want to make my pa…

The example in the article had the same environmental engineer question, is it just memorizing responses and spitting them back?

Perhaps it has a narrow “on-ramp” to get started? This example certainly does not paint it in a very good light.

I remember the text-based adventure example from a few months ago seemed both more interesting and immersive, and certainly more artistic.

Post reply on HN