Live data from Hacker News

Google cut a deal with Reddit for AI training data

theverge.com

121–130 of 154 posts

Re: Google cut a deal with Reddit for AI training data

#121

Earlier quoted context omitted.

This worked amazingly for me when I stopped using Reddit.

Who knows if the "deleted" data can still be sold off. Sure you can't view it on the website, but nobody knows what Reddit will hand to Google. Then again certain laws passed recently might make such a thing illegal...

The second option I linked also replaces all comments and posts with random scrambled gibberish, not only "deleting" it. I would agree that deleted != actually deleted, but scrambling all content might actually work to prevent further siphoning

Re: Google cut a deal with Reddit for AI training data

#123
post #72

Earlier quoted context omitted.

Hate on HN all you want, I’ve been without my ADHD meds (warning: the company “Done” is not technically a scam) and spending way too much time on each for the past few days, and I can say this for sure: at least people on HN pay heed to the concept of premises and careful, non-combative argument. Most responses on Reddit are “no, that’s dumb” or “yes, that reminds me of my metaphysical takes”…

Not only that, reddit hive mind is plain wrong in most of the cases. Plus in number of occasions the "le reddit investigation", "we did it reddit" excrement caused real-world issues for people that they were targeting, and those people were innocent. Reddit is ok and quite cool for targeted discussion on targeted sub-reddits. But all the general subreddits visited by general population and everything that pops once i…

Try saying anywhere on Reddit "WD-40 is a lubricant" and be prepared to face a tsunami of incorrect information. Or say anything about glyphosate.

Re: Google cut a deal with Reddit for AI training data

#124
post #111

Reddit blocks all search engines from crawling their comments, essentially blocking all user generated contents. See this directive in https://www.reddit.com/robots.txt `Disallow: /r/ /comments/ / / /*` But I can still search Reddit on Google. How does Google manage to get the data?

I think you're misinterpreting the rule. The relevant robots.txt rules are: User-Agent: * Disallow: */comment/* Disallow: /r/*/comments/*/*/*/* The url for the comment sections looks like: https://www.reddit.com/r/[subreddit]/comments/[id]/[slug]/ . This doesn't match the above rules because there's only 2 parts after "comments", not 4 parts as specified by the rules.

Now it says "Disallow: /"

Re: Google cut a deal with Reddit for AI training data

#125

Earlier quoted context omitted.

What would you say is the agenda of Reddit (the company, not the users)?

Any ideas that make you uncomfortable are an "agenda" now.

Attempting to rewrite history and racism regardless of the race involved makes me uncomfortable yes. Do you disagree that Google had an agenda when they programmed their AI to do the things it does?

Re: Google cut a deal with Reddit for AI training data

#126

Earlier quoted context omitted.

A. Nothing said it was exclusive. B. Reddit's value is in the community more than the data. A company with Reddit's data but not Reddit's users is worth relatively little.

How do you extract the value from the community?

Well, I've known several people over the years who have created a multitude of accounts for the purpose of selling them when they reach around 100K karma to various activist and political groups. He didn't care what they were doing with the accounts, but it seems there's been a black market for user data and user profiles with high Karma counts for a long time now.

The only things I can think of Google using that data for are for nefarious purposes.

Re: Google cut a deal with Reddit for AI training data

#128
post #29

$60M/year for GOOG to access all their data when they purport to be targeting a $5B valuation at IPO is really cheap. Arguably Reddit's value is it's data, and GOOG is renting it for 1.2%/year?

1. It is pointless to compare total company value with an annual payment. 2. It isn't an exclusive license. There are dozens of companies training language models, and if they are all forced to pay the same amount then that's some serious cash.

Companies are generally valued based on cash flows, and some multiples based on profit/growth/etc.

Re: Google cut a deal with Reddit for AI training data

#129
When I first learned that AI companies are vacuuming up all internet content without any regard for permission, attribution or compensation for the content creator, I found that deeply immoral.

I figured they should pay for it. But now that they do (in this instance), I'm realizing this might be even worse. They can just buy the entire market in the same way they buy Google search users by paying Apple billions a year.

And still the actual content creator, a Reddit user in this case, is not compensated.

It's truly wild how lax regulation is. This is probably the most important technology ever created and we just let 2 companies have it all: the data and the compute.

Re: Google cut a deal with Reddit for AI training data

#130

> Google Search is currently expanding the test of a "forums" filter that lets you browse through results from sites with human discussion, like Reddit Yeah, about that whole "human discussion" thing...

Kagi already has such a filter by default, although I've never tested it.
Post reply on HN