Live data from Hacker News

Google is the only search engine that works on Reddit now, thanks to AI deal

404media.co

201–210 of 379 posts

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#201

Earlier quoted context omitted.

We did. As in we, the Internet, existed for a long time without anyone making money and we paid for the privilege. Websites were built and hosted at owner's expense, for years, with no expectation that they be financially rewarded. Sure some would run donation drives, or work with sponsors relevant to the community in question, but a whole ton, mine included, just cost me a lot of money over many years. Those website…

>Websites don't cost that much to run. Popular websites that allow user content to be uploaded or linked do cost that much to run, due to content moderation. There might be a small (relatively) forum here and there that a few good moderators are willing to slave away at keeping clean, but you will never see a website that allows user content with as many users as Reddit/Youtube/Instagram/etc be cheap. Although, due t…

Although it is quite surprising that mainly text websites (Reddit, Twitter) are hard to run sustainably but video and image websites (YouTube, Instagram, TikTok) can because it is easier to sell ads against them.

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#203

IANAL but as far as I understand the current legal status (in the US) a change in robots.txt or terms and conditions is not binding for web scrapers since the data is publicly accessible. Neither does displaying a banner "By using this site you accept our terms and conditions" change anything about that. The only thing that can make these kinds of terms binding is if the data is only accessible after proactively acce…

Quite sure they are also enforcing these with some technical measures to limit scraping.

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#204

This is an interesting development. How many other sites might have leverage to charge to be indexed? I don't want to live in a world where you have to use X search engine to get answers from Y site - but this seems like the beginning of that world. From an efficiency perspective - it's obviously better for websites to just lease their data to search engines then both sides paying tons of bandwidth and compute to get…

Kagi uses at least Google and Mojeek edit: > Realistically, there are only 2 search engines now. https://seirdy.one/posts/2021/03/10/search-engines-with-own-...

Doesn't it list three major ones, Google, Bing, and Yandex, plus Mojeek and a few other small ones? That's a bit more than two.

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#206
post #78
post #5

# Welcome to Reddit's robots.txt # Reddit believes in an open internet, but not the misuse of public content. # See https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy Reddit's Public Content Policy for access and use restrictions to Reddit content. # See https://www.reddit.com/r/reddit4researchers/ for details on how Reddit continues to support research and non-commercial use. # pol…

Nobody who wants to be successful obeys robots.txt. And I do mean nobody.

[dead]

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#207
post #197

Earlier quoted context omitted.

> # Reddit believes in an open internet, but not the misuse of public content. Calling it "public" content in the very act of exercising their ownership over it. The balls on whoever wrote that.

it's even worse. it's not theirs (it's the users'), they are merely hosting it and using it (ToS gives them a fancy irrevocable license I guess). so they can do whatever they want with it and the actual owners/authors have no chance to really influence Reddit at all to make it crawlable. (the GDPR-like data takeout is nice, but ... completely useless in these cases where the value is in the composition and aggregatio…

On top of that, a sizable chunk of Reddit content is ripped from elsewhere, whether videos, images, etc.

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#208

FWIW, we inquired to the reddit sales team about paying for data sometime last year, as we do similar elsewhere for use cases like helping emergency responders, and even though they were launching the program and asking for customers... no email back. Nor on our second and I think third attempt. I'm not sure what to make of that.

How much were you willing to pay? Still, rude of them not to even discuss the issue. Every time I've gone to buy data, if I'm too small of a fish, vendors have always been happy refer me to a reseller.

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#209
post #86

Boy, the LLMs have really been an apocalypse moment for the web, haven’t they? Between hoovering up and monetizing every bit of content they can without any attribution or compensation and the absolute flood of mediocre generated content, they’ve really done in the last straggling remains of the open internet. It’s not like everyone wasn’t already pulling the same grift, but quantity really does have a quality all it…

Of course, we have to be careful not to villainize a neutral tech. Instead let's call it what it is: unchecked capitalism and monopolistic behaviors.

Capitalism seems to work ok for the common good until you remove all the protections. LLMs provide a defacto monopoly for the owner which must already be a near monopoly: they take vast resources to train; only a giant corp can afford to buy all the content and provision enough resources to train one.

LLM did not enshittify what's left of the internet, greed did it.

Re: Google is the only search engine that works on Reddit now, thanks to AI deal

#210
post #78

Earlier quoted context omitted.

Nobody who wants to be successful obeys robots.txt. And I do mean nobody.

They changed it to disallow so that scrapers can't just claim the robots.txt gave them permission.

According to the US court systems the robots.txt file is meaningless. If they respond with a 200 status code giving you the access then you can legally scrape it all you want. If they require that you log in then you have to follow the terms you agree to when creating an account. Public means public though, and if Reddit doesn't want to make the content private (put it behind a login) then we can scrape away.

Note that scraping, regardless of the level of permission, doesn't mean you can do anything you want with the content. Copyright still applies. But you can scrape it, and if your use falls under Fair Use or another caveat to the copyright laws then you can do ahead and do it without needing any permission from the authors.

Post reply on HN