Live data from Hacker News

Cloudflare's new marketplace lets websites charge AI bots for scraping

techcrunch.com

41–50 of 280 posts

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#41
post #37

Earlier quoted context omitted.

[flagged]

> This kind of cynicism is boring. IMHO, this kind of thinking is only cynicism iff you're only looking for your angle to profit, and someone is peeing on your parade, every time they boorishly mention irrelevant, imaginary concerns like "ethics", "legality", or "Geneva Convention".

[flagged]

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#42
post #11

Cloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection rack…

As an actual content provider I see this as an opportunity. We pay our journalists real money to write real stories. If AI results haven't started affecting our search traffic they will start to soon. Up until now we've had two choices: block AI-based crawlers and fall completely out of that market, or continue to let AI companies train off of our hard-won content and take it as a loss that still generates a little bit of traffic. Cloudflare now offers a third option if we can figure out how to use it.

Dissing on Cloudflare is the new thing, and I get it. They're big and powerful and they influence a massive amount of the traffic on the web. Like the saying goes though, don't blame the player, blame the game. Ask yourself if you'd rather have Alphabet, Microsoft, Amazon or Apple in their place, because probably one of them would be.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#43
post #11

Cloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection rack…

I distinctly remember Cloudfare being accused here of hosting spammers and selling protection against them a decade ago. Then suddenly the name became associated with positive things only, and the whole thing have been memory-holed.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#44
post #35

Earlier quoted context omitted.

> Common Crawl runs once and exposes the data in industry standard formats like WARC for other consumers And what stops companies from using this data for model training? Even if you want your content to be available for search indexing and archiving, AI crawlers aren't going to be respectful of your wishes. Hence the need for restrictive gatekeeping.

Licensing. Common Crawl could change the license of how the data it produces is used. Common Crawl already talks about allowed use of the data in their FAQ, and in their terms of use: https://commoncrawl.org/terms-of-use/ https://commoncrawl.org/faq While this doesn't currently discuss AI, they could. This would allow non-AI downstream consumers to not be penalized.

Licensing doesn't mean shit when no court in the country is actually willing to prosecute violations. Who have OpenAI, Anthropic, Microsoft, Google, Meta licensed all their training data from?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#46
post #35

Common Crawl is shown in their screen shot of "Providers" along side OpenAI and Antropic. The challenge is that Common Crawl is used for a lot of things that are not AI training. For example, it's a major source of content for the Wayback machine. In fact, that's the entire point of the Common Crawl project. Instead of dozens of companies writing and running their (poorly) designed crawlers and hitting everyone's sit…

> Common Crawl runs once and exposes the data in industry standard formats like WARC for other consumers And what stops companies from using this data for model training? Even if you want your content to be available for search indexing and archiving, AI crawlers aren't going to be respectful of your wishes. Hence the need for restrictive gatekeeping.

Either AI training is fair use or it isn't. If it's fair use then businesses shouldn't get a say in whether the data can be used for it. If it isn't, then the answer to your question is copyright law.

Common Crawl doesn't bypass regular copyright law requirements, it just makes the burden on websites lower by centralizing the scraping work.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#47
post #45

What's wrong with AI agents accessing website content? We seem to have been happy with Google doing that for ages in exchange for displaying the website in search results.

And AI agents scrape your content in exchange for what exactly?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#48

This seems like a gimmick. Isn't preventing crawling a sisyphean task? The only real difference this will make is further entrenching big players who have already crawled a ton of data. And if this feature comes at the cost of false positives and overbearing captchas, it will start to affect users.

The risk of getting sued prevents companies from using pirated software.

The big players might just pay the fee because they might one day need to prove where they got the data from.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#49

This seems like a gimmick. Isn't preventing crawling a sisyphean task? The only real difference this will make is further entrenching big players who have already crawled a ton of data. And if this feature comes at the cost of false positives and overbearing captchas, it will start to affect users.

My website contains millions of pages. It's not hard to notice the difference between a bot (or network) that wants to access all pages and a regular user.
Post reply on HN