Live data from Hacker News

Cloudflare's new marketplace lets websites charge AI bots for scraping

techcrunch.com

121–130 of 280 posts

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#121
post #107

I was recently speaking with people from OpenFoodFacts and OpenStreetMap, and I guess Wikipedia as the same issue. They are under constantly DDoS by bots which are scraping everything, even if the full dataset can be downloaded for free with a single HTTP request. They said this useless traffic was a huge cost for them. This is not about copyright, just about bots being stupid and people behind them not caring at all…

I’ve just taken to blocking entire swaths of cloud services IP networks. I don’t care what the intentions are, my personal sites don’t get the infinite bandwidth to put up with a thousands of poorly written spiders.

Is there a public list of those address blocks, which you'd recommend?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#122
post #104
post #10

Earlier quoted context omitted.

The problem is that soft technical measures like HTTP 402 and robots.txt aren't legally binding, so there's nothing stopping scrapers from just ignoring them. Cloudflares value proposition here is they will play the cat-and-mouse game of detecting things like spoofed user agents and residential proxies on your behalf, and actively block what appears to be scraper traffic unless they pay up. Unfortunately this probabl…

Sure it's not legally binding, but if I see >100000 requests coming from 1 IP address within a week, I'm also not legally bound to make that 402 error go away. By having an automated payment mechanism, the two parties could come to an agreement they're both happy about > there's nothing stopping scrapers from just ignoring them Feel free to ignore HTTP errors, but those pages don't contain the content you're looking…

I mean it's not legally binding in the sense that if you start sending 402s or 403s to a scraper it can just take that as a signal to try again from a different IP address until it works - your servers clearly stated intent that the bot should pay up or go away isn't legally actionable. With enough effort you can chase the bots until they run out of resources, but few people have time to win that battle by themselves, hence delegating it to Cloudflare or similar.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#123

Earlier quoted context omitted.

So Cloudflare now wants to collect money to not block people. Is that about the gist of it?

> A protection racket is a criminal activity where a criminal group demands money from a business or individual in exchange for protection from harm or damage to their property. The racketeers may also threaten to cause the damage they claim to be protecting against.

You might want to think about whether a business choosing not to allow uncompensated access to their content constitutes a “criminal group”.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#124
post #42
post #11

Cloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection rack…

As an actual content provider I see this as an opportunity. We pay our journalists real money to write real stories. If AI results haven't started affecting our search traffic they will start to soon. Up until now we've had two choices: block AI-based crawlers and fall completely out of that market, or continue to let AI companies train off of our hard-won content and take it as a loss that still generates a little b…

> don't blame the player, blame the game

You make it sound like this is OK. "It's not their fault that a protection racket didn't already exist. They just filled the market's need for one."

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#125
post #102

Earlier quoted context omitted.

Lately I’ve been noticing captchas have been increasingly difficult day by day on Firefox. Checking the box use to go through without issue, but now it’s been starting to pop up challenges with the boxes that fade after clicking. Just like your experience, chrome has no hiccups on the same machine.

Those "keep clicking until we stop fading in more results" challenges mean they're fairly confident you're a bot and this is the highest difficulty level to prove your lack of guilt. I get these only when using a browser that isn't already full of advertising cookies (edit: which, to be clear, I hope is still considered an acceptable state to have your browser in)

FWIW, it can't be cookies alone that gives you an inordinate number of bot challenges. I use private tabs on Firefox (for Linux and Android) for most of my browsing, and I rarely get any challenges regardless of what I do. The only issues tend to be when I make repeated searches for things with "quotes" and whatnot on Google or on Stack Exchange sites. But for the most part, those challenges aren't particularly drawn-out: I've only ever gotten the "fading" ones when I'm using Tor or a VPN.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#126
post #123

Earlier quoted context omitted.

> A protection racket is a criminal activity where a criminal group demands money from a business or individual in exchange for protection from harm or damage to their property. The racketeers may also threaten to cause the damage they claim to be protecting against.

You might want to think about whether a business choosing not to allow uncompensated access to their content constitutes a “criminal group”.

Don’t put your stuff on the internet then, or put it behind a paywall/registration.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#128
post #44

Earlier quoted context omitted.

Licensing. Common Crawl could change the license of how the data it produces is used. Common Crawl already talks about allowed use of the data in their FAQ, and in their terms of use: https://commoncrawl.org/terms-of-use/ https://commoncrawl.org/faq While this doesn't currently discuss AI, they could. This would allow non-AI downstream consumers to not be penalized.

Licensing doesn't mean shit when no court in the country is actually willing to prosecute violations. Who have OpenAI, Anthropic, Microsoft, Google, Meta licensed all their training data from?

Copyright infringement is a civil matter.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#129
post #105

Earlier quoted context omitted.

Lately I’ve been noticing captchas have been increasingly difficult day by day on Firefox. Checking the box use to go through without issue, but now it’s been starting to pop up challenges with the boxes that fade after clicking. Just like your experience, chrome has no hiccups on the same machine.

I wonder how many of those captchas are controlled by competitors of Firefox?

ReCAPTCHA absolutely hammers Firefox compared to Chrome for me. On sites that use it for login I rarely just get the "check the box" challenge anymore, and am instead being asked to train their CV algorithms by picking 5+ images of stoplights or motorcycles. Punishment for avoiding the Chrome universe I guess.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#130
post #102

Earlier quoted context omitted.

Lately I’ve been noticing captchas have been increasingly difficult day by day on Firefox. Checking the box use to go through without issue, but now it’s been starting to pop up challenges with the boxes that fade after clicking. Just like your experience, chrome has no hiccups on the same machine.

Those "keep clicking until we stop fading in more results" challenges mean they're fairly confident you're a bot and this is the highest difficulty level to prove your lack of guilt. I get these only when using a browser that isn't already full of advertising cookies (edit: which, to be clear, I hope is still considered an acceptable state to have your browser in)

> Those "keep clicking until we stop fading in more results" challenges mean they're fairly confident you're a bot

Those ones are the fucking worst. I've noticed that if I try to succeed in these captchas too quickly, it'll just say "Sorry, try again" even when every click was correct, so instead, I've started going in slow motion and faking "misclicking" which makes it much more likely to accept me as human.

I cannot stand the idea that I have to pretend to be slower than I am, in order for a computer to not think I'm a computer. Thanks CloudFlare and Google.

Post reply on HN