I was recently speaking with people from OpenFoodFacts and OpenStreetMap, and I guess Wikipedia as the same issue. They are under constantly DDoS by bots which are scraping everything, even if the full dataset can be downloaded for free with a single HTTP request. They said this useless traffic was a huge cost for them. This is not about copyright, just about bots being stupid and people behind them not caring at all…
I’ve just taken to blocking entire swaths of cloud services IP networks. I don’t care what the intentions are, my personal sites don’t get the infinite bandwidth to put up with a thousands of poorly written spiders.
Cloudflare's new marketplace lets websites charge AI bots for scraping
121–130 of 280 posts
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#122Earlier quoted context omitted.
The problem is that soft technical measures like HTTP 402 and robots.txt aren't legally binding, so there's nothing stopping scrapers from just ignoring them. Cloudflares value proposition here is they will play the cat-and-mouse game of detecting things like spoofed user agents and residential proxies on your behalf, and actively block what appears to be scraper traffic unless they pay up. Unfortunately this probabl…
Sure it's not legally binding, but if I see >100000 requests coming from 1 IP address within a week, I'm also not legally bound to make that 402 error go away. By having an automated payment mechanism, the two parties could come to an agreement they're both happy about > there's nothing stopping scrapers from just ignoring them Feel free to ignore HTTP errors, but those pages don't contain the content you're looking…
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#123Earlier quoted context omitted.
So Cloudflare now wants to collect money to not block people. Is that about the gist of it?
> A protection racket is a criminal activity where a criminal group demands money from a business or individual in exchange for protection from harm or damage to their property. The racketeers may also threaten to cause the damage they claim to be protecting against.
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#124Cloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection rack…
As an actual content provider I see this as an opportunity. We pay our journalists real money to write real stories. If AI results haven't started affecting our search traffic they will start to soon. Up until now we've had two choices: block AI-based crawlers and fall completely out of that market, or continue to let AI companies train off of our hard-won content and take it as a loss that still generates a little b…
You make it sound like this is OK. "It's not their fault that a protection racket didn't already exist. They just filled the market's need for one."
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#125Earlier quoted context omitted.
Lately I’ve been noticing captchas have been increasingly difficult day by day on Firefox. Checking the box use to go through without issue, but now it’s been starting to pop up challenges with the boxes that fade after clicking. Just like your experience, chrome has no hiccups on the same machine.
Those "keep clicking until we stop fading in more results" challenges mean they're fairly confident you're a bot and this is the highest difficulty level to prove your lack of guilt. I get these only when using a browser that isn't already full of advertising cookies (edit: which, to be clear, I hope is still considered an acceptable state to have your browser in)
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#126Earlier quoted context omitted.
> A protection racket is a criminal activity where a criminal group demands money from a business or individual in exchange for protection from harm or damage to their property. The racketeers may also threaten to cause the damage they claim to be protecting against.
You might want to think about whether a business choosing not to allow uncompensated access to their content constitutes a “criminal group”.
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#127but if they are only tracking the bot via the user agent
then can't i piggyback on that user agent?
no ai scraper is going to include an auth header when accessing your website...
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#128Earlier quoted context omitted.
Licensing. Common Crawl could change the license of how the data it produces is used. Common Crawl already talks about allowed use of the data in their FAQ, and in their terms of use: https://commoncrawl.org/terms-of-use/ https://commoncrawl.org/faq While this doesn't currently discuss AI, they could. This would allow non-AI downstream consumers to not be penalized.
Licensing doesn't mean shit when no court in the country is actually willing to prosecute violations. Who have OpenAI, Anthropic, Microsoft, Google, Meta licensed all their training data from?
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#129Earlier quoted context omitted.
Lately I’ve been noticing captchas have been increasingly difficult day by day on Firefox. Checking the box use to go through without issue, but now it’s been starting to pop up challenges with the boxes that fade after clicking. Just like your experience, chrome has no hiccups on the same machine.
I wonder how many of those captchas are controlled by competitors of Firefox?
Re: Cloudflare's new marketplace lets websites charge AI bots for scraping
#130Earlier quoted context omitted.
Lately I’ve been noticing captchas have been increasingly difficult day by day on Firefox. Checking the box use to go through without issue, but now it’s been starting to pop up challenges with the boxes that fade after clicking. Just like your experience, chrome has no hiccups on the same machine.
Those "keep clicking until we stop fading in more results" challenges mean they're fairly confident you're a bot and this is the highest difficulty level to prove your lack of guilt. I get these only when using a browser that isn't already full of advertising cookies (edit: which, to be clear, I hope is still considered an acceptable state to have your browser in)
Those ones are the fucking worst. I've noticed that if I try to succeed in these captchas too quickly, it'll just say "Sorry, try again" even when every click was correct, so instead, I've started going in slow motion and faking "misclicking" which makes it much more likely to accept me as human.
I cannot stand the idea that I have to pretend to be slower than I am, in order for a computer to not think I'm a computer. Thanks CloudFlare and Google.