Live data from Hacker News

Cloudflare's new marketplace lets websites charge AI bots for scraping

techcrunch.com

131–140 of 280 posts

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#131
post #42
post #11

Cloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection rack…

As an actual content provider I see this as an opportunity. We pay our journalists real money to write real stories. If AI results haven't started affecting our search traffic they will start to soon. Up until now we've had two choices: block AI-based crawlers and fall completely out of that market, or continue to let AI companies train off of our hard-won content and take it as a loss that still generates a little b…

> If AI results haven't started affecting our search traffic they will start to soon. Up until now we've had two choices: block AI-based crawlers and fall completely out of that market, or continue to let AI companies train off of our hard-won content and take it as a loss that still generates a little bit of traffic

You have another option, one that iFixit chose: poison[1] the data sent to AI crawlers, you may even use GenAI to generate the fake content for maximum efficiency.

1. https://www.ifixit.com/Guide/Data+Connector++Replacement/147...

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#132

This seems like a gimmick. Isn't preventing crawling a sisyphean task? The only real difference this will make is further entrenching big players who have already crawled a ton of data. And if this feature comes at the cost of false positives and overbearing captchas, it will start to affect users.

Companies have been trying and failing to prevent large scale crawling for 25 years. It’s a constant arms race and the scrapers always win. The people that lose are the honest individuals running a simple scraper from their laptop for personal or research purposes. Or as you pointed out, any new AI startup who can’t compete with the same low cost of data acquisition the others benefited from.

> The people that lose ...

are also everyone who makes (literally) any effort in the direction of digital privacy, whose internet experience is degraded and frustrating due to increasingly bad captchas or just outright refusal of service.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#133
post #83

Common Crawl is shown in their screen shot of "Providers" along side OpenAI and Antropic. The challenge is that Common Crawl is used for a lot of things that are not AI training. For example, it's a major source of content for the Wayback machine. In fact, that's the entire point of the Common Crawl project. Instead of dozens of companies writing and running their (poorly) designed crawlers and hitting everyone's sit…

> gatekeep access to those who pay and those who don't, and that applied whether they are bots or people. I'm already constantly being classified as bot. Just today: To check if something is included in a subscription that we already pay for, I opened some product page on the Microsoft website this morning. Full-page error: "We are currently experiencing high demand. Please try again later." It's static content but i…

Likely you're in a blocked IP address range.

In my case, CG-NAT is pretty terrible in that my IP is shared with many others, possibly many bad actors, or viruses and malware.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#134
post #95

Earlier quoted context omitted.

Either AI training is fair use or it isn't. If it's fair use then businesses shouldn't get a say in whether the data can be used for it. If it isn't, then the answer to your question is copyright law. Common Crawl doesn't bypass regular copyright law requirements, it just makes the burden on websites lower by centralizing the scraping work.

Copyright is only part of the equation, there's also the use of other people's resources If what a government receptionist says is copyright-free, you still can't walk into their office thousands of times per day and ask various questions to learn what human answers are like in order to train your artificial neural network The amount of scraping that happened in ~2020 as compared to 2024 is orders of magnitude differ…

> you still can't walk into their office thousands of times per day

why not?

Esp. if that receptionist is an automaton, and isn't bothered by you. Of course, if you end up taking more resources and block others from asking as well, then you need to observe some etiquette (aka, throttle etc).

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#135
post #128
post #44

Earlier quoted context omitted.

Licensing doesn't mean shit when no court in the country is actually willing to prosecute violations. Who have OpenAI, Anthropic, Microsoft, Google, Meta licensed all their training data from?

Copyright infringement is a civil matter.

And where do you think civil matters are handled?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#136
post #119
post #103

Earlier quoted context omitted.

> because the bot operators have unlimited money I rather think the cause is that inbound bandwidth is usually free, so they need maybe 1/100th of the money because requests are smaller than responses (plus discounts they get for being big customers)

> I rather think the cause is that inbound bandwidth is usually free, so they need maybe 1/100th of the money because requests are smaller than responses (plus discounts they get for being big customers) Seems like there's the potential to take advantage of this for a semi-custom protocol, if there's a desire to balance costs for serving data while still making things available to end users. We'd have the server repl…

May work for small pages, like most of my webpages besides some downloadable files, but megabytes of JavaScript on an average (mobile?) connection are going to take very significantly longer to load, cost more battery, and take twice as much from your data bundle

Perhaps it's effective as bot deterrent when someone incurs, say, a ten times higher than median load (as measured in something like CPU time per hour or bandwidth per week or so). It will not prevent anyone from seeing your pages so information is still free, but it levels the playing field -- at least, for those with free inbound bandwidth dealing with bots that pay for outgoing bandwidth

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#137
post #134
post #95

Earlier quoted context omitted.

Copyright is only part of the equation, there's also the use of other people's resources If what a government receptionist says is copyright-free, you still can't walk into their office thousands of times per day and ask various questions to learn what human answers are like in order to train your artificial neural network The amount of scraping that happened in ~2020 as compared to 2024 is orders of magnitude differ…

> you still can't walk into their office thousands of times per day why not? Esp. if that receptionist is an automaton, and isn't bothered by you. Of course, if you end up taking more resources and block others from asking as well, then you need to observe some etiquette (aka, throttle etc).

> why not? Esp. if that receptionist is an automaton, and isn't bothered by you

I chose "thousands" to keep it within the realm of possibility while making it clear that it would bother a human receptionist precisely because humans aren't automatons, making the use of resources very obvious.

If you need an analogy to understand how an automated system could suffer from resources being consumed, perhaps picture a web server and billions of requests using a certain amount of bandwidth and CPU time each. Wait, now we're back to the original scenario!

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#138
post #126
post #123

Earlier quoted context omitted.

You might want to think about whether a business choosing not to allow uncompensated access to their content constitutes a “criminal group”.

Don’t put your stuff on the internet then, or put it behind a paywall/registration.

So … it’s okay if they build their own system but you find it upsetting when they pay Cloudflare for a service?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#139

Common Crawl is shown in their screen shot of "Providers" along side OpenAI and Antropic. The challenge is that Common Crawl is used for a lot of things that are not AI training. For example, it's a major source of content for the Wayback machine. In fact, that's the entire point of the Common Crawl project. Instead of dozens of companies writing and running their (poorly) designed crawlers and hitting everyone's sit…

> There are significant knock-on effects

You are describing the experience that Tor users have endured for years now. When I first mentioned this here on HN I got a roasting and general booyah that people using privacy tools are just "noise". Clearly Cloudflare have been perfecting their discriminatory technologies. I guess what goes around comes around. "first they came for the...." etc etc.

Anyway, I see a potential upside to this, so we might be optimistic. Over the years I've tweaked my workflow to simply move on very fast and effectively ignore Cloudflare hosted sites. I know... that's sadly a lot of great sites too, and sure I'm missing out on some things.

On the other hand, it seems to cut out a vast amount of rubbish. Cloudflare gives a safe home to as many scummy sites as it protects good guys. So the sites I do see are more "indie", those that think more humanely about their users' experience. Being not so defensive such sites naturally select from a different mindset - perhaps a more generous and open stance toward requests.

So what effect will this have on AI training?

Maybe a good one. Maybe tragic. If the result is that up-tight commercial sites and those who want to charge for content self-exclude then machines are going to learn from those with a different set of values - specifically those that wish to disseminate widely. That will include propaganda and disinformation for sure. It will also tend to filter out well curated good journalism. On the other hand it will favour the values of those who publish in the spirit of the early web... just to put their own thing up there for the world.

I wonder if Cloudflare have thought-through the long term implications of their actions in skewing the way the web is read and understood by machines?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#140
Boy I'm sick of clicking "Verify you are human" on everything from GitLab to banking apps running Cloudflare.

Sick enough that I hope someone prominent at the EFF or similar takes Cloudflare to court over it.

One company shouldn't be allowed to police access to the internet. And certainly shouldn't be allowed to start gatekeeping what is viewable by discriminating against the person or software doing the viewing.

I worry that Cloudflare will keep escalating this unless they're sent a strong signal that it's not supported by the tech community. If you work there, it might be time to consider getting a different job. If you own stock, maybe divest. If you're connected, perhaps your associates can buy from competitors. That's probably the only way to get the board and CEO replaced these days.

Post reply on HN