Live data from Hacker News

Cloudflare to introduce pay-per-crawl for AI bots

blog.cloudflare.com

91–100 of 317 posts

Re: Cloudflare to introduce pay-per-crawl for AI bots

#91

This seems like it’s going about things in entirely the wrong way. What this does is say “okay, you still do all the work of crawling, you just pay more now”. There’s no attempt by Cloudflare to offer value for this extra cost. Crawling the web is not a competitive advantage for any of these AI companies, nor challenger search engines. It’s a cost and a massive distraction. They should collaborate on shared infrastru…

I am not sure how or why you are throwing shade at cloudflare. Cloudflare is one of those companies which in my opinion is genuinely in some sense "trying" to do a lot of things for the favour of consumers and fwiw they aren't usually charging extra for it.

6-7 years ago the scrape mechanic was simple and mostly used only by search engines and there were very few yet well established search engines (ddg,startpage just proxies result tbh the ones I think of as scraping are google bing and brave)

And these did genuinely value robots.txt and such because, well there were more cons than pros. Cons are a reputational hurt and just bad image in media tbh. Pros are what? "Better content?" So what. These search engines are on a lose basis model. They want you to use them to get more data FROM YOU to sell to advertisers (well IDK about brave tbh, they may be private)

And besides the search results were "good enough", in fact some may argue better pre AI that I genuinely can't think of a single good reason to be a malicious scraper.

Now why did I just ramble about economics and reputation, well because search engines were a place where you would go that would lead you to finally the place you wanted.

Now AI has become the place you go that would directly answer. And AI has shifted economics in that manner. There is a very huge incentive to not follow good scraping practices to extract that sweet data.

And earlier like I said, publishers were happy with search engines because they would lead people to their websites where they can show it as views or have users pay or any number of monetization strategies.

Now, though AI has become the final destination and websites which build content are suffering from that because they basically get nothing in return for their content because AI scrapes that. So, I guess now we need a better way to solve the evil scrapers.

Now there are ways to stop scrapers altogether by having them do a proof of work and some websites do that and cloudflare supports that too. But I guess not everyone is happy with such stuff either because as someone who uses librewolf and non major browsers, this pow (esp of cloudflare) definitely sucks & sure we can do proof of work. There's Anubis which is great at it.

But is that the only option? Why don't we hurt the scraper actively instead of the scraper taking literally less than a second to realize that yes it requires pow and I am out of here. What if we can waste the "scrapers time"

Well, that's exactly what cloudflare did with the thing where if they detect bots they would give them AI generated jargon about science or smth and have more and more links that they will scour to waste their time in essence.

I think that's pretty cool. Using AI to defeat AI. It is poetic and one of the best HN posts I ever saw.

Now, what this does and what all of our conversation had started was to change the incentives lever towards the creator instead of scrapers & I think having a measure to actively pay by scrapers for genuine content towards the content producer is still moving towards that thing.

Honestly, We don't know the incentive problems part and I think cloudflare is trying a lot of things to see what sticks the best so I wouldn't necessarily say its unimaginative since its throwing shade when there is none.

Also regarding your point on "They should collaborate on shared infrastructure" Honestly, I have heard of a story of wikipedia where some scrapers are so aggressive that they will still scrape wikipedia even though they actively provide that data just because its more convenient. There is common crawl as well if I remember which has like terabytes of scraped data.

Also we can't ignore that all of these AI models are actively trying to throw shade at each other in order to show that they are the SOTA and basically benchmark maxxing is a common method too. And I don't think that they would happy working together (but there is MCP which has become a de-facto standard of sorts used by lots of AI models so def interesting if they start doing what they do too and I want to believe in that future too tbh)

Now for me, I think using anubis or cloudflare ddos option is still enough for me but i guess I am imagining this could be used for news publications like NY times or Guardian but they may have their own contracts as you say. Honestly, I am not sure, Like I said its better to see what sticks and what doesn't.

Re: Cloudflare to introduce pay-per-crawl for AI bots

#92

That sounds reasonable for access to actual content, but it produces a huge new incentive to constantly produce vast amounts of AI-generated slop served via Cloudflare. Is there a way to disincentivize this?

Thats a more general problem. As content gets cheaper to produce with AI, how do consumers discriminate between good content and slop. We already have this problem with youtube and twitter and reddit

Its interesting that the AI companies will now be on the other end of this issue

Re: Cloudflare to introduce pay-per-crawl for AI bots

#93

Earlier quoted context omitted.

Cloudflare has not used CAPTCHAs since 2023: https://blog.cloudflare.com/turnstile-ga/

To save a visit: they use turnstile, a captcha replacement. The checkbox with verify you are a human. I would call that a captcha, but it is debatable if a non-puzzle check is.

[deleted]

Re: Cloudflare to introduce pay-per-crawl for AI bots

#94
So we used to have this company that did good things for the internet... like usable search...

Now we have this company that does good things for the internet... like ddos protection, cdns, and now protecting us from "AI"...

How long will the second one last before it also becomes universally hated?

Re: Cloudflare to introduce pay-per-crawl for AI bots

#95

In theory, why not, in practice welcome to the world where neutrality of internet explode... Soon they could decide if your requests come from a specific company IP or networks, because you look suspicious... In addition, bot fighting was never supposed to be about blocking automatic users but about blocking abusers, like spammers and co. So now it means that bad actor can have a free pass if they pay (with stolen cr…

> Soon they could decide if your requests come from a specific company IP or networks, because you look suspicious...

They already are. You probably can't browse half the internet without Cloudflare's approval.

Re: Cloudflare to introduce pay-per-crawl for AI bots

#96
post #55
post #50

Earlier quoted context omitted.

The Web Monetisation protocol solved this by doing a prorata based on how long you spent on the page.

What if I save the page locally on my machine? Or archive it?

What if my phone rings and I forget to close the page? Or leave it open to read it later?

Plus as one of the parent comments said, I am not paying before I get an idea what I'm paying for.

Re: Cloudflare to introduce pay-per-crawl for AI bots

#97

Earlier quoted context omitted.

If you don't want payment, there is: https://anubis.techaro.lol/ Used by https://gcc.gnu.org/bugzilla/ for example. It is less annoying than CAPTCHA/Turnstile/whatever because the proof of work runs automatically.

See also (AFAIK most of these support JSless challenges out of the box): haproxy-protection, go-away, anticrawl

Anubis does too: https://anubis.techaro.lol/docs/admin/configuration/challeng...

Re: Cloudflare to introduce pay-per-crawl for AI bots

#98
I really like the idea that crawlers who are profiting should have to pay content owners/creators per crawl.

On principal though, I think Cloudflare doing this is just one more thing to create the perception that you can't put something on the internet unless it's through Cloudflare. This harms a transparent and decentralised web and makes selfhosting seem even less appealing to those who don't know any better.

This should be implemented as a web protocol with crypto though so anyone can charge bots without having to be Cloudflare fronted. Not really a fanboi of 99% of crypto stuff, but IMO, a purely technical, open and decentralised solution to this sort of problem was the crypto dream.

We can all guess the people who will make the most money off this, and one of them is Cloudflare. A bunch of the other winners probably also run some of the more aggressive crawlers.

Re: Cloudflare to introduce pay-per-crawl for AI bots

#99

What about if somebody uses artificial intelligence crawler to help them navigate the web as an accessibility tool? Enabling UI automation. It already throws up a lot of... uh... troublesome verifications.

We already have ARIA, which is far more deterministic and should already be present on all major sites. AI should not be used, or necessary, as an accessibility tool.

Re: Cloudflare to introduce pay-per-crawl for AI bots

#100

It’s a step in the right direction but I think there’s a long ways to go. Even better would be pay-for-usage. So if you want to crawl a site for research, then it should be practically free, for example. If you want to crawl a site to train a bot that will be sold then it should cost a lot. I am truly sorry to even be thinking along these lines, but the alternative mindset has been made practically illegal in the mod…

> I would 100% be fine with there being a world library that strives to provide access to any and all information for free, while also aiming to find a fair way to compensate ip owners… technology has removed most of the technical limitations to making this a reality AND I think the net benefit to humanity would be vastly superior to the cartel approach we see today.

I can't help but wonder if this isn't actually true. As you've noted, if there's a system where it's 100% free to access and share information, then it's also 100% free to abuse such a system to the point of ruining it.

It seems the biggest limitations aren't actually whether such a system can technically be built, but whether it can be economically sustainable. The effect of technology removing too many barriers at once is actually to create economic incentives that make such a system impossible, rather than enabling such a system to be built.

Maybe there's an optimal amount level of information propagation that maximizes useful availability without shifting the equilibrium towards bots and spam, but we've gone past it. Arguably, large public libraries were just as close to that as using the Internet as a virtual library, I think.

I've explored this elsewhere through an evolutionary lens. When the genetic/memetic reproduction rate is too high, evolution creates r-strategists— Spamming lots of low-quality offspring/ideas that cannibalize each other, because it doesn't cost anything to do so. Adding limits actually results in K-strategists, incentivizing cooperation and investment in high-quality offspring/ideas because each one is worth more.

Post reply on HN