Live data from Hacker News

Cloudflare's new marketplace lets websites charge AI bots for scraping

techcrunch.com

61–70 of 280 posts

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#61
post #35

Earlier quoted context omitted.

> Common Crawl runs once and exposes the data in industry standard formats like WARC for other consumers And what stops companies from using this data for model training? Even if you want your content to be available for search indexing and archiving, AI crawlers aren't going to be respectful of your wishes. Hence the need for restrictive gatekeeping.

Either AI training is fair use or it isn't. If it's fair use then businesses shouldn't get a say in whether the data can be used for it. If it isn't, then the answer to your question is copyright law. Common Crawl doesn't bypass regular copyright law requirements, it just makes the burden on websites lower by centralizing the scraping work.

There is no objective black and white is or is not in this situation.

There is litigation of multiple cases and a judge making a judgement on each one.

Until then, and even after then, publishers can set the terms and enforce those terms using technical means like this.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#62
post #42
post #11

Cloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection rack…

As an actual content provider I see this as an opportunity. We pay our journalists real money to write real stories. If AI results haven't started affecting our search traffic they will start to soon. Up until now we've had two choices: block AI-based crawlers and fall completely out of that market, or continue to let AI companies train off of our hard-won content and take it as a loss that still generates a little b…

Not dissing any company; just pointing out a real concern to be considered, in this freshly disrupted and rapidly evolving environment.

We all know that someone is going to try to slip one past the regulators, and they're probably on HN, and we know from the past that this can pay off hugely for them.

Maybe, this time, the HN people who grumble about past exploiters and abusers in retrospect, can be more proactive, and help inform lawmakers and regulators in time.

And for those of us who don't want to be activists, but also don't want to be abusers -- just run honest businesses -- we're reminded to think twice about what we do and how we do it, when we're operating in what seems like novel space.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#63
post #59

Earlier quoted context omitted.

I don't believe that companies have the right to say that my user agent must run their ads. They can politely request that it does and I can tell my agent whether to show them or not.

True, but by the same measure your user agent can politely request a webpage and the server has the right to say 403 Forbidden. Nobody is required to play by the other parties rules here.

Exactly. The trouble is that companies want the benefits of being on the open web without the trade-offs. They're more than welcome to turn me down entirely, but they don't do that because that would have undesirable knock-on effects. So instead they try to make it sound like I have a moral obligation to render their ads.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#64
post #33
post #11

Cloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection rack…

> I'd need to ask the lawyer how close this is to technically being a protection racket, or other no-no. Wait 'til you find out how many of the DDoS-for-hire services that Cloudflare offers to protect you from are themselves protected by Cloudflare.

I hear this pretty often. I am curious what do you think Cloudfare should do?

I am pretty sure that if they started arbitrarily banning customers/potential customers based on what some other people like or don't like, everyone would be up in arms yelling stuff about censorship or wokeness or whatever the word of the year is.

As an example, what if I'm not a DDoS-for-hire, but just a website that sells some software capable of launching DDoS attacks? Should I be able to buy Cloudfare protection? Should a site like Metasploit be allowed to purchase protection?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#65
post #45

What's wrong with AI agents accessing website content? We seem to have been happy with Google doing that for ages in exchange for displaying the website in search results.

For traditional search indexing the interests of the aggregator and the content creator were aligned. AIs on the other hand are adversarial to the interest of content creators, a sufficiently advanced AI can replace the creator of the content it was trained on.

We're talking in this subthread about an AI agent accessing content, not training a model on content.

Training has copyright implications that are working their way through courts. AI agent access cannot be banned without fundamentally breaking the User Agent model of the web.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#66
post #35

Earlier quoted context omitted.

> Common Crawl runs once and exposes the data in industry standard formats like WARC for other consumers And what stops companies from using this data for model training? Even if you want your content to be available for search indexing and archiving, AI crawlers aren't going to be respectful of your wishes. Hence the need for restrictive gatekeeping.

Either AI training is fair use or it isn't. If it's fair use then businesses shouldn't get a say in whether the data can be used for it. If it isn't, then the answer to your question is copyright law. Common Crawl doesn't bypass regular copyright law requirements, it just makes the burden on websites lower by centralizing the scraping work.

Its not a legal question but a behavior and sustainability question. If it is fair use, but is undesirable for content makers, then they’re still not under any obligation to allow scraping. So they’ll try stuff like this, and other more restrictive bot blockers.

Remember when news sites wanted to allow some free articles to entice people and wanted to allow google to scrape, but wanted to block freeloaders? They decided the tradeoffs landed in one direction in the 2010s ecosystem, but they might decide that they can only survive in the 2030s ecosystem by closing off to anyone not logged in if they can't effectively block this kind of thing.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#67
post #60

Earlier quoted context omitted.

[flagged]

Interesting, as my theory for why cynicism is so common nowadays is that it's a coping mechanism for people who understand perfectly well what's happening in the modern world

You don't need to be a cynic if you have a grasp on reality If your truly understand something you are capable of evaluating it on a case by case basis without resorting to pathos.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#69

Earlier quoted context omitted.

For traditional search indexing the interests of the aggregator and the content creator were aligned. AIs on the other hand are adversarial to the interest of content creators, a sufficiently advanced AI can replace the creator of the content it was trained on.

We're talking in this subthread about an AI agent accessing content, not training a model on content. Training has copyright implications that are working their way through courts. AI agent access cannot be banned without fundamentally breaking the User Agent model of the web.

Ok, fine, let's restrict it to AI agents only, without training. It's still an adversarial relationship with the content creator. When you take an AI agent an ask it "find me the best italian restaurant in city xyz" it scans all the restaurant review sites and gives you back a recommendation. The content creator bears all the burden of creating and hosting the content and reaps non of the reward as the AI agent has now inserted itself as a middleman.

The above is also a much clearer / more obvious case of copyright infringement than AI training.

> AI agent access cannot be banned without fundamentally breaking the User Agent model of the web.

This is a non-sequitur but yes you are right, everything in the future will be behind a login screen and search engines will die.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#70

Earlier quoted context omitted.

We're talking in this subthread about an AI agent accessing content, not training a model on content. Training has copyright implications that are working their way through courts. AI agent access cannot be banned without fundamentally breaking the User Agent model of the web.

Ok, fine, let's restrict it to AI agents only, without training. It's still an adversarial relationship with the content creator. When you take an AI agent an ask it "find me the best italian restaurant in city xyz" it scans all the restaurant review sites and gives you back a recommendation. The content creator bears all the burden of creating and hosting the content and reaps non of the reward as the AI agent has n…

> reaps non of the reward

Just to be clear what we're talking about: the reward in question is advertising dollars earned by manipulating people's attention for profit, right?

I frankly don't think that people have the right to that as a business model and would be more than happy to see AI agents kill off that kind of "free" content.

Post reply on HN