Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

481–490 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#481

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

> 1. If I as a human request a website, then I should be shown the content. Everyone agrees. I disagree. The website should have the right to say that the user can be shown the content under specific conditions (usage terms, presented how they designed, shown with ads, etc). If the software can't comply with those terms, then the human shouldn't be shown the content. Both parties did not agree in good faith.

You want the website to be able to force the user to see ads?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#482
post #353

Earlier quoted context omitted.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

Unreasonable is to use such incompetent companies like Cloudflare, which are absolutely incapable of distinguishing between the normal usage of a Web site by humans and DDOS attacks or accesses done by bots. Only this week I have witnessed several dozen cases when Cloudflare has blocked normal Web page accesses without any possible correct reason, and this besides the normal annoyance of slowing every single access t…

I don’t know seems like it was working as intended to me.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#483

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Right, but the LLM isn't really being used for that. It's being used for marketing and advertising purposes most of the time. The AI companies also let you play with it from time to time so you'll be a shill for them, but mostly it's the advertising people you claim to not like.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#484

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

These are more like a store putting up a billboard or catalog and asking people to turn off their meta AI glasses nearby because the store doesn't want AI translating it on your behalf as a tourist.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#485
post #415
post #123

[flagged]

> No, he (Matthew) opted everyone in by default Now you're just lying. I checked several of my Cloudflare sites and none have it enabled by default: "No robots.txt file found. Consider enabling Cloudflare managed robots.txt or generate one for your website" "A robots.txt was found and is not managed by Cloudflare" "Instruct AI bot traffic with robots.txt" disabled

I think lying is a bit strong, I think they're potentially incorrect at worst.

The Cloudflare blog post where they announced this a few weeks ago stated "Cloudflare, Inc. (NYSE: NET), the leading connectivity cloud company, today announced it is now the first Internet infrastructure provider to block AI crawlers accessing content without permission or compensation, by default." [1]

I was also a bit confused by this wording and took it to mean Cloudflare was blocking AI traffic by default. What does it mean exactly?

Third party folks seemingly also interpreted it in the same way, eg The Verge reporting it with the title "Cloudflare will now block AI crawlers by default" [2]

I think what it actually means is that they'll offer new folks a default-enabled option to block ai traffic, so existing folks won't see any change. That aligns with text deeper in their blog post:

> Upon sign-up with Cloudflare, every new domain will now be asked if they want to allow AI crawlers, giving customers the choice upfront to explicitly allow or deny AI crawlers access. This significant shift means that every new domain starts with the default of control, and eliminates the need for webpage owners to manually configure their settings to opt out. Customers can easily check their settings and enable crawling at any time if they want their content to be freely accessed.

Not sure what this looks like in practice, or whether existing customers will be notified of the new option or something. But I also wouldn't fault someone for misinterpreting the headlines; they were a bit misleading.

[1]: https://www.cloudflare.com/en-ca/press-releases/2025/cloudfl...

[2]: https://www.theverge.com/news/695501/cloudflare-block-ai-cra...

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#486

Earlier quoted context omitted.

If you allow Googlebot to crawl your website and train Gemini, but you don't allow smaller AI companies to do the same thing, then you're contributing to Google's hegemony. Given that AI is likely to be an increasingly important part of society in the future, that kind of discrimination is anti-social. I don't want a future where everything is run by Google even more than it currently is. Crawling is legal. Training…

Googlebot respects robots.txt. And Google doesn't use the fetched data from users of Chrome to supplement their search index (as a2128 is speculating that Perplexity might do when they fetch pages on the user's behalf).

Yes, but there's no way to say "allow indexing for search, but not for AI use", right?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#488

Earlier quoted context omitted.

Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?

If I were DOSing your blog, you'd ask me to stop. I run server ops for multiple online communities that are being severely negatively impacted and DOSed by these AI scrapers, and we have very few ways to stop them.

That is a problem, but is not related to my comment. The person I'm replying to is acting as if consent is a relevant aspect of the public web, I am saying it isn't. That is not the same as saying "you can do whatever you want to a public server". It is just that what you are allowed to do is not related to the arbitrary whim of the server operator.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#489
post #260

[flagged]

what if i want the rate set to zero?

Then turn off the server?

You don't have a right to say who or what can read your public website (this is a normative statement). You do have a right not to be DoS'd. If you pretend not to know what that means, it sounds the same as saying "you have an arbitrary right to decide who gets to make requests to your service", but it does not mean that.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#490
post #481

Earlier quoted context omitted.

> 1. If I as a human request a website, then I should be shown the content. Everyone agrees. I disagree. The website should have the right to say that the user can be shown the content under specific conditions (usage terms, presented how they designed, shown with ads, etc). If the software can't comply with those terms, then the human shouldn't be shown the content. Both parties did not agree in good faith.

You want the website to be able to force the user to see ads?

no, I think a fair + just world, both parties agree before they transact. There is no force in either direction (don't force creators to give their content on terms they don't want, don't force users to view ads they don't want). It's perfectly fine if people with strict preferences don't match. It's a big web, there are plenty of creators and consumers.

If the user doesn't want to view content with ads, that's okay and they can go elsewhere.

Post reply on HN