Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

621–630 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#621

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

The problem is not about personal use. It's about big corporations scrapping billions of pages to make money.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#622
post #486

Earlier quoted context omitted.

Yes, but there's no way to say "allow indexing for search, but not for AI use", right?

But there is: https://developers.google.com/search/docs/crawling-indexing/... There is an user agent for search that you can control in robots.txt. user-agent: Googlebot There is another user agent for AI training. user-agent: Google-Extended

Wow, I had no idea this page existed, thanks for the reference!

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#623
post #585

Earlier quoted context omitted.

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

With all the crypto development how come we haven't got to HTTP/1.1 402 Payment Required WWW-price: 0.0000001 BTC, 0.000001 ETH, 0.00001 DOGE > You are less likely to participate in discussion you (or AI on your behalf) paid instead. Many sites would probably like it better.

It's not a development problem, it's an adoption problem. Publishers are desperate to sell us on a $20+/month subscription, they don't want to offer convenient affordable access to single articles.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#624

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

> why would the LLM accessing the website on my behalf be in a different legal category as my Firefox web browser accessing the website on my behalf?

It is illegal to copy stuff from the internet and then make it available from your own servers, especially when those sources have expressly asked you not to do it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#625
post #624

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

> why would the LLM accessing the website on my behalf be in a different legal category as my Firefox web browser accessing the website on my behalf? It is illegal to copy stuff from the internet and then make it available from your own servers, especially when those sources have expressly asked you not to do it.

You need more qualifiers for this to be true. archive.org routinely archives content that the site operators would prefer to be taken down.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#626

Earlier quoted context omitted.

Who cares what Hacker News wants? You’re not obliged to participate in discussion. Am I supposed to spend money on Amazon.com when I visit the website just because Amazon wants me to?

Who cares what you want?

Most humans place the desires of human beings over the desires of companies.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#627

Earlier quoted context omitted.

As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?

the fact that it would be discovered almost immediately. If you give them a URL that does not appear in Google, ask them to visit that URL specifically, and then notice the content from that URL in the training data, it's proof that they're doing this, which would be quite damaging to them.

> […] it's proof that they're doing this, which would be quite damaging to them.

Is it? It's damning, but is it damaging at all?

I'm not getting the impression that anyone's data being available for training if some bot can get to it is just how things are now, rather than an unsettled point of contention. There's too much money invested in this thing for any other outcome, and with the present decline of the rule of law…

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#628

Earlier quoted context omitted.

This honor system mostly worked at scale because interests align, which seems to be no longer the case. Does information no longer wants to be free now? Maybe internet, just like social media was just a social experiment at the end, albeit a successful one. Thanks GenAI.

Can the Terms of Service of individual content creators leverage a "death of a thousand cuts" model to produce a legal honeypot which would require organizations like Perplexity to be bound up in 10s of thousands of conciliation court cases? Big Tech has hidden behind ToS for years. Now, it seems as though it only works for them, but not against. It seems as though this would be easy to orchestrate and prove forcing…

Because lawyers are expensive and big tech companies have lots of them. Because it takes a ton of time and effort to sue someone. Because you need to show standing, which means you need to be able to demonstrate you lost something of value by their actions. Because the power imbalance is heavily weighted towards a corporation. Because the way to deal with such things should be legislation and not court decisions. And lots more reasons...

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#630
post #327

Earlier quoted context omitted.

Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.

In the previous years, I did not have many problems with Cloudflare. However, in the last few months, Cloudflare has become increasingly annoying. I suspect that they might have implemented some "AI" "threat" detection, which gives much more false positives than before. For instance, this week I have frequently been blocked when trying to access the home page of some sites where I am a paid subscriber, with a complet…

I'm running into this as well (Firefox Debian). I suspect it may be Firefox's tracker blocking combined with the older extended support release.

Sometimes just refreshing the page seems to work too. Disabling the tracker blocking allows cross-site requests to Cloudflare endpoints which seems to be enough. Maybe worth allow-listing CF domains, but I didn't look into if that is possible yet.

Post reply on HN