Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

581–590 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#581
post #107

Earlier quoted context omitted.

Maybe we should just institutionalize and explicitly legalize the Internet Archive and Archive Team. Then, I can download a complete and halfway current crawl of domain X from the IA and that way, no additional costs are incurred for domain X. But of course, most website publishers would hate that. Because they don't want people to access their content, they want people to look at the ads that pay them. That's why to…

Or websites can monetize their data via paid apis and downloadable archives. That's what makes Reddit the most valuable data trove for regular users.

[flagged]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#582
post #537
post #491

Earlier quoted context omitted.

Won't work with biometric attestation. For example, banks in China require periodic facial recognition to continue the banking session.

yea but those are not open sites, try imposing that on an open site you'd want to actually attract human traffic to

see, literally, reddit requiring teenagers to open their mouth and roll their heads around to enter.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#584

This is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the L…

Hi, website operator here. I don't want my content to be accessible to you through Perplexity. I want my work to be freely available to any person who wants it. Feel free to transform my material as you see fit. Hell, do it with LLMs! I don't care. The LLM isn't the problem, it's what companies like Perplexity are doing with the LLM. Do not create commercial products that regurgitate my work as if it was your own. It…

Another very important reason why I prefer Perplexity is due to attribution. It actually does cite the sources that it bases it's output on (unless it's something generic or calculated), so if I suspect something is off or want to look deeper into some particular aspect I can easily click through. And I've done enough click-throughs to be confident that Perplexity faithfully represents sourced content, and accurately gets exactly the bits I'm interested in, maybe 98% or more of the time.

It is your prerogative to tune your servers as you see fit, but as LLM adoption increases you'll merely find that your site has fewer and fewer visits overall, so your content will only be utilized by you and a vanishingly small group of other persons. Perhaps you're OK with that, and that's also fine for the rest of us.

It's strange you mention theft, and then say it isn't about money. For me, and many others, it's about practicality and efficiency. We went from having to visit physical libraries to using search engines, and now we're entering the era of increasingly intelligent content fetch+preprocess tools.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#585

Earlier quoted context omitted.

Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

With all the crypto development how come we haven't got to

  HTTP/1.1 402 Payment Required
  WWW-price: 0.0000001 BTC, 0.000001 ETH, 0.00001 DOGE
> You are less likely to participate in discussion

you (or AI on your behalf) paid instead. Many sites would probably like it better.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#586

Earlier quoted context omitted.

Hi, website operator here. I don't want my content to be accessible to you through Perplexity. I want my work to be freely available to any person who wants it. Feel free to transform my material as you see fit. Hell, do it with LLMs! I don't care. The LLM isn't the problem, it's what companies like Perplexity are doing with the LLM. Do not create commercial products that regurgitate my work as if it was your own. It…

Another very important reason why I prefer Perplexity is due to attribution. It actually does cite the sources that it bases it's output on (unless it's something generic or calculated), so if I suspect something is off or want to look deeper into some particular aspect I can easily click through. And I've done enough click-throughs to be confident that Perplexity faithfully represents sourced content, and accurately…

> as LLM adoption increases you'll merely find that your site has fewer and fewer visits overall, so your content will only be utilized by you and a vanishingly small group of other persons.

So far, AI has had the opposite effect on my site. I've now been featured on both Hackaday and Adafruit's blog. Both features were clearly AI-generated. Both posts coincided with an influx of emails from folks interested in my work.

Perplexity is good at citing things when it decides to cite things and when you tell it to cite things. It can and does spit out plain expository text with no indication of the information's origin. I do appreciate that you have better-than-usual habits about validating sources.

I think you may have misinterpreted my remark about money. With the direction conversations around AI have been going lately, I was expecting a backhanded accusation that I was farming ad revenue.

"It's not about money" meant that I have nothing to lose financially by losing direct human traffic to my websites. Instead, I stand to lose those aforementioned email conversations.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#587

Earlier quoted context omitted.

The way to prevent people from downloading your pages and using them is to take them off the public internet. There are laws to prevent people from violating your copyright or from preventing access to your service (by excessive traffic). But there is (thankfully) no magical right that stops people from reading your content and describing it.

Many site operators want people to access their content, but prevent AI companies from scraping their sites for training data. People who think like that made tools like Anubis, and it works. I also want to keep this distinction on the sites I own. I also use licenses to signal that this site is not good to use for AI training, because it's CC BY-NC-SA-2.0. So, I license my content appropriately (No derivative, Non-c…

I don't want AI companies to scrape my sites (or use the files I wrote) for training data either, but that is not specifically what I am trying to stop (unless the files are supposed to be private and unpublished). I should not stop them from using the files for what they want, once they have them. (I also specifically do not want to block use of lynx, curl, Dillo, etc.)

What I want to stop is excessive crawling and scraping of my server. Once they have the file they can do what they want with it. Another comment (44786237) mentions that robots.txt is only for restricting recursive access; I agree and that is what should be blocked. They also should not access the same file several times quickly even though it should be unnecessary to do so, just as much as they should not access all of the files. (If someone wants to make a mirror of the files, there may be other ways, e.g. in case there is a archive file available to download many at once (possibly, in case if the site operator made their own index and then did it this way). If it is a git repository, then it can be cloned.)

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#588

Earlier quoted context omitted.

If magazines and newspapers were once able to be funded by native ads, so can websites. The spying industry doesn't want you to know this, but ads work without spying too - just look at all the IRL billboards still around.

I never said anything about spying. Magazines and newspapers were able to by funded by native ads because you couldn't auto-remove ads from their printed media and nobody could clone their content and give it away for free.

If ads were more respectful I wouldn’t have to remove them. Alas they can’t help themselves and so I do.

When ads were far less invasive, I had a lot more tolerance.

Now they want my data, they want to play audio, video, hijack the content, page etc.

Advertising scum can not be trusted to forever take more and more and more.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#589

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

We have a faceted search that creates billions of unique URLs by combinations of the facets. As such, we block all crawlers from it in robots.txt, which saves us AND them from a bunch of pointless indexing load.

But a stealth bot has been crawling all these URLs for weeks. Thus wasting a shitload of our resources AND a shitload of their resources too.

Whoever it is (and I now suspect it is Perplexity based on this Cloudflare post), they thought they were being so clever by ignoring our robots.txt. Instead they have been wasting money for weeks. Our block was there for a reason.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#590

Their test seems flawed: > We created multiple brand-new domains, similar to testexample.com and secretexample.com. These domains were newly purchased and had not yet been indexed by any search engine nor made publicly accessible in any discoverable way. We implemented a robots.txt file with directives to stop any respectful bots from accessing any part of a website: > We conducted an experiment by querying Perplexit…

Yeah I'm not so sure about that. If Perplexity are visiting that page on your behalf to give you some information and aren't doing anything else with it, and just throw away that data afterwards, then you may have a point. As a site owner, I feel it's still my decision what I do and don't let you do, because you're visiting a page that I own and serve. But if, as I suspect, Perplexity are visiting that page and then…

> But if, as I suspect, Perplexity are visiting that page and then using information from that webpage in order to train their model then sorry mate, you're a crawler, you're just using a user as a proxy for your crawling activity.

If it is not recursive access, and is only one file, then it hopefully should be OK (except for issues with HTML where common browsers will usually also download CSS, JavaScripts, WebAssembly, pictures, favicons (even if the web page does not declare any favicons), etc; many "small web" formats deliberately avoid this), especially if it is just used only since you requested it.

However, if they do then use it to train their model, without documenting that, that can be a problem, especially if the file being accessed is not intended to be public; but this is a different issue than the above.

Post reply on HN