Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

181–190 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#182

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

I believe you're being disingenuous. Perplexity is running a set of crawlers that do not respect robots.txt and take steps to actively evade detection.

They are running a service and this is not a user taking steps to modify their own content for their own use.

Perplexity is not acting as a user proxy and they need to learn to stick to the rules, even when it interferes with their business model.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#183

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

For 1, 2, and 3, the website owner can choose to block you completely based on IP address or your User Agent. It's not nice, but the best reaction would be to find another website.

Perplexity is choosing to come back "on a VPN" with new IP addresses to evade the block.

#2 and #3 are about modifying data where access has been granted, I think Cloudflare is really complaining about #1.

Evading an IP address ban doesn't violate my principles in some cases, and does in others.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#184

Earlier quoted context omitted.

To stretch the analogy to the breaking point: If you send 10,000 personal shoppers all at once to the same store just to check prices, the store's going to be rightfully annoyed that they aren't making sales because legit buyers can't get in.

Too bad. Build a bigger store or publish this information so we don't need 10,000 personal shoppers. Was this not the whole point of having a website? Who distorted that simple idea into the garbage websites we have now?

Weird take. The store doesn't owe your personal shippers anything.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#185

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Websites should be able to request payment. Who cares if it is a human or an agent of a human if it is paying for the request?

They are able to request payment.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#186

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

I think it's an issue of scale. The next step in your progression here might be: If / when people have personal research bots that go and look for answers across a number of sites, requesting many pages much faster than humans do - what's the tipping point? Is personal web crawling ok? What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to gather information regularly to…

I saw someone suggest in another post, if only one crawler was visiting and scraping and everyone else reused from that copy I think most websites would be ok with it. But the problem is every billionaire backed startup draining your resources with something similar to a DOS attack.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#187

Earlier quoted context omitted.

If the AI archives/caches all the results it accesses and enough people use it, doesn't it become a scraper? Just learn off the cached data. Being the man-in-the-middle seems like a pretty easy way to scrape salient content while also getting signals about that content's value.

No. The key difference is that if a user asks about a specific page, when Perplexity fetches that page, it is being operated by a human not acting as a crawler. It doesn’t matter how many times this happens or what they do with the result. If they aren’t recursively fetching pages, then they aren’t a crawler and robots.txt does not apply to them. robots.txt is not a generic access control mechanism, it is designed so…

> It doesn’t matter how many times this happens or what they do with the result.

That's where you lost me, as this is key to GP's point above and it takes more than a mere out-of-left-field declaration that "it doesn't matter" to settle the question of whether it matters.

I think they raised an important point about using cached data to support functions beyond the scope of simple at-request page retrieval.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#188

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

It’s possible to violate all sorts of social norms. Societies that celebrate people that do so are on the far opposite end of the spectrum from high trust ones. They are rather unpleasant.

[flagged]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#189
post #52
post #37

Earlier quoted context omitted.

Ads are a problematic business model, and I think your point there is kind of interesting. But AI companies disintermediating content creators from their users is NOT the web I want to replace it with. Let’s imagine you have a content creator that runs a paid newsletter. They put in lots of effort to make well-researched and compelling content. They give some of it away to entice interested parties to their site, whe…

Ofttimes people are sufficiently anti-ad that this point won't resonate well. I'm personally mostly in that camp in that with relatively few exceptions money seems to make the parts of the web I care about worse (it's hard to replace passion, and wading through SEO-optimized AI drivel to find a good site is a lot of work). Giving them concrete examples of sites which would go away can help make your point. E.g., Shel…

> Sheldon Brown (July 14, 1944 – February 4, 2008)

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#190

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process.

> If you want to gatekeep your content, use authentication.

Are there no limits on what you use the content for? I can start my own search engine that just scrapes Google results?

Post reply on HN