Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

11–20 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#11

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

How about I open a proxy, replace all ads with my ads, redirect the content to you and we share the ad revenue?

That's somewhat antisocial, but perfectly legal in the US. It's called PayPal Honey, for example, and has been running for 13 years now.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#12
post #5

I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? It’s a different UI, sure, but there should be no discrimination towards it as there should be no discrimination towards, say, Links terminal browser, or some exotic Firefox derivative.

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#14
Cloudflare screaming into the void desperate to insert themselves as a middleman, in a market ( that they will never succeed in creating) where they extort scrapers for access to websites they cover.

Sorry CF, give up. the courts are on our sides here

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#15
post #9

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

If the LLM were running this sort of thing at the user's explicit request this would be fine. The problem is training. Every AI startup on the planet right now is aggressively crawling everything that will let them crawl. The server isn't seeing occasional summaries from interested users, but thousands upon thousands of bots repeatedly requesting every link they can find as fast as they can.

But that's not what this article is about. From, what I understand, this articles is about a user requesting information about a specific domain and not general scraping.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#16
post #9

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

If the LLM were running this sort of thing at the user's explicit request this would be fine. The problem is training. Every AI startup on the planet right now is aggressively crawling everything that will let them crawl. The server isn't seeing occasional summaries from interested users, but thousands upon thousands of bots repeatedly requesting every link they can find as fast as they can.

Then what if I ask the LLM 10 questions about the same domain and ask it to research further? Any human would then click through 50 - 100 articles to make sure they know what that domain contains. If that part is automated by using an LLM, does that make any legal change? How many page URLs do you think one should be allowed to access per LLM prompt?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#17
post #6

If you put info on the web, it should be available to everyone or everything with access.

Not according to CF. They are desperate to turn web sites into newspaper dispensers, where you should give them a quarter to see the content, on the basis that a bot is somehow different than a normal human vistor o a legal basis. Cf has been trying this psyop for years.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#18

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

>2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into modifying the software you run locally.

If I put time and effort into a website and it's content, I should expect no compensation despite bearing all costs.

Is that something everyone would agree with?

The internet should be entirely behind paywalls, besides content that is already provided ad free.

Is that something everyone would agree with?

I think the problem you need to be thinking about is "How can the internet work if no one wants to pay anything for anything?"

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#19

Cloudflare screaming into the void desperate to insert themselves as a middleman, in a market ( that they will never succeed in creating) where they extort scrapers for access to websites they cover. Sorry CF, give up. the courts are on our sides here

Are you sure? I'm surprised they haven't jumped in on the "scan your face to see the webpage" madness that's taking off around the world

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#20
post #3

>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…

> I think most people would draw a distinction between the two, and would at least agree the latter is more acceptable than the former. No. I should be able to control which automated retrieval tools can scrape my site, regardless of who commands it. We can play cat and mouse all day, but I control the content and I will always win: I can just take it down when annoyed badly enough. Then nobody gets the content, and…

> Then nobody gets the content, and we can all thank upstanding companies like Perplexity for that collapse of trust.

But they didn't take down the content, you did. When people running websites take down content because people use Firefox with ad-blockers, I don't blame Firefox either, I blame the website.

Post reply on HN