Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

601–610 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#601

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

This is a hypothetical so give me a little rope here, but what if robots.txt wasn't a suggestion? What if it were binding (leaving aside for a moment how one would enforce / mandate / guarantee that)?

Would that solve the whole problem? Folks who ran webservers declared what they consent to, and that happens?

I think it's useful to just see if there's a consensus on that: actually making that happen is a whole can of worms itself, but it's strictly simpler than devising a good outcome without the consensus.

(And such things are not impossible, merely difficult, we have other systems ranging from BGP to the TLD mechanism that get honored in real life).

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#602

Earlier quoted context omitted.

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

Who cares what Hacker News wants? You’re not obliged to participate in discussion. Am I supposed to spend money on Amazon.com when I visit the website just because Amazon wants me to?

Whats the point of a human coming to a site if all the threads and empty and its front page is a glorified RSS feed for lazy peoples AI agents?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#603

Earlier quoted context omitted.

Another very important reason why I prefer Perplexity is due to attribution. It actually does cite the sources that it bases it's output on (unless it's something generic or calculated), so if I suspect something is off or want to look deeper into some particular aspect I can easily click through. And I've done enough click-throughs to be confident that Perplexity faithfully represents sourced content, and accurately…

> as LLM adoption increases you'll merely find that your site has fewer and fewer visits overall, so your content will only be utilized by you and a vanishingly small group of other persons. So far, AI has had the opposite effect on my site. I've now been featured on both Hackaday and Adafruit's blog. Both features were clearly AI-generated. Both posts coincided with an influx of emails from folks interested in my wo…

> So far, AI has had the opposite effect on my site. I've now been featured on both Hackaday and Adafruit's blog. Both features were clearly AI-generated. Both posts coincided with an influx of emails from folks interested in my work.

This may be missing some context, but it seems as though you're saying that you made something with AI and it led to traction. That's great! Seems off the point that blocking LLM service will lead to less exposure over time though.

> Perplexity is good at citing things when it decides to cite things and when you tell it to cite things.

Maybe I'm just lucky, but a quick skim of my Perplexity history yielded only 2 instances of no citations, and they were for general coding queries. I've never had to ask it to cite anything, as that's built into the default prompt.

> lose those aforementioned email conversations.

I think those will remain a possibility as long as LLM users, or services, ensure citations are included in output.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#604
post #54

Earlier quoted context omitted.

I like the terminology "crawler" vs. "fetcher" to distinguish between mass scraping and something more targeted as a user agent. I've been working on AI agent detection recently (see https://stytch.com/blog/introducing-is-agent/ ) and I think there's genuine value in website owners being able to identify AI agents to e.g. nudge them towards scoped access flows instead of fully impersonating a user with no controls. O…

A/ i love this distinction. B/ my brother used to use "fetcher" as a non-swear for "fucker"

Fetcher? Damn near killed'er!

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#605

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

The simple answer to #3 is advertising, including telemetry, tracking and other forms of web-based surveillance. These usually rely on certain browser "features" and/or default settings.

The goal is not to make the content usable. The goal is to get the traffic.

When advertising alone is the "business model", e.g., not the value of the "content", then even Cloudflare is going to try to protect it (the advertising, not the content). Anything to get www users to turn on Javascript so the surveillance capitalism can proceed. Hence all the "challenges" to frustrate and filter out software thatis not advertising-friendly, e.g., graphical.

Cloudflare's ruminations on user-agent strings are perplexing. It has been an expectation that the user-agent HTTP header will be spoofed since the earliest web browsers. The user-agent header is a joke.

This is from circa 1993, the year the www was opened to public access:

https://raw.githubusercontent.com/alandipert/ncsa-mosaic/mas...

Cloudflare's "bot protections" are not to ensure human use of a website but to ensure use of specific software to access a website. Software that facilitates data collection and advertising services. For example, advertising-sponsored browsers. Any other software is labeled "bot". It does not matter if a human is operating it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#606

Earlier quoted context omitted.

I don't think people have a problem with an LLM issuing GET website.com and then summarising that, each and every time it uses that information (or atleast, save a citation to it and refer to that citation). Except ad ecosystem, ignoring them for now, please refer to last paragraph. The problem is with the LLM then training on that data _once_ and then storing it forever and regurgitating it N times in the future wit…

Yes, this is the crux of the matter. The "social contract" that has been established over the last 25+ years is that site owners don't mind their site being crawled reasonably provided that the indexing that results from it links back to their content. So when AltaVista/Yahoo/Google do it and then score and list your website, interspersing that with a few ads, then it's a sensible quid pro quo for everyone. LLM AI ou…

Anything but expanding copyright laws. Tbh, a pay per citation with an opt in database to add your info (think music streaming style monetization) would be reasonable to me. Not that I think it's a good scheme for music but I think it's fitting for web crawling. Though it does inevitably lead to enshitification. Pick your poison I guess.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#607

Earlier quoted context omitted.

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

Who cares what Hacker News wants? You’re not obliged to participate in discussion. Am I supposed to spend money on Amazon.com when I visit the website just because Amazon wants me to?

Who cares what you want?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#608
post #485

Earlier quoted context omitted.

I think lying is a bit strong, I think they're potentially incorrect at worst. The Cloudflare blog post where they announced this a few weeks ago stated "Cloudflare, Inc. (NYSE: NET), the leading connectivity cloud company, today announced it is now the first Internet infrastructure provider to block AI crawlers accessing content without permission or compensation, by default." [1] I was also a bit confused by this w…

> I think lying is a bit strong, I think they're potentially incorrect at worst. I understand that you're trying to be generous, but the claim that "Matthew opted everyone in by default" is flat out incorrect.

We are in agreement; I just think saying "lying" implies a level malintent that isn't present -- it's rather overly ungenerous. The poster is at worst incorrect. And their misconception is understandable given the company's own confusing marketing.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#609

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

What if my local ai model and system crawls, indexes and trains itself on content that only I can see and work with?
Post reply on HN