Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

121–130 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#121

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

I don't think people have a problem with an LLM issuing GET website.com and then summarising that, each and every time it uses that information (or atleast, save a citation to it and refer to that citation). Except ad ecosystem, ignoring them for now, please refer to last paragraph. The problem is with the LLM then training on that data _once_ and then storing it forever and regurgitating it N times in the future wit…

Yes, this is the crux of the matter.

The "social contract" that has been established over the last 25+ years is that site owners don't mind their site being crawled reasonably provided that the indexing that results from it links back to their content. So when AltaVista/Yahoo/Google do it and then score and list your website, interspersing that with a few ads, then it's a sensible quid pro quo for everyone.

LLM AI outfits are abusing this social contract by stuffing the crawled data into their models, summarising/remixing/learning from this content, claiming "fair use" and then not providing the quid pro quo back to the originating data. This is quite likely terminal for many content-oriented businesses, which ironically means it will also be terminal for those who will ultimately depend on additions, changes and corrections to that content - LLM AI outfits.

IMO: copyright law needs an update to mandate no training on content without explicit permission from the holder of the copyright of that content. And perhaps, as others have pointed out, an llms.txt to augment robots.txt that covers this for llm digestion purposes.

EDIT: Apparently llms.txt has been suggested, but from what I can tell this isn't about restricting access: https://llmstxt.org/

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#122

Earlier quoted context omitted.

Technically correct (the best kind of correct), but if I set a thousand users on to a website to each download a single page and then feed the information they retrieve from that one page into my AI model, then are those thousand users not performing the same function as a crawler, even though they are (technically) not one? If it looks like a duck, quacks like a duck and surfs a website like a duck, then perhaps we…

> if I set a thousand users on to a website to each download a single page and then feed the information they retrieve from that one page into my AI model, then are those thousand users not performing the same function as a crawler, even though they are (technically) not one? No. robots.txt is designed to stop recursive fetching. It is not designed to stop AI companies from getting your content. Devising scenarios in…

I don’t disagree with you about robots.txt… however, what _is_ the right tool for the job?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#124

Earlier quoted context omitted.

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

But I can send my personal shopper and you'll be none the wiser.

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#125

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Flip it around, why would you go to the trouble of creating a web page and content for it, if some AI bot is going to scrape it and save people the trouble of visiting your site? The value of your work has been captured by some AI company (by somewhat nefarious means too).

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#126
Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process.

If you want to gatekeep your content, use authentication.

Robots.txt is not a technical solution, it's a social nicety.

Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control.

On the technical side, we could use CRC mechanisms and differential content loading with offline caching and storage, but this puts control of content in the hands of the user, mitigates the value of surveillance and tracking, and has other side effects unpalatable to those currently exploiting user data.

Adtech companies want their public reach cake and their mass surveillance meals, too, with all sorts of malignant parties and incentives behind perpetuating the worst of all possible worlds.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#127

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

>2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into modifying the software you run locally. If I put time and effort into a website and it's content, I should expect no compensation despi…

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#128
post #116

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

This analogy doesn't map to the actual problem here. Perplexity is not visiting a website everytime a user asks about it. It's frequently crawling and indexing the web, thus redirecting traffic away from websites. This crawling reduces costs and improves latency for Perplexity and its users. But it's a major threat to crawled websites

I have never created a website that I would not mind being fully crawled and indexed into another dataset that was divorced from the source (other than such divorcement makes it much harder to check pedigree, which is an academic concern, not a data-content concern: if people want to trust information from sources they can't know and they can't verify I can't fix that for them).

In fact, the "old web" people sometimes pine for was mostly a place where people were putting things online so they were online, not because it would translate directly to money.

Perhaps AI crawlers are a harbinger for the death of the web 2.0 pay-for-info model... And perhaps that's okay.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#130

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Legal category?
Post reply on HN