Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

671–680 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#671
Funny enough, Perplexity blocks the bots themselves. Imagine I develop an "agent" called Merplexity, which simulates an anonymous client browsing on Perplexity and injects my ads into the output without paying for the Sonar API. Would that be OK with Perplexity?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#672

Why single out Perplexity? Pretty much no crawler out there fetches robots.txt. robots.txt is not a blocking mechanism; it's a hint to indicate which parts of a site might be of interest to indexing. People started using robots.txt to lie and declare things like no part of their site is interesting, and so of course that gets ignored.

This is objectively wrong. Take it straight from the source: https://www.rfc-editor.org/rfc/rfc9309.html

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#673

Earlier quoted context omitted.

I don't know how to make this any clearer. You - website owner - your consent does not matter . You are publishing information on the internet. I do not think you have a right to decide who is allowed to read it or not, or how they use what they read. You have exactly two legal rights: the right not to be DoS'd/hacked, and the right not to have your copyright infringed. Neither of those rights have anything to do wit…

You can repeat it, but we fundamentally disagree, it's not a matter of understanding. Fundamentally it's not true that the moment I publish something on the internet, I lose control of who can consume my intellectual property. Licensing, for example, is a way we regulate the way that code or prose can be consumed even if public. Also expressing my consent is not in any way a way to control others, is a way to control…

Ok, that wasn't clear before since you just kept saying how you expressed your consent rather than why your consent should be taken into account.

Licensing is much much more limited than you seem to be thinking of it. For instance, you said explicitly you want a way to control your ideas. The only thing this can mean is a way to control who gets to use your ideas, or what they get to use them for. So if I express a political idea in a novel way or tell a funny joke or something I should be able to dictate who gets to repeat it, or in this case with LLMs who gets to summarise and describe it.

This kind of control is antithetical to the spirit of the internet and would be frankly evil if people were actually able to assert it. Luckily in most cases it's impossible, nobody can actually stop me from describing a movie to my friends or from reposting a meme. Just copying and reposting what you wrote verbatim is something we can probably agree is wrong, but that isn't what's up for questioning here. The idea I was actually replying to in the first place was that you can decide somebody can't read your ideas - even if they're public - just because you don't like them or you don't like what they will do with them. It is hard to think of a more egregious kind of 1984-style censorship, really.

There is a place for regulation of LLM companies, they are doing a lot of harm that I wish governments would effectively rein in. It would not be hard if the political will existed. But this idea of saying I should be able to "control my ideas" is way, way worse.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#674

Earlier quoted context omitted.

The problem in your logic is that all points starts wit "I". You're not the only stakeholder in any of those interactions. There's you, a mediator (search or LLM), and the website owner. The website owner (or its users) basically do all the work and provide all the value. They produce the content and carry the costs and risks. The pre-LLM "deal" was that at least some traffic was sent their way, which helps with reac…

But the entire reason that the web is so frustrating is that visitors don't want to pay for anything. They are already paying, it is the way they are paying that causes the mess. When you buy a product, some fraction of the price is the ad budget that gets then distributed to websites showing ads. Therefore there is also nothing wrong with blocking ads, they have already been paid for, whether you look at them or not…

Publishers don't get paid a dime if you block the ad unless they are doing a direct ad transaction. Adtech has largely made that transaction a rarity for like 30 years.

It's not like newspapers where advertising is paid in full before publishers put stories online. It has not been that way for a long time.

Your reasoning for not accessing advertising reminds me of that scene in Arrested Development where, to hide the money they've taken out of the till, they throw away the bananas. It doesn't hide the transaction, it compounds the problem.

If publishers were getting paid before any ads ran the publishing business would be a hell of a lot stronger.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#675

Earlier quoted context omitted.

Can the Terms of Service of individual content creators leverage a "death of a thousand cuts" model to produce a legal honeypot which would require organizations like Perplexity to be bound up in 10s of thousands of conciliation court cases? Big Tech has hidden behind ToS for years. Now, it seems as though it only works for them, but not against. It seems as though this would be easy to orchestrate and prove forcing…

Because lawyers are expensive and big tech companies have lots of them. Because it takes a ton of time and effort to sue someone. Because you need to show standing, which means you need to be able to demonstrate you lost something of value by their actions. Because the power imbalance is heavily weighted towards a corporation. Because the way to deal with such things should be legislation and not court decisions. And…

That's exactly why I said conciliation court. None of what you've outlined is required nor is it expensive. But, for each case, the defendant is still required to show up.

I've successfully used conciliation court against large corporations in the past which is why I question it here.

And while this should be able to be handled via legislation it won't be. Beyond that a workaround could force that to happen.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#676

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

Cloudflare is growing more and more vile with each passing year. Half the tools they're building now should never have existed in the first place.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#677
post #614

Earlier quoted context omitted.

> You asked the web-enabled AI to look at the domains. Right, and the domain was configured to disallow crawlers, but Perplexity crawled it anyway. I am really struggling to see how this is hard to understand. If you mean to say "I don't think there is anything wrong with ignoring robots.txt" then just say that . Don't pretend they didn't make it clear what they're objecting to, because they spell it out repeatedly.

> Perplexity crawled it anyway No, they did not. Crawling = recursive fetching, which wasn't what was happening here. But also, I don't think there is anything wrong with ignoring robots.txt. In fact, I believe it is discriminatory and people should ignore it. See: https://wiki.archiveteam.org/index.php/Robots.txt

> I don't think there is anything wrong with ignoring robots.txt

Neither do I, I just thought your reply was disingenuous.

> Crawling = recursive fetching

I do not find this convincing. I am ok with using the word crawler for recursive fetching only. But robots.txt is not only for excluding crawlers and never has been. From the very beginning it was used to exclude specific automated clients, whether they only fetch one page or many, and that is certainly how the vast majority of people think about it today.

Like I implied in my first comment, I have no problem with you saying you dislike robots.txt, but it is not reasonable to pretend the article is unclear in some way.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#678

Earlier quoted context omitted.

Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?

Yes. Asking for consent is built into the HTTP protocol. The issue at hand here is that Perplexity scrapers lie about who they are by providing a false user agent. Thus consent was given on a false pretense.

> Asking for consent is built into the HTTP protocol.

The HTTP protocol does not specify what is right and wrong. The fact a protocol encodes or permits a particular kind of behaviour does not mean that every use of the protocol is ethically justified. I am sure you would agree with me that "black people can't visit this server" would be such an unethical rule, even though HTTP permits you to enforce such a rule. So let's forget about the protocol for a minute.

Is it morally wrong to lie about your User Agent in order to visit a website. Well, that depends on whether it is legitimate for the server operator to discriminate according to the User Agent. If it is not legitimate, then lying about your User Agent to circumvent the restriction is morally justified.

So we are back at square one: is it legitimate for a server operator to discriminate what sort of a client is used to visit them. Since the service is public, the person is allowed to visit the service and to read the content. If the client is misbehaved in some way (some LLM scrapers are) then this is a legitimate difference. But if this is controlled for, so the LLM scraper can't be easily distinguished from a human doing the same thing, then the service is not harmed any more than would be ordinary. Therefore the discrimination is not legitimate.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#679
post #642

We (humanity) need to invent a simple GPLv3 style license “You can derive any data on the data you see here, any derived data you sell or share should mention this place as a source and is subject to the same copyright as the source”. This will imply scraped datasets should become public and the law enforcement bodies will be able to work in an established framework to fight copyright and license crimes. Just blockin…

Because scrapers would certainly comply with that /s

More like have easier to assess legality status.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#680

This is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the L…

Hi, website operator here. I don't want my content to be accessible to you through Perplexity. I want my work to be freely available to any person who wants it. Feel free to transform my material as you see fit. Hell, do it with LLMs! I don't care. The LLM isn't the problem, it's what companies like Perplexity are doing with the LLM. Do not create commercial products that regurgitate my work as if it was your own. It…

To be honest I don't care.
Post reply on HN