Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

681–690 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#681

Earlier quoted context omitted.

You can repeat it, but we fundamentally disagree, it's not a matter of understanding. Fundamentally it's not true that the moment I publish something on the internet, I lose control of who can consume my intellectual property. Licensing, for example, is a way we regulate the way that code or prose can be consumed even if public. Also expressing my consent is not in any way a way to control others, is a way to control…

Ok, that wasn't clear before since you just kept saying how you expressed your consent rather than why your consent should be taken into account. Licensing is much much more limited than you seem to be thinking of it. For instance, you said explicitly you want a way to control your ideas. The only thing this can mean is a way to control who gets to use your ideas, or what they get to use them for. So if I express a p…

LLMs are not "someone", LLMs are something, and they don't "read content", they by definition acquire and reuse that content (for example, by summarizing it), as part of their product.

So here the consent is indeed about what can be done with the data.

In general, it's absolutely the norm that public websites (I.e., unauthenticated) restrict even who can access the data. The simplest example that comes to mind is geoblocking. I have all the rights to say that my website is not made available to anybody in the US, for example. Would you still call that website "public"? Would bypassing the block via a VPN be a violation of my consent? This is mostly a moral discussion I suppose.

But anyway, it's not what's happening here. LLMs access content for the sole purpose of doing something with that content, either training or providing the service to their customers. They are not humans, they are not consumers, they don't simply fetch the content and present it to the users (a much more neutral action, like curl or the browser does). It's impossible to distinguish, in the case of LLMs the act of accessing and the act of using, so the difference you make doesn't apply in my opinion.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#683

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

It's quite easy to solve. Hold companies legally accountable for computer fraud and abuse.

The problem is that those in the position to do that are not interested.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#684
post #125

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Flip it around, why would you go to the trouble of creating a web page and content for it, if some AI bot is going to scrape it and save people the trouble of visiting your site? The value of your work has been captured by some AI company (by somewhat nefarious means too).

I don't see the problem. I want AI agents to learn my website and have it available as part of its corpus of knowledge to users. If asked for something it can answer the question, and if user requests make a parallel web search for sources which would bring up my page. The latter is only a bonus not a necessity for me. Getting the information out there by whatever means is my first priority.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#685

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

> I think that it’s pretty unambiguously reasonable to choose to not allow an unrelated business to operate inside of your physical storefront. I also think that maps onto digital services.

The line is drawn for me on my own computer. Even if I am in your building, my phone remains mine.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#687
I've jyst asked perplexity ai itself: this is the answer

In summary: Officially, Perplexity claims its bots honor robots.txt. In practice, outside investigators and hosting providers document persistent circumvention of such directives by undeclared or disguised crawlers acting on Perplexity’s behalf, especially for real-time user queries

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#688

Earlier quoted context omitted.

Foo news wants you to visit the site, look at the main page, watch the ads, click on them and buy the products advertised by third parties which will give money to Foo news in exchange for this service. And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads. They claim that since they are free to not buy an advertised product, why would they…

It's not ads. We have ads in paper magazines and newspapers and no one went around with scissors to remove them. It's obnoxious ads, designed to violently grabs your attention and trackers (malware). It's like a newspapers giving your address to a whole crew of salemens that intrudes on your property at 3am and looking at you sleeping and installing cameras in your bathroom. All so that they can jump at you in the st…

This is one of the dumbest things about ad networks. Google has enough data about your watching habits on Youtube and their algorithm is basically as good as it gets in terms of showing you what you want to watch and getting you hooked on it, but the moment they show you ads, all that technical expertise appears to have vanished into thin air and all they show you is fake mobile ads?

People hate obnoxious ads because the money that pays for them is essentially a bribe to artificially elevate content above its deserved ranking. It feels like you're being manipulated into an unfavorable trade.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#689

Earlier quoted context omitted.

The way to prevent people from downloading your pages and using them is to take them off the public internet. There are laws to prevent people from violating your copyright or from preventing access to your service (by excessive traffic). But there is (thankfully) no magical right that stops people from reading your content and describing it.

Many site operators want people to access their content, but prevent AI companies from scraping their sites for training data. People who think like that made tools like Anubis, and it works. I also want to keep this distinction on the sites I own. I also use licenses to signal that this site is not good to use for AI training, because it's CC BY-NC-SA-2.0. So, I license my content appropriately (No derivative, Non-c…

> Many site operators want people to access their content, but prevent AI companies from scraping their sites for training data.

That is unfortunately not a distinction that is currently legally enforceable. Until that changes all other "solutions" are pointless and only cause more harm.

> People who think like that made tools like Anubis, and it works.

It works to get real humans like myself to stop visiting your site while scrapers will have people whose entire job is to work around such "protections". Just like traditional DRM inconveniences honest customers and not pirates. And to be clear, what you are advocating for is DRM.

> I also want to keep this distinction on the sites I own. I also use licenses to signal that this site is not good to use for AI training, because it's CC BY-NC-SA-2.0.

If AI crawlers cared about that we wouldn't be talking about this issue. A license and only give more permissions than there are without one.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#690
post #512
post #86

In unrelated news, Fedora (the Linux distro) has been taken down by a DDoS today which I understand is AI-scraping related: https://pagure.io/fedora-infrastructure/issue/12703

The last comment there now reads: "It was actually a caching issue on our end. ;) I just fixed it a few min ago..." Lets not go on a witch hunt and blame everything on AI scrapers.

How many requests are LLMs typically making in order for people to accuse them of doing a DoS attack?
Post reply on HN