Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

171–180 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#171
post #38
post #5

I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? It’s a different UI, sure, but there should be no discrimination towards it as there should be no discrimination towards, say, Links terminal browser, or some exotic Firefox derivative.

>I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? AI broke the brains of many people. The internet isn't a monolith, but prior to the AI boom you'd be hard pressed to find people who were pro-copyright (except maybe a few who wanted to use it to force companies to comply with copyleft obligations), pro user-agent restrictions, or anti-scraping. Now such positio…

It's the hypocrisy you're seeing - why are AIs allowed to profit from violating copyright, while people wanting to do actually useful things have been consistently blocked? Either resolution would be fine, but we can't have it both ways.

Regardless, the bigger AI problem is spam, and that has never been acceptable.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#172

Earlier quoted context omitted.

I think it's an issue of scale. The next step in your progression here might be: If / when people have personal research bots that go and look for answers across a number of sites, requesting many pages much faster than humans do - what's the tipping point? Is personal web crawling ok? What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to gather information regularly to…

Maybe we should just institutionalize and explicitly legalize the Internet Archive and Archive Team. Then, I can download a complete and halfway current crawl of domain X from the IA and that way, no additional costs are incurred for domain X. But of course, most website publishers would hate that. Because they don't want people to access their content, they want people to look at the ads that pay them. That's why to…

I have mixed feelings on this.

Many websites (especially the bigger ones) are just businesses. They pay people to produce content, hopefully make enough ad revenue to make a profit, and repeat. Anything that reproduces their content and steals their views has a direct effect on their income and their ability to stay in business.

Maybe IA should have a way for websites to register to collect payment for lost views or something. I think it’s negligible now, there are likely no websites losing meaningful revenue from people using IA instead, but it might be a way to get better buy in if it were institutionalized.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#173

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

It's somebody's else content and resources and they are free to ban you or your bots as much as they please.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#174
post #118

Earlier quoted context omitted.

I think it’s basically impossible to prevent AI crawlers. It is like video game cheating, at the extreme they could literally point a camera at the screen and have it do image processing, and talk to the computer through the USB port emulating, a mouse and keyboard outside the machine. They don’t do that, of course, because it is much easier to do it all in software, but that is the ultimate circumvention of any atte…

I don’t subscribe to technological inevitabilism. Cloudflare banning bad actors has at least made scraping more expensive, and changes the economics of it - more sophisticated deception is necessarily more expensive. If the cost is high enough to force entry, scrapers might be willing to pay for access. But I can imagine more extreme measures. e.g. old web of trust style request signing[0]. I don’t see any easy way f…

It is inevitable, not because of some technological predestination but because if these services get hard-blocked and unable to perform their duties they will ship the agent as a web browser or browser add-on just like all the VSCode forks and then the requests will happen locally through the same pipe as the user's normal browser. It will be functionally indistinguishable from normal web traffic since it will be normal web traffic.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#175

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

1. I actually disagree. I think teasers should be free but websites should charge micropayments for their content. Here is how it can be done seamlessly, without individuals making decisions to pay every minute: https://qbix.com/ecosystem

2. This also intersects with copyright law. Ingesting content to your servers en masse through automation and transforming it there is not the same as giving people a tool (like Safari Reader) they can run on their client for specific sites they visit. Examples of companies that lost court cases about this:

  Aereo, Inc. v. American Broadcasting Companies (2014)
  TVEyes, Inc. v. Fox News Network, LLC (2018)
  UMG Recordings, Inc. v. MP3.com, Inc. (2000)
  Capitol Records, LLC v. ReDigi Inc. (2018)
  Cartoon Network v. CSC Holdings (Cablevision) (2008)
  Image Search Engines: Perfect 10 v. Google (2007)
That last one is very instructive. Caching thumbnails and previews may be OK. The rest is not. AMP is in a copyright grey area, because publishers choose to make their content available for AMP companies to redisplay. (@tptacek may have more on this)

3. Putting copyright law aside, that's the point. Decentralization vs Centralization. If a bunch of people want to come eat at an all-you-can-eat buffet, they can, because we know they have limited appetites. If you bring a giant truck and load up all the food from all all-you-can-eat buffets in the city, that's not OK, even if you later give the food away to homeless people for free. You're going to bankrupt the restaurants! https://xkcd.com/1499/

So no. The difference is that people have come to expect "free" for everything, and this is how we got into ad-supported platforms that dominate our lives.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#177
Every major AI platform is doing this right now, it's effectively impossible to avoid having your content vacuumed up by LLMs if you operate on the public web.

I've given up and restored to IP based rate-limiting to stay sane. I can't stop it, but I can (mostly) stop it from hurting my servers.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#178

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

To stretch the analogy to the breaking point: If you send 10,000 personal shoppers all at once to the same store just to check prices, the store's going to be rightfully annoyed that they aren't making sales because legit buyers can't get in.

Too bad. Build a bigger store or publish this information so we don't need 10,000 personal shoppers. Was this not the whole point of having a website? Who distorted that simple idea into the garbage websites we have now?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#179
post #156

Earlier quoted context omitted.

> companies who want AI to recommend their products need to turn this off before it starts hurting them financially Content marketing, gamified SEO, and obtrusive ads significantly hurt the quality of Google search. For all its flaws, LLMs don’t feel this gamified yet. It’s disappointing that this is probably where we’re headed. But I hope OpenAI and Anthropic realize that this drop in search result quality might be…

This has already started with people using special tags also people making content just for llms.

There is a standard for making content just for LLMs: https://llmstxt.org

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#180

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

I don't think criticizing the business practices of Cloudfare does the work of excusing Perplexity's disregard for norms.
Post reply on HN