Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

641–650 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#641

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

> If I now go one step further and use an LLM to summarize content because the authentic presentation is so riddled with ads, JavaScript, and pop-ups, that the content becomes borderline unusable, then why would the LLM accessing the website on my behalf be in a different legal category as my Firefox web browser accessing the website on my behalf?

Because quantity has a quality of its own.

I say this as someone who is on the side of pro local user commands how local compute works, but understand why companies are reacting to how cheap LLMs are making information discovery against their own datasets

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#642
We (humanity) need to invent a simple GPLv3 style license “You can derive any data on the data you see here, any derived data you sell or share should mention this place as a source and is subject to the same copyright as the source”. This will imply scraped datasets should become public and the law enforcement bodies will be able to work in an established framework to fight copyright and license crimes. Just blocking me from using any tools I want to make sense of the world around me (data on the internet sites being part of it) with crawlers and whatnot, is inherently evil, and is not logically consistent.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#643

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

The problem in your logic is that all points starts wit "I".

You're not the only stakeholder in any of those interactions. There's you, a mediator (search or LLM), and the website owner.

The website owner (or its users) basically do all the work and provide all the value. They produce the content and carry the costs and risks.

The pre-LLM "deal" was that at least some traffic was sent their way, which helps with reach and attempts at monetization. This too is largely a broken and asymmetrical deal where the search engine holds all the cards but it's better than nothing.

A full LLM model that no longer sends traffic to websites means there's zero incentive to have a website in the first place, or it is encouraged to put it behind a login.

I get that users prefer an uncluttered direct answer over manually scanning a puzzling web. But the entire reason that the web is so frustrating is that visitors don't want to pay for anything.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#644

Earlier quoted context omitted.

If magazines and newspapers were once able to be funded by native ads, so can websites. The spying industry doesn't want you to know this, but ads work without spying too - just look at all the IRL billboards still around.

I never said anything about spying. Magazines and newspapers were able to by funded by native ads because you couldn't auto-remove ads from their printed media and nobody could clone their content and give it away for free.

Newspapers sell information. Information is now trivial to copy and send across the globe, when 50 years ago it wasnt. And youre wrong about "nobody could clone their content", because they absolutely could, different editions were pressed throughout the day (morning, lunch, evening newspapers) at the peak of print media. The barrier to entry used to be a printing press, now its just an internet connection, print media has a hard time accepting that

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#645
post #80

Earlier quoted context omitted.

I think it's an issue of scale. The next step in your progression here might be: If / when people have personal research bots that go and look for answers across a number of sites, requesting many pages much faster than humans do - what's the tipping point? Is personal web crawling ok? What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to gather information regularly to…

Doesn't o3 sort of already do this? Whenever I ask it something, it makes it look like it simultaneously opens 3-8 pages (something a human can't do). Seems like a reasonable stance would be something like "Following the no crawl directive is especially necessary when navigating websites faster than humans can." > What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to ga…

how do you propose we do anything about this? any law you propose would have to be global

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#646
Maybe we can just configure webservers to block anyone who requests robots.txt, regular browsers don't do it, but robots do to get list of urls to crawl (while ignoring rules). Just create simple PHP/CGI script that adds client IP addres to iptables once /robots.txt is accessed.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#647
post #37

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Ads are a problematic business model, and I think your point there is kind of interesting. But AI companies disintermediating content creators from their users is NOT the web I want to replace it with. Let’s imagine you have a content creator that runs a paid newsletter. They put in lots of effort to make well-researched and compelling content. They give some of it away to entice interested parties to their site, whe…

there are companies that already do this, and the ONE thing none of them do is place the information they are selling on THE PUBLIC INTERNET. so your point is moot

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#648
post #118

Earlier quoted context omitted.

I don’t subscribe to technological inevitabilism. Cloudflare banning bad actors has at least made scraping more expensive, and changes the economics of it - more sophisticated deception is necessarily more expensive. If the cost is high enough to force entry, scrapers might be willing to pay for access. But I can imagine more extreme measures. e.g. old web of trust style request signing[0]. I don’t see any easy way f…

> Cloudflare banning bad actors has at least made scraping more expensive, and changes the economics of it - more sophisticated deception is necessarily more expensive. If the cost is high enough to force entry, scrapers might be willing to pay for access. I think this might actually point at the end state. Scraping bots will eventually get good enough to emulate a person well enough to be indistinguishable (are we t…

Netflix CAN "stop you from pointing a camera at your TV and distributing it" because of copyright law.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#649

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

The problem in your logic is that all points starts wit "I". You're not the only stakeholder in any of those interactions. There's you, a mediator (search or LLM), and the website owner. The website owner (or its users) basically do all the work and provide all the value. They produce the content and carry the costs and risks. The pre-LLM "deal" was that at least some traffic was sent their way, which helps with reac…

Agreed.

Cloudflare released these insights showing the disparity between crawling/scraping and visits referred from the AI platforms.

https://radar.cloudflare.com/ai-insights#crawl-to-refer-rati...

Post reply on HN