Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

321–330 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#321

Earlier quoted context omitted.

You understand that HN is ad supported too, right?

No, I don't. But what is your point? Is the value in HN primarily in its hosting, or the non-ad-supported community?

Outside of Wikipedia, I'm not sure what content you are thinking of.

Taking HN as a potential one of these places, it doesn't even qualify. HN is funded entirely to be a place for advertising ycombinator companies to a large crowd of developers. HN is literally a developer honey pot that they get exclusive ad rights to.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#322
post #22

Earlier quoted context omitted.

You're free to deny access to your site arbitrarily, including for lack of compensation.

>and the website should not be notified about it.

My user agent and its handling of your content once it's on my computer are not your concern. You don't need to know if the data is parsed by a screen reader, an AI agent, or just piped to /dev/null. It's simply not your concern and never will be.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#323
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> why does perplexity even need to crawl websites?

I was recently working on a project where I needed to find out the published date for a lot of article links and this came helpful. Not sure if it's changed recently but asking ChatGPT, Gemini etc didn't work and it said that it doesn't have access to the current websites. However, asking perplexity, it fetched the website in real time and gave me the info I needed.

I do agree with the rest of your comment that this is not a random robot crawling. It was doing what a real user (me) asked it to fetch.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#324
post #190

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. > If you want to gatekeep your content, use authentication. Are there no limits on what you use the content for? I can start my own search engine that just scrapes Google results?

I tried to scrape Google results once using an automated process, and quickly got banned from all of Google. They banned my IP address completely. It kind of really sucked for a while, until my ISP assigned a new IP address. Funny enough, this was about 15 years ago and I was exploring developing something very similar to what LLMs are today.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#325
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior The Big Tech bots provide proven value to most sites. They have also through the years proven themselves to respect robots.txt, including crawl speed directives. If you manage…

it's because those SEO bots keep crawling over and over, which perplexity does not seem to do (considering that the URLS are user-requested). Those are different cases and robots.txt is only about the former. Cloudflare in this case is not doing "ddos protection" because i presume Perplexity does not constantly refetch or crawl or ddos the website (If perplexity does those things then they are guilty)

https://www.robotstxt.org/faq/what.html

I wonder if cloudflare users explicitly have to allow google or if it's pre-allowed for them when setting up cloudflare.

Despite what Cloudflare wants us to think here, the web was always meant to be an open information network , and spam protection should not fundamentally change that characteristic.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#326
post #246

Earlier quoted context omitted.

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

> We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. But it's done through a bait and switch. They serve the full article to Google, which allows Google to show you excerpts that you have to pay for. It would be better if Google shows something like PAYMENT REQUIRED on…

> They serve the full article to Google, which allows Google to show you excerpts that you have to pay for.

I'm old enough to remember when that was grounds for getting your site removed from Google results - "cloaking" was against the rules. You couldn't return one result for Googlebot, and another for humans.

No idea when they stopped doing that, but they obviously have let go of that principle.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#327
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#328
post #296

Earlier quoted context omitted.

> "Stealth" crawlers are always going to win the game. no, because we'll end up with remote attestation needed to access any site of value

Almost no site of value will use remote attestation because an alternative that works will all of your devices, operating systems, ad blockers and extensions will attract more users than your locked-down site.

> alternative that works will all of your devices, operating systems, ad blockers and extensions

When 99.9% of users are using the same few types of locked down devices, operating systems, and browsers that all support remote attestation, the 0.1% doesn't matter. This is already the case on mobile devices, it's only a matter of time until computers become just as locked down.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#329

Earlier quoted context omitted.

Wouldn't this lead to pirated page clones where customer pays less for same-ish content, and less, all the way down to essentially free? Because I as an user would be glad to have "free sites only" filter, and then just steal content :)) But it's an interesting idea and thought experiment.

That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.

You ignore the issue of motivation. Most web content exists because someone wants to make money on it. If the content creator can't do that, they will stop producing content.

These AI web crawlers (Google, Perplexity, etc) are self-cannibalizing robots. They eat the goose that laid the golden egg for breakfast, and lose money doing it most of the time.

If something isn't done to incentivize content creators again eventually there will be only walled-gardens and obsolete content left for the cannibals.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#330

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

> 1. If I as a human request a website, then I should be shown the content. Everyone agrees.

I disagree. The website should have the right to say that the user can be shown the content under specific conditions (usage terms, presented how they designed, shown with ads, etc). If the software can't comply with those terms, then the human shouldn't be shown the content. Both parties did not agree in good faith.

Post reply on HN