Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

461–470 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#461
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> the "perplexity bots" arent crawling websites, they fetch URLs that the users explicitly asked. This shouldnt count as something that needs robots.txt access. It's not a robot randomly crawling, it's the user asking for a specific page and basically a shortcut for copy/pasting the content You say "shouldn't" here, but why? There seems to be a fundamental conflict between two groups who each assert they have "rights…

"Creators" need to eat, OK, but there's no right to get paid to paste yesterday's recycled newspapers on my laptop screen. Making that unprofitable seems incredibly good for by and large everyone.

It'd likely be a fantastic good if "content creators" stopped being able to eat from the slop they shovel. In the meantime, the smarter the tools that let folks never encounter that form of "content", the more they will pay for them.

There remain legitimate information creation or information discovery activities that nobody used to call "content". One can tell which they are by whether they have names pre-existing SEO, like "research" or "journalism" or "creative writing".

Ad-scaffolding, what the word "content" came to mean, costs money to make, ideally less than the ads it provides a place for generate. This simple equation means the whole ecosystem, together with the technology attempting to perpetuate it as viable, is an ouroboros, eating its own effluvia.

It is, I would argue, undetermined that advertising-driven content as a business model has a "right" to exist in today's form, rather than any number of other business models that sufficed for millennia of information and artistry before.

Today LLMs serve both the generation of additional literally brain-less content, and the sifting of such from information worth using. Both sides are up in arms, but in the long run, it sure seems some other form of information origination and creativity is likely to serve everyone better.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#462

Earlier quoted context omitted.

The HTTP spec draws such a distinction, albeit implicitly, in the form (and name) of its concept of "user agent."

Over time it degraded into declaring compatibility with a bunch of different browser engines and doesn't reflect the actual agent anymore. And very likely Perplexity is in fact using a Chrome-compatible engine to render the page.

The header to which you refer was named for the concept.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#463

Earlier quoted context omitted.

[flagged]

Why bring up capitalism? I don't get it. What's stopping people from lying and cheating under any other system?

When lying and cheating doesn't get you ahead, there is no reason to do it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#464

Earlier quoted context omitted.

Many people don't want their data used for free/any training. AI developers have been so repeatedly unethical that the well-earned Baysian prior is high probability that you cannot trust AI developers to not cross the training/inference streams.

> Many people don't want their data used for free/any training. That is true. But robots.txt is not designed to give them the ability to prevent this.

It is in the name, rules for the robots. Any scraping ai or not, and even mass recrsive or single page, should abide by the rules.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#465

Earlier quoted context omitted.

How is everyone having a personal shopper a problem of scale? I was going to shop myself, but I sent someone else to do it for me. At this moment I am using Perplexity's Comet browser to take a spotify playlist and add all the tracks to my youtube music playlist. I love it.

Let's look at the opposite benefit to a store if a mom that would need to bring her 3 kids to the store vs that mom having a personal shopper. In this case, the personal shopper is "better" for the store as far as physical space. However, I'm sure the store would still rather have the mom and 3 kids physically in the store so that the kids can nag mom into buying unneeded items that are placed specifically to attract…

>o that the kids can nag mom into buying unneeded items

Excellent. Personal shoppers are 'adblock for IRL'.

>You owe the companies nothing. You especially don't owe them any courtesy. They have re-arranged the world to put themselves in front of you. They never asked for your permission, don't even start asking for theirs.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#466

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

If it was just one human requesting one summary of the page nobody would ever notice. The typical watermark for junk traffic is pretty high as it was.

I have a dinky little txt site on my email domain. There is nothing of value on it, and the content changes less than once a year. So why are AI scrapers hitting it to the tune of dozens of GB per month?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#467

Earlier quoted context omitted.

Then they can't complain if they're barred entry.

http is neutral. it's up to the client to ignore robots.txt You can block IP's at the host level but there's pretty easy ways around that with proxy networks.

> http is neutral.

Who misled you with that statement?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#468

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Websites should be able to request payment. Who cares if it is a human or an agent of a human if it is paying for the request?

Cloudflare launched a mechanism for this: https://blog.cloudflare.com/introducing-pay-per-crawl/

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#469
post #290
post #22

Earlier quoted context omitted.

You're free to deny access to your site arbitrarily, including for lack of compensation.

This article is about Cloudflare attempting to deny Perplexity access to their demo site by blocking Perplexity's declared user-agent and official IP range. Perplexity responded to this denial by impersonating Google Chrome on macOS and rotating through IPs not listed in their published IP range to access the site anyway. This means it's not just "you're free to deny access to your site arbitrarily", it's "you're fre…

The comment I'm responding to established a slightly different context by asking a specific question about getting compensation from site visitors.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#470

Earlier quoted context omitted.

[flagged]

[flagged]

> High trust is prima facie incompatible with capitalism

Quite compatible

> If you want a high trust society, you don't want capitalism.

There is nothing at all in capitalism that would prevent a high level of trust in society.

> Capitalism is inherently low trust

But that's not true. The thing about capitalism is that it's RESILENT to low trust. It does not require low levels of trust, but is capable of functioning in such conditions.

> If the penalty for deceit was greater than the penalty for non-deceit

Who are the judges? Capitalism is the most resistant to deception, deceivers under capitalism receive fewer benefits than under any other economic system. Simply because capitalism is based on the premise that people cheat, act out of greed, try to get the most for themselves at the expense of others. These qualities exist in people regardless of the existence of capitalism, it is just that capitalism ensures prosperity in society even when people have these qualities.

Post reply on HN