Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

221–230 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#221

> How can you protect yourself? Put your valuable content behind a paywall.

A combination of "Bypass Paywalls Clean for Firefox" and archive.is usually get past these.

Isn't that only because they offer unpaywalled versions to web crawlers in the first place, so they still get ranked in search results?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#222
post #139

Earlier quoted context omitted.

It's all about scale. The impact of your personal shopper is insignificant unless you manage to scale it up into a business where everyone has a personal shopper by default.

How is everyone having a personal shopper a problem of scale? I was going to shop myself, but I sent someone else to do it for me. At this moment I am using Perplexity's Comet browser to take a spotify playlist and add all the tracks to my youtube music playlist. I love it.

We'll see more of this sort of thing as AI agents become more popular and capable. They will do things that the site or app should be able to do (or rather, things that users want to be able to do) but don't offer. The YouTube music playlist is a good example. One thing I'd like to be able to do is make a playlist of some specific artists. But you can't. You have to select specific songs.

If sites want to avoid people using agents, they should offer the functionality that people are using the agents to accomplish.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#223
post #37

Earlier quoted context omitted.

Ads are a problematic business model, and I think your point there is kind of interesting. But AI companies disintermediating content creators from their users is NOT the web I want to replace it with. Let’s imagine you have a content creator that runs a paid newsletter. They put in lots of effort to make well-researched and compelling content. They give some of it away to entice interested parties to their site, whe…

> Otherwise there is literally no reason for them to make any of it available on the open web This is the hypothesis I always personally find fascinating in light of the army of semi-anonymous Wikipedia volunteers continuously gathering and curating information without pay. If it became functionally impossible to upsell a little information for more paid information, I'm sure some people would stop creating informati…

Any information that requires something approximating a full-time job worth of effort to produce will necessarily go away, barring the small number of independently wealthy creators.

Existing subject-matter experts who blog for fun may or may not stick around, depending on what part of it is “fun” for them.

While some must derive satisfaction from increasing the total sum of human knowledge, others are probably blogging to engage with readers or build their own personal brand, neither of which is served by AI scrapers.

Wikipedia is an interesting case. I still don’t entirely understand why it works, though I think it’s telling that 24 years later no one has replicated their success.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#224
post #146

It's ironic Perplexity itself blocks crawlers: $ curl -sI https://www.perplexity.ai | head -1 HTTP/2 403 Edit: trying to fake a browser user agent with curl also doesn't work, they're using a more sophisticated method to detect crawlers.

someone asked this already to the CEO: https://x.com/AravSrinivas/status/1819610286036488625

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#225

Earlier quoted context omitted.

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

But I can send my personal shopper and you'll be none the wiser.

And you can be trespassed and prosecuted if you continue to violate.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#226
post #116

Earlier quoted context omitted.

This analogy doesn't map to the actual problem here. Perplexity is not visiting a website everytime a user asks about it. It's frequently crawling and indexing the web, thus redirecting traffic away from websites. This crawling reduces costs and improves latency for Perplexity and its users. But it's a major threat to crawled websites

I have never created a website that I would not mind being fully crawled and indexed into another dataset that was divorced from the source (other than such divorcement makes it much harder to check pedigree, which is an academic concern, not a data-content concern: if people want to trust information from sources they can't know and they can't verify I can't fix that for them). In fact, the "old web" people sometime…

There's an important distinction that we are glossing over I think. In the times of the "old web", people were putting things online to interact with a (large) online audience. If people found your content interesting, they'd keep coming back and some of them would email you, there'd be discussions on forums, IRC chatrooms, mailing lists, etc. Communities were built around interesting topics, and websites that started out as just some personal blog that someone used to write down their thoughts would grow into fonts of information for a large number of people.

Then came the social networks and walled gardens, SEO, and all the other cancer of the last 20 years and all of these disappeared for un-searchable videos, content farms and discord communities which are basically informational black holes.

And now AI is eating that cancer, but IMO it's just one cancer being replaced by an even more insidious cancer. If all the information is accessed via AI, then the last semblance of interaction between content creators and content consumers disappears. There are no more communities, just disconnected consumers interacting with a massive aggregating AI.

Instead of discussing an interesting topic with a human, we will discuss with AI...

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#227

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. > Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. How does one follow the other? It's my web server and I can gatekeep access to my content however I want (eg Cloudflare). How is that an "abuse" of internet…

most users of cloudflare assume it's for spam control. They don't realize that they are blocking their content for everyone except for Faangs

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#228
post #175

Earlier quoted context omitted.

1. I actually disagree. I think teasers should be free but websites should charge micropayments for their content. Here is how it can be done seamlessly, without individuals making decisions to pay every minute: https://qbix.com/ecosystem 2. This also intersects with copyright law. Ingesting content to your servers en masse through automation and transforming it there is not the same as giving people a tool (like Saf…

I would love micropayments as a kind of baked-in ecosystem support. You can crawl if you want, but it's pay to play. Which hopefully drives motivation for robust norms for content access and content scraping that makes everyone happy.

I want to bring Ted Nelson on my channel and interview him about Xanadu. Does anyone here know him?

https://xanadu.com.au/ted/XU/XuPageKeio.html

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#230
I’m just curious at what point ai is a crawler and at what point ai is a client when the user is directing the searches and the ai is executing them.

Perplexity Comet sort of blurs the lines there as does typing quesitons into Claude.

Post reply on HN