> How can you protect yourself? Put your valuable content behind a paywall.
A combination of "Bypass Paywalls Clean for Firefox" and archive.is usually get past these.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
221–230 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#222Earlier quoted context omitted.
It's all about scale. The impact of your personal shopper is insignificant unless you manage to scale it up into a business where everyone has a personal shopper by default.
How is everyone having a personal shopper a problem of scale? I was going to shop myself, but I sent someone else to do it for me. At this moment I am using Perplexity's Comet browser to take a spotify playlist and add all the tracks to my youtube music playlist. I love it.
If sites want to avoid people using agents, they should offer the functionality that people are using the agents to accomplish.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#223Earlier quoted context omitted.
Ads are a problematic business model, and I think your point there is kind of interesting. But AI companies disintermediating content creators from their users is NOT the web I want to replace it with. Let’s imagine you have a content creator that runs a paid newsletter. They put in lots of effort to make well-researched and compelling content. They give some of it away to entice interested parties to their site, whe…
> Otherwise there is literally no reason for them to make any of it available on the open web This is the hypothesis I always personally find fascinating in light of the army of semi-anonymous Wikipedia volunteers continuously gathering and curating information without pay. If it became functionally impossible to upsell a little information for more paid information, I'm sure some people would stop creating informati…
Existing subject-matter experts who blog for fun may or may not stick around, depending on what part of it is “fun” for them.
While some must derive satisfaction from increasing the total sum of human knowledge, others are probably blogging to engage with readers or build their own personal brand, neither of which is served by AI scrapers.
Wikipedia is an interesting case. I still don’t entirely understand why it works, though I think it’s telling that 24 years later no one has replicated their success.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#224It's ironic Perplexity itself blocks crawlers: $ curl -sI https://www.perplexity.ai | head -1 HTTP/2 403 Edit: trying to fake a browser user agent with curl also doesn't work, they're using a more sophisticated method to detect crawlers.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#225Earlier quoted context omitted.
Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…
But I can send my personal shopper and you'll be none the wiser.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#226Earlier quoted context omitted.
This analogy doesn't map to the actual problem here. Perplexity is not visiting a website everytime a user asks about it. It's frequently crawling and indexing the web, thus redirecting traffic away from websites. This crawling reduces costs and improves latency for Perplexity and its users. But it's a major threat to crawled websites
I have never created a website that I would not mind being fully crawled and indexed into another dataset that was divorced from the source (other than such divorcement makes it much harder to check pedigree, which is an academic concern, not a data-content concern: if people want to trust information from sources they can't know and they can't verify I can't fix that for them). In fact, the "old web" people sometime…
Then came the social networks and walled gardens, SEO, and all the other cancer of the last 20 years and all of these disappeared for un-searchable videos, content farms and discord communities which are basically informational black holes.
And now AI is eating that cancer, but IMO it's just one cancer being replaced by an even more insidious cancer. If all the information is accessed via AI, then the last semblance of interaction between content creators and content consumers disappears. There are no more communities, just disconnected consumers interacting with a massive aggregating AI.
Instead of discussing an interesting topic with a human, we will discuss with AI...
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#227Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…
> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. > Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. How does one follow the other? It's my web server and I can gatekeep access to my content however I want (eg Cloudflare). How is that an "abuse" of internet…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#228Earlier quoted context omitted.
1. I actually disagree. I think teasers should be free but websites should charge micropayments for their content. Here is how it can be done seamlessly, without individuals making decisions to pay every minute: https://qbix.com/ecosystem 2. This also intersects with copyright law. Ingesting content to your servers en masse through automation and transforming it there is not the same as giving people a tool (like Saf…
I would love micropayments as a kind of baked-in ecosystem support. You can crawl if you want, but it's pay to play. Which hopefully drives motivation for robust norms for content access and content scraping that makes everyone happy.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#229Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#230Perplexity Comet sort of blurs the lines there as does typing quesitons into Claude.