Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

371–380 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#371

Earlier quoted context omitted.

But users depend on major sites like google [insert service] still and will prioritize their usage accordingly like limited minutes and texts back in the day, right?

Networking is so cheap, unless ISPs drastically inflate their price, users won’t care. The average American allegedly* downloads 650-700GB/month, or >20GB/day. 10MB is more than enough for a webpage (honestly, 1MB is usually enough), so that means on average, ISPs serve over 2000 webpages worth of data per day. And the average internet plan is allegedly** $73/month, or That’s cheap enough, wrapped in a monthly bill,…

[dead]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#372

Earlier quoted context omitted.

It’s possible to violate all sorts of social norms. Societies that celebrate people that do so are on the far opposite end of the spectrum from high trust ones. They are rather unpleasant.

[flagged]

[flagged]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#373

Earlier quoted context omitted.

If there's an article you want to read, and the ToS says that in between reading each paragraph, you must switch to their YouTube channel and look at their ads about cat food for 5 minutes, are your going to do that?

Hacker News has collectively answered this question by consistently voting up the archive.is links in the comments of every paywalled article posted here.

New sites have collectively decided to require people use those services because they can't fathom not enshittifying everything until it's an unusable transaction hellscape.

I never really minded magazine ads or even television ads. They might have tried to make me associate boobs with a brand of soda but they didn't data mine my life and track me everywhere. I'd much rather have old fashioned manipulation than pervasive and dangerous surveillance capitalism.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#374
post #80

Earlier quoted context omitted.

I think it's an issue of scale. The next step in your progression here might be: If / when people have personal research bots that go and look for answers across a number of sites, requesting many pages much faster than humans do - what's the tipping point? Is personal web crawling ok? What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to gather information regularly to…

Doesn't o3 sort of already do this? Whenever I ask it something, it makes it look like it simultaneously opens 3-8 pages (something a human can't do). Seems like a reasonable stance would be something like "Following the no crawl directive is especially necessary when navigating websites faster than humans can." > What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to ga…

>Doesn't o3 sort of already do this?

ChatGPT probably uses a cache though. Theoretically, the average load on the original sites could be far less than users accessing them directly.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#375

Earlier quoted context omitted.

That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.

You ignore the issue of motivation. Most web content exists because someone wants to make money on it. If the content creator can't do that, they will stop producing content. These AI web crawlers (Google, Perplexity, etc) are self-cannibalizing robots. They eat the goose that laid the golden egg for breakfast, and lose money doing it most of the time. If something isn't done to incentivize content creators again eve…

AFAIK, currently creators get money while not charging for users because of ads.

While I don’t blame creators for using ads now, I don’t think they’re a long-term solution. Ads are already blocked when people visit the site with ad blockers, which are becoming more popular. Obvious sponsored content may be blocked with the ads, and non-obvious sponsored content turns these “creators” into “shills” who are inauthentic and untrustworthy. Even without Google summaries, ad revenue may decrease over time as advertisers realize they aren’t effective or want more profit; even if it doesn’t, it’s my personal opinion that society should decrease the overall amount of ads.

Not everyone creates only for money, the best only create for enough money to sustain themselves. A long-term solution is to expand art funding (e.g. creators apply for grants with their ideas and, if accepted, get paid a fixed rate to execute them) or UBI. Then media can be redistributed, remixed, etc. without impacting creators’ finances.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#376
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> the "perplexity bots" arent crawling websites, they fetch URLs that the users explicitly asked. This shouldnt count as something that needs robots.txt access. It's not a robot randomly crawling, it's the user asking for a specific page and basically a shortcut for copy/pasting the content You say "shouldn't" here, but why? There seems to be a fundamental conflict between two groups who each assert they have "rights…

The web browsers that the AI companies are about to ship will make requests that are indistinguishable from user requests. The ship on trying to save minimization has sailed.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#377
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

Spam and DDOS are serious problems, it's not fair to suggest Cloudflare is just doing this to gatekeep the Internet for its own sake.

ovh does a good job with ddos

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#378
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. Am I misunderstanding something. I (the site owner) pay Cloudflare to do this. It is my fault this happens, not Cloudflare's.

You’re paying Cloudflare to not get DDoS-attacked or swamped by illegitimate requests. GP is implying that Cloudflare could do a better job of not blocking legitimate, benign requests.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#379
post #175

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

1. I actually disagree. I think teasers should be free but websites should charge micropayments for their content. Here is how it can be done seamlessly, without individuals making decisions to pay every minute: https://qbix.com/ecosystem 2. This also intersects with copyright law. Ingesting content to your servers en masse through automation and transforming it there is not the same as giving people a tool (like Saf…

I think this is the world we are going to. I'm not going to get mired in the details of how it would happen, but I see this end result as inevitable (and we are already moving that way).

I expect a lot more paywalls for valuable content. General information is commoditized and offered in aggregated form through models. But when an AI is fetching information for you from a website, the publisher is still paying the cost of producing that content and hosting that content. The AI models are increasing the cost of hosting the content and then they are also removing the value of producing the content since you are just essentially offering value to the AI model. The user never sees your site.

I know Ads are unpopular here, but the truth is that is how publishers were compensated for your attention. When an AI model views the information that a publisher produces, then modifies it from its published form, and removes all ad content. Then you now have increased costs for producers, reduced compensation in producing content (since they are not getting ad traffic), and the content isn't even delivered in the original form.

The end result is that publishers now have to paywall their content.

Maybe an interesting middle-ground is if the AI Model companies compensated for content that they access similar to how Spotify compensates for plays of music. So if an AI model uses information from your site, they pay that publisher a fraction of a cent. People pay the AI models, and the AI models distribute that to the producers of content that feed and add value to the models.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#380
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

As a website owner I definitely want the capability allow and block certain crawlers. If I say I don’t want crawlers from Perplexity they should respect that. This sneaky evasion just highlights that company is not to be trusted, and I would definitely pay any hosting provider that helps me enforce blocking parasitic companies like perplexity.
Post reply on HN