Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

381–390 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#381
post #152

Earlier quoted context omitted.

They hate AI it seems. I don’t see them offering any AI products or embracing it in any way. Seems like they’ll get left behind in the AI race.

Cloudflare literally publishes documentation pages and prompts for the single purpose of enabling better AI usage of their products and services [1,2] They offer many products for the sole purpose of enabling their customers to use AI as a part of their product offers, as even the most cursory inquiry would have uncovered. We're out here critiquing shit based on vibes vs. reality now. [1] https://developers.cloudflar…

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#382
post #246
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

I get your thinking, but x.com is proof that simply making users pay (quite a lot) does not eliminate bots.

The amount of "verified" paying "users" with a blue checkmark that are just total LLM bots is incredible on there.

As long as spamming and DDOS'ing pays more than whatever the request costs, it will keep existing.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#383
post #356

Earlier quoted context omitted.

robots.txt is for bots and I am not one though. As a user I can access anything regardless of it being blocked to bots. There are other mechanisms like status codes to rate limit or authenticate if that is an issue.

I'm talking about perplexity's behavior. Perhaps there's a point of contention on perplexity downloading a document on a person's behalf. I view this as if there is a service running that does it for multiple people, then it's a bot.

Perplexity makes requests on behalf of its users. I would argue that’s only illegitimate if the combined volume of the requests exceeds what the users would do by an order of magnitude or two. Maybe that’s what’s happening.

But “for multiple people” isn’t an argument IMO, since each of those people could run a separate service doing the same. Using the same service, on the contrary, provides an opportunity to reduce the request volume by caching.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#384

Earlier quoted context omitted.

In the same token the personal shoppers don't owe the store anything either.

Surely they owe them money for the goods and service, no? I thought that's how stores worked.

Context friend. This article and entire comments sections is about questionable web page access. Context.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#386
post #246

Earlier quoted context omitted.

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

A scary observation in light of another front page article right now: https://news.ycombinator.com/item?id=44783566 If pages can't be served for free, all internet content is at the mercy of payment processors and their ideas of "brand safety".

“Free” could have a number of meanings here. Free to the viewer, free to the hoster, free to the creator, etc…

That content can't be served entirely for free doesn't mean that all content will require payment, and so is subject to issues with payment processors, just that some things may gravitate back to a model where it costs a small amount to host something (i.e. pay for home internet and host bits off that, or you might have VPS out there that runs tools and costs a few $ /yr or /month). I pay for resources to host my bits & bobs instead of relying on services provided in exchange for stalking the people looking at them, this is free for the viewer as they aren't even paying indirectly.

Most things are paid for anyway, even if the person hosting it nor my looking at it are paying directly: adtech arseholes give services to people hosting content in exchange for the ability to stalk us and attempt to divert our attention. Very few sites/apps, other than play/hobby ones like mine or those from more actively privacy focused types, are free of that.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#387
post #342

Earlier quoted context omitted.

do you think that every well-meaning GET request should be treated the same way as a distributed attack ? The latter is the reason why people use CF not the former.

The line can be extremely blurry (that's putting it mildly), and "the latter" is not the only reason people use CF (actually, I wouldn't be surprised at all if it wasn't even the biggest reason).

The reason people use Cloudflare is because they provide free CDN, and we have at least 10 years of content marketing out there telling aspiring bloggers that, if they use a CDN in front of their website, their shitty WordPress website hosted on a shady shared hosting will become fast.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#388
post #246
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

Why would I pay for a page if I don't know if the content is what I asked for? How much are you going to pay? How much are you going to charge? This will end up in SEO hell, especially with AI-generated pages farming paid clicks.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#389

Earlier quoted context omitted.

Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…

As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?

Thanks for sharing your experience. A little off-topic but I'd like to start hosting some personal content, guides/tutorials, etc.

Do you still see authentic human traffic on your domains, is it easy to discern?

I feel like I missed the bus on running a blog pre-AI.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#390
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

I crawl 3000 RSS feeds once a week. Let me tell you! Cloudflare sucks. What business is it of theirs to block something that is meant to be accessed by everyone. Like an RSS feeds. FU Cloudflare.
Post reply on HN