Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

401–410 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#401
Question for those in this thread who are okay with this: If I have endpoints that are computationally expensive server-side, what mechanism do you propose I could use to avoid being overwhelmed?

The web will be a much worse place if such services are all forced behind captchas or logins.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#402
post #246
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

Or, flip this, don't expect to get paid for pamphleteering?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#403
post #95

AI companies continuing to have problems with the concept of "consent" is increasingly alarming god help us if they ever manage to build anything more than shitty chatbots

Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?

I am not told I cannot access. And, yes, I would, because I'd be breaking the law otherwise.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#404

Why single out Perplexity? Pretty much no crawler out there fetches robots.txt. robots.txt is not a blocking mechanism; it's a hint to indicate which parts of a site might be of interest to indexing. People started using robots.txt to lie and declare things like no part of their site is interesting, and so of course that gets ignored.

That's not true, at all.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#405

Earlier quoted context omitted.

Wouldn't this lead to pirated page clones where customer pays less for same-ish content, and less, all the way down to essentially free? Because I as an user would be glad to have "free sites only" filter, and then just steal content :)) But it's an interesting idea and thought experiment.

That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.

Strongly agree with this armchair POV. Btw it doesn't cost much to host markdown.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#407
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

Ask yourself why so many content hosting platforms utilize CLoudflare's services and then contrast that perspective with your posted one. Might enlighten you a bit to think about that for a second.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#408

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Websites should be able to request payment. Who cares if it is a human or an agent of a human if it is paying for the request?

What if the agent is reselling the request?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#409
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

I crawl 3000 RSS feeds once a week. Let me tell you! Cloudflare sucks. What business is it of theirs to block something that is meant to be accessed by everyone. Like an RSS feeds. FU Cloudflare.

That's not Cloudflare's fault, that's the website owner's fault.

If they want the RSS feeds to be accessible then they should configure it to allow those requests.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#410

Earlier quoted context omitted.

That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.

You ignore the issue of motivation. Most web content exists because someone wants to make money on it. If the content creator can't do that, they will stop producing content. These AI web crawlers (Google, Perplexity, etc) are self-cannibalizing robots. They eat the goose that laid the golden egg for breakfast, and lose money doing it most of the time. If something isn't done to incentivize content creators again eve…

Pretty sure this "most" motivation means it's not a golden egg. It's SEO slop.

If only the one in ten thousand with something to share are left standing to share it, no manufactured content, that's a fine thing.

Post reply on HN