Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

441–450 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#441
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

Ironically, cloudflare is also the reason OpenAI agent mode with web use isn’t very usable right now. Every second time I asked it to do a mundane task like checking me in for a flight it couldn’t because of cloudflare.

what ironic with this???

we seeing many post about site owner that got hit by millions request because of LLM, we cant blame cloudflare for this because it literally neccessary evil

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#442
Is it just me or is it rage bait? Switching up marketing a notch when the AI paywall did not get much media attention so far? Cloudflare seems to focus on enterprise marketing nowadays, currently geared towards the media industry, rather than the technical marketing suited for the HN audience. They have no horse in the AI race, so they’re betting on the anti-AI horse instead to gain market share in the media sector?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#443

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

[flagged]

Whoa, please don't post like this. We end up banning accounts that do.

https://news.ycombinator.com/newsguidelines.html

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#444
post #246
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

Your theory does not match the practice of Cloudflare.

Whatever method is used by Cloudflare for detecting "threats" has nothing to do with consuming resources on the "protected" servers.

The so-called "threats" are identified in users that may make a few accesses per day to a site, transferring perhaps a few kilobytes of useful data on the viewed pages (besides whatever amount of stupid scripts the site designer has implemented).

So certainly Cloudflare does not meter the consumed resources.

Moreover, Cloudflare preemptively annoys any user who accesses for the first time a site, having never consumed any resources, perhaps based on irrational profiling based on the used browser and operating system, and geographical location.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#445
post #157

Earlier quoted context omitted.

> specific to robots.txt > poor summarization of data I'm not really addressing the issue raised in the article. I am noting that the LLM, when asked, is either lying to the user or making a statement that it does not know to be true (that there is no robots.txt). This is way beyond poor summarization.

I would say it's orthogonal to it. LLMs being unable to judge their capabilities is a separate issue to summarization quality.

I'm not critiquing its ability to judge its own capability, I am pointing out that it is providing false information to the user.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#446
Internet was built on trust, but not anymore. It's a Darwinian system; everyone has to find their own way to survive.

Cloudflare will help their publisher to block more aggresively, and AI companies will up their game too. Harvest information online is hard labor that needs to be paid for, either to AI, or to human.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#447

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

[flagged]

> Eat a dick.

Could you please stop breaking the HN guidelines? Your account has unfortunately done that repeatedly, and we've asked you several times to stop.

Your comment would be just fine without that bit.

https://news.ycombinator.com/newsguidelines.html

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#448
post #378

Earlier quoted context omitted.

> This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. Am I misunderstanding something. I (the site owner) pay Cloudflare to do this. It is my fault this happens, not Cloudflare's.

You’re paying Cloudflare to not get DDoS-attacked or swamped by illegitimate requests. GP is implying that Cloudflare could do a better job of not blocking legitimate, benign requests.

Then we're all operating with very different definitions of legitimate or benign!

I've only ever seen a Cloudflare interstitial when viewing a page with my VPN on, for example -- something I'm happy about as a site owner and accept quite willingly as a VPN user knowing the kinds of abuse that occur over VPN.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#449

Earlier quoted context omitted.

But users depend on major sites like google [insert service] still and will prioritize their usage accordingly like limited minutes and texts back in the day, right?

Networking is so cheap, unless ISPs drastically inflate their price, users won’t care. The average American allegedly* downloads 650-700GB/month, or >20GB/day. 10MB is more than enough for a webpage (honestly, 1MB is usually enough), so that means on average, ISPs serve over 2000 webpages worth of data per day. And the average internet plan is allegedly** $73/month, or That’s cheap enough, wrapped in a monthly bill,…

Wait, so the ISPs do from taking $73/user home today to taking $0/user home tomorrow under this plan?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#450

So, this calls for a new type of honeytrap, content that appears to be human generated, and high quality, but subtly wrong, preferably on a commercially catastrophic way. Behind settings that prohibit commercial usage. It really shouldn't be hard to generate gigantic quantities of the stuff. Simulate old forum posts, or academic papers.

They did that too https://blog.cloudflare.com/ai-labyrinth/
Post reply on HN