Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

391–400 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#391
post #105

Seems a win. CF being internet police is a problem too but someone credible publicly shaming a company for shady scraping is good. Even if it just creates conversation Somehow this needs to go back to search era where all players at least attempt to behave. This scrapping Ddos stuff and I don’t care if it kills your site (while “borrowing” content) is unethical bullshit

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#392

Earlier quoted context omitted.

In the same token the personal shoppers don't owe the store anything either.

Then they can't complain if they're barred entry.

http is neutral. it's up to the client to ignore robots.txt

You can block IP's at the host level but there's pretty easy ways around that with proxy networks.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#393
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior The Big Tech bots provide proven value to most sites. They have also through the years proven themselves to respect robots.txt, including crawl speed directives. If you manage…

> The Big Tech bots provide proven value to most sites.

They provide valeu for their companies. If you get some value from them it's just a side effect.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#394
post #289

C'mon CF. What are you doing? You are literally breaking the internet with your police behaviour. Starts to look like the Great Firewall.

Not affiliated with CF in any way. Respectfully disagree. Calling out bad actors is in the public interest.

CF is a bad actor. They ruin internet. They own more and more parts of it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#396
post #326

Earlier quoted context omitted.

> We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. But it's done through a bait and switch. They serve the full article to Google, which allows Google to show you excerpts that you have to pay for. It would be better if Google shows something like PAYMENT REQUIRED on…

> They serve the full article to Google, which allows Google to show you excerpts that you have to pay for. I'm old enough to remember when that was grounds for getting your site removed from Google results - "cloaking" was against the rules. You couldn't return one result for Googlebot, and another for humans. No idea when they stopped doing that, but they obviously have let go of that principle.

I remember that too, along with high-profile punishments for sites that were keyword stuffing (IIRC a couple of decades ago BMW were completely unlisted for a time for this reason).

I think it died largely because it became impossible top police with any reliability, and being strict about it would remove too much from Google's index because many sites are not easily indexable without them providing a “this is the version without all the extra round-trips for ad impressions and maybe a login needed” variant to common search engines.

Applying the rule strictly would mean that sites implementing PoW tricks like Anubis to reduce unwanted bot traffic would not be included in the index if they serve to Google without the PoW step.

I can't say I like that this has been legitimised even for the (arguably more common) deliberate bait & switch tricks is something I don't like, but (I think) I understand why the rule was allowed to slide.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#397
Has anyone bothered to properly quantify the worst case load (i.e., requests per second) that has been incurred by these scraping tools? I recall a post on HN a few weeks/months ago about something similar, but it seemed very light on figures.

It seems to me that ~50% of the discourse occurring around AI providers involves the idea that a machine reading webpages on a regular schedule is tantamount to a DDOS attack. The other half seems to be regarding IP and capitalism concerns - which seem like far more viable arguments.

If someone requesting your site map once per day is crippling operations, the simplest solution is to make the service not run like shit. There is a point where your web server becomes so fast you stop caring about locking everyone into a draconian content prison. If you can serve an average page in 200uS and your competition takes 200ms to do it, you have roughly 1000x the capacity to mitigate an aggressive scraper (or actual DDOS attack) in terms of CPU time.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#399
post #327

Earlier quoted context omitted.

Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.

Ethics-free organizations and individuals like Perplexity are why Cloudflare exists . If you have a better way to solve the problems that they solve, the marketplace would reward you handsomely.

Do you think users shouldn't get to have user agents or that "content farm ads scaffold" as a business model has a right to be viable? Forcing users to reward either stance seems unsustainable.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#400
post #246

Earlier quoted context omitted.

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

A scary observation in light of another front page article right now: https://news.ycombinator.com/item?id=44783566 If pages can't be served for free, all internet content is at the mercy of payment processors and their ideas of "brand safety".

That's already a deep problem for all of society. If we don't want that to be an ongoing issue, we need to make sure money is a neutral infrastructure.

It doesn't just apply to the web, it applies to literally everything that we spend money on via a third party service. Which is... most everything these days.

Post reply on HN