Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

431–440 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#431
post #353

Earlier quoted context omitted.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

No, they're supposed to rally together and fight for better laws and enforcement of those laws. Which is, arguably, exactly what they've done just in a way that you and I don't like.

What kind of laws and enforcement would stop a foreign actor from effectively DDoSing your site? What if the actor has (illegally) hacked tech-illiterate users so they have domestic residential IP addresses?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#432

Earlier quoted context omitted.

Ethics-free organizations and individuals like Perplexity are why Cloudflare exists . If you have a better way to solve the problems that they solve, the marketplace would reward you handsomely.

Do you think users shouldn't get to have user agents or that "content farm ads scaffold" as a business model has a right to be viable? Forcing users to reward either stance seems unsustainable.

> Do you think users shouldn't get to have user agents or that "content farm ads scaffold" as a business model has a right to be viable?

Users should get to have authenticated, anonymous proxy user agents. Because companies like Perplexity just ignore `robots.txt`, maybe something like Private Access Tokens (PATs) with a new class for autonomous agents could be a solution for this.

By "content farm ads scaffold", I'm not sure if you had Perplexity and their ads business in mind, or those crappy little single-serving garbage sites. In any case, they shouldn't be treated differently. I have no problem with the business model, other than that the scam only works because it's currently trivial to parasitically strip-mine and monetize other people's IP.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#433

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

This is 100% incorrect.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#434
post #353
post #327

Earlier quoted context omitted.

Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

Unreasonable is to use such incompetent companies like Cloudflare, which are absolutely incapable of distinguishing between the normal usage of a Web site by humans and DDOS attacks or accesses done by bots.

Only this week I have witnessed several dozen cases when Cloudflare has blocked normal Web page accesses without any possible correct reason, and this besides the normal annoyance of slowing every single access to any page on their "protected" sites with a bot check popup window.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#435
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

Ironically, cloudflare is also the reason OpenAI agent mode with web use isn’t very usable right now. Every second time I asked it to do a mundane task like checking me in for a flight it couldn’t because of cloudflare.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#436
post #338

Earlier quoted context omitted.

> The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior Plenty of site/service owners explicitly want Google, Meta and Apple bots (because they believe they have a symbiotic relationship with it) and don't want your bot because the…

they didnt seem to mind when openai et al. took all their content to train LLMs when they were still parasites that didn't have a symbiotic relationship. This thinking is kind of too pro-monopolist for me

That’s a good thing. You want an LLM to know about product or service you are selling and promote it to its users. Getting into the training data is the new SEO.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#437

Earlier quoted context omitted.

The line can be extremely blurry (that's putting it mildly), and "the latter" is not the only reason people use CF (actually, I wouldn't be surprised at all if it wasn't even the biggest reason).

The reason people use Cloudflare is because they provide free CDN, and we have at least 10 years of content marketing out there telling aspiring bloggers that, if they use a CDN in front of their website, their shitty WordPress website hosted on a shady shared hosting will become fast.

well they aren't wrong

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#438
post #423
post #353

Earlier quoted context omitted.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

There is a difference between blocking abusive behavior and blocking all bots. No one really cared about bot scraping to this degree before AI scraping for training purposes became a concern. This is fearmongering by Cloudflare for website maintainers who haven't figured out how to adapt to the AI era so they'll buy more Cloudflare.

> No one really cared about bot scraping to this degree before AI scraping for training purposes became a concern. This is fearmongering by Cloudflare for website maintainers who haven't figured out how to adapt to the AI era so they'll buy more Cloudflare.

I think this is an overly harsh take. I run a fairly niche website which collates some info which isn't available anywhere else on the internet. As it happens I don't mind companies scraping the content, but I could totally undrestand if someone didn't want a company profiting from their work in that way. No one is under an obligation to provide a free service to AI companies.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#439
post #327

Earlier quoted context omitted.

Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.

Ethics-free organizations and individuals like Perplexity are why Cloudflare exists . If you have a better way to solve the problems that they solve, the marketplace would reward you handsomely.

While the existence of Perplexity may justify the existence of Cloudflare, it does not justify the incompetence of Cloudflare, which is unable to distinguish accesses done by Perplexity and the like from normal accesses done by humans, who use those sites exactly for the purpose they exist, so there cannot be any excuse for the failure of Cloudflare to recognize this.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#440
post #246

Earlier quoted context omitted.

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

> We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. I agree, but your idea below that is overly complicated. You can't micro-transact the whole internet. That idea feels like those episodes of Star Trek DS9 that take place on Feregenar - where you have to pay admission…

The presented solution has invisible UX via layering it into existing metered billing.

And, the whole internet is already micro-transactioned! Every page with ads is doing a bidding war and spending money on your attention. The only person not allowed to bid is yourself!

Post reply on HN