Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

411–420 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#411
post #353
post #327

Earlier quoted context omitted.

Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

No they're supposed to allow scraping and information aggregation. That's the essence of the web: it's all text, crawlable, machine-readable (sort of) and parseable. Feel free to block ddos'es.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#412
post #257
post #156

Earlier quoted context omitted.

This has already started with people using special tags also people making content just for llms.

I hope they realize Cloudflare opted them in to blocking LLMs.

I hope you realise that lying is bad.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#413

Earlier quoted context omitted.

> We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. I agree, but your idea below that is overly complicated. You can't micro-transact the whole internet. That idea feels like those episodes of Star Trek DS9 that take place on Feregenar - where you have to pay admission…

> You can't micro-transact the whole internet. I agree that end-users cannot handle micro transactions across the whole internet. That said, I would like to point out that most of the internet is blanketed in ads and ads involve tons of tiny quick auctions and micro transactions that occur on each page load. It is totally possible for a system to evolve involving tons of tiny transactions across page loads.

Remember Flattr?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#414
post #353
post #327

Earlier quoted context omitted.

Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

No, they're supposed to rally together and fight for better laws and enforcement of those laws. Which is, arguably, exactly what they've done just in a way that you and I don't like.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#415
post #123

[flagged]

> No, he (Matthew) opted everyone in by default

Now you're just lying.

I checked several of my Cloudflare sites and none have it enabled by default:

"No robots.txt file found. Consider enabling Cloudflare managed robots.txt or generate one for your website"

"A robots.txt was found and is not managed by Cloudflare"

"Instruct AI bot traffic with robots.txt" disabled

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#416
post #157

Earlier quoted context omitted.

The article did not test if the issue was specific to robots.txt or if it can not find other files. There is a difference between doing a poor summarization of data, and failing to even be able to get the data to summarize in the first place.

> specific to robots.txt > poor summarization of data I'm not really addressing the issue raised in the article. I am noting that the LLM, when asked, is either lying to the user or making a statement that it does not know to be true (that there is no robots.txt). This is way beyond poor summarization.

I would say it's orthogonal to it. LLMs being unable to judge their capabilities is a separate issue to summarization quality.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#417
post #84
post #3

>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…

In theory retrieving a page on behalf of a user would be acceptable, but these are AI companies who have disregarded all norms surrounding copyright, etc. It would be stupid of them not to also save contents of the page and use it for future AI training or further crawling

If you allow Googlebot to crawl your website and train Gemini, but you don't allow smaller AI companies to do the same thing, then you're contributing to Google's hegemony. Given that AI is likely to be an increasingly important part of society in the future, that kind of discrimination is anti-social. I don't want a future where everything is run by Google even more than it currently is.

Crawling is legal. Training is presumably legal. Long may the little guys do both.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#418
post #411
post #353

Earlier quoted context omitted.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

No they're supposed to allow scraping and information aggregation. That's the essence of the web: it's all text, crawlable, machine-readable (sort of) and parseable. Feel free to block ddos'es.

Feel free to crawl paywalled sites and republish them with discoverable links.

Also after starting the crawl, you can read about Aaron Swartz while waiting.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#419
post #353

Earlier quoted context omitted.

So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.

No, they're supposed to rally together and fight for better laws and enforcement of those laws. Which is, arguably, exactly what they've done just in a way that you and I don't like.

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#420
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

I could not keep my website up without Cloudflare given the level of bot and AI crawlers hammering things. I try whenever to do challenges, but sometimes I have to block entire AS blocks.
Post reply on HN