Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

501–510 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#501
post #401

Question for those in this thread who are okay with this: If I have endpoints that are computationally expensive server-side, what mechanism do you propose I could use to avoid being overwhelmed? The web will be a much worse place if such services are all forced behind captchas or logins.

How do you make the money you need to finance these computationally expensive server-side endpoints?

Maybe I'm a community-driven project funded by donations and volunteer time. Maybe I'm a local government with extremely limited IT budget and no in-house skills. Maybe I'm just some dude who maintains a hobby project that lives on a NUC under my desk.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#502

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Intellectual property laws are what creates the entitlement that someone else besides you can tell you what to do with the things Internet connected computers and phones download, because almost everything you download is copy of something a person created, therefore its copyrighted for the life of the author + 75 years or whatever by default.

Therefore artifices like "you don't have the right to view this website without ads" or "you can't use your phone, computer, or LLM to download or process this outside of my terms because copyright" become possible, institutionalizable, enforceable, and eventually unbypassable by technology.

If we reverted back to the Constitutional purpose of copyright (to Progress the Science and Useful Arts) then things might be more free. That's probably not happening in my lifetime or yours.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#503

Earlier quoted context omitted.

A place where you can lose you wallet and get it back with all the cash inside. The horror!!

[flagged]

Let me guess: A low violence society is bad because people get attacked and beat up?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#504
post #492

Earlier quoted context omitted.

Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?

Repeat after me - intentional discrimination of computer programs over humans is a good and praise worthy thing. We can and should make execution of computer programs harder and harder, even disproportionately so, if that makes lives of humans better and easier. LLM programs does not have human rights.

"if that makes lives of humans better" is doing a lot of heavy lifting, and remains to be explained.

Computer programs don't take actions, people do. If I use a web browser, or scrape some site to make an LLM, that's me doing it, not the program. And I have human rights.

If you think training LLMs should be illegal, just say that. If you think LLM companies are putting an undue strain on computer networks and they should be forced to pay for it, say that. But don't act like it's a virtue to try and capriciously gatekeep access to a public resource.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#505
post #485
post #415

Earlier quoted context omitted.

> No, he (Matthew) opted everyone in by default Now you're just lying. I checked several of my Cloudflare sites and none have it enabled by default: "No robots.txt file found. Consider enabling Cloudflare managed robots.txt or generate one for your website" "A robots.txt was found and is not managed by Cloudflare" "Instruct AI bot traffic with robots.txt" disabled

I think lying is a bit strong, I think they're potentially incorrect at worst. The Cloudflare blog post where they announced this a few weeks ago stated "Cloudflare, Inc. (NYSE: NET), the leading connectivity cloud company, today announced it is now the first Internet infrastructure provider to block AI crawlers accessing content without permission or compensation, by default." [1] I was also a bit confused by this w…

> I think lying is a bit strong, I think they're potentially incorrect at worst.

I understand that you're trying to be generous, but the claim that "Matthew opted everyone in by default" is flat out incorrect.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#507
post #39

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Unless I am misunderstanding you, you are talking about something different than the article. The article is talking about web-crawling. You are talking about local / personal LLM usage. No one has any problems with local / personal LLM usage. It's when Perplexity uses web crawlers that an issue arises.

Is the article really talking about crawling? Because in one of their screenshots where they ask information about the "honeypot" website you can see that the model requested pages from the website. But that is most definitely "fetching by proxy because I asked a question about the website" and not random crawling.

It is confusing.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#508

Earlier quoted context omitted.

> The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior The Big Tech bots provide proven value to most sites. They have also through the years proven themselves to respect robots.txt, including crawl speed directives. If you manage…

> The Big Tech bots provide proven value to most sites. They provide valeu for their companies. If you get some value from them it's just a side effect.

It goes without saying that they are profit-oriented. The point is that they historically offered a clear trade: let us crawl you, and we will refer traffic to you. An AI crawler does not provide clear value back. An AI user request agent might or might not provide enough clear value back for sites to want to participate. (Same goes for the search incumbents if they go all-in on LLM search results and don't refer much traffic back).

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#509
post #486

Earlier quoted context omitted.

Googlebot respects robots.txt. And Google doesn't use the fetched data from users of Chrome to supplement their search index (as a2128 is speculating that Perplexity might do when they fetch pages on the user's behalf).

Yes, but there's no way to say "allow indexing for search, but not for AI use", right?

But there is: https://developers.google.com/search/docs/crawling-indexing/...

There is an user agent for search that you can control in robots.txt.

    user-agent: Googlebot
There is another user agent for AI training.

    user-agent: Google-Extended

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#510
post #463

Earlier quoted context omitted.

Why bring up capitalism? I don't get it. What's stopping people from lying and cheating under any other system?

When lying and cheating doesn't get you ahead, there is no reason to do it.

You seriously think that mankind wasn't lying and cheating long before inventing capitalism?
Post reply on HN