Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

251–260 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#251

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Not only is it difficult to solve, it's the next step in the process of harvesting content to train AIs: companies will pay humans (probably in some flavor of "company scrip," such as extra queries on their AI engine) to install a browser extension that will piggy-back on their human access to sites and scrape the data from their human-controlled client. At the limit, this problem is the problem of "keeping secrets w…

> companies will pay humans (probably in some flavor of "company scrip," such as extra queries on their AI engine) to install a browser extension that will piggy-back on their human access to sites and scrape the data from their human-controlled client.

Proprietary web browsers are in a really good position to do something like this, especially if they offer a free VPN. The browser would connect to the "VPN servers", but it would be just to signal that this browser instance has an internet connection, while the requests are just proxied through another browser user.

That way the company that owns this browser gets a free network of residential IP address ready to make requests (in background) using a real web browser instance. If one of those background requests requires a CAPTCHA, they can just show it to the real user, e.g. the real user visits a Google page and they see a Cloudflare CAPTCHA, but that CAPTCHA is actually from one of the background requests (while lying in its UI and still showing the user a Google URL in the address bar).

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#252

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

There is a significant distinction between 2 and 3 that you glossed over. In 1 and 2, you the human may be forced to prove that you are human via a captcha. You are present at the time of the request. Once you’ve performed the exchange, then the HTML is on your computer and so you can do what you want to it.

In 3, although you do not specify, I assume you mean that a bot requests the page, as opposed to you visiting the page like in scenario 2 and then an LLM processes the downloaded data (similarly to an adblocker). It is the former case that is a problem, the latter case is much harder to stop and there is much less reason to stop it.

This is the distinction: is a human present at the time of request.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#253
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> The internet we knew was open and not trusted ... monopolistic behavior

Monopolistic is the wrong word, because you have the problem backwards. Cloudflare isnt helping Apple/Google... It's helping its paying consumers and those are the only services those consumers want to let through.

Do you know how I can predict that AI agents, the sort that end users use to accomplish real tasks, will never take off? Because the people your agent would interact with want your EYEBALLS for ads, build anti patterns on purpose, want to make it hard to unsubscribe, cancel, get a refund, do a return.

AI that is useful to people will fail. For the same reason that no one has great public API's any more. Because every public companies real customers are its stock holders, and the consumers are simply a source of revenue. One that is modeled, marked to, and manipulated all in the name of returns on investment.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#254
Like many other generative AI companies, Perplexity exploits the good faith of the old Internet by extracting the content created almost entirely by normal folks (i.e. those who depend on a wage for subsistence) and reproducing it for a profit while removing the creators from the loop - even when normal folks are explicitly asking them to not do this.

If you don't understand why this is at least slightly controversial I imagine you are not a normal folk.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#255
post #152

Earlier quoted context omitted.

They hate AI it seems. I don’t see them offering any AI products or embracing it in any way. Seems like they’ll get left behind in the AI race.

? https://developers.cloudflare.com/workers-ai/ ? https://ai.cloudflare.com/

Ah TIL. These are tiny models though but maybe it’s a good sign.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#256

Earlier quoted context omitted.

Well then. Seems like you would be a fool to not allow personal shoppers then. The point is the web is changing, and people use a different type of browser now. Ans that browser happens to be LLMs. Anybody complaining about the new browser has just not got it yet, or has and is trying to keep things the old way because they don’t know how or won’t change with the times. We have seen it before, Kodak, blockbuster, wha…

> Anybody complaining about the new browser has just not got it yet, or has and is trying to keep things the old way because they don’t know how or won’t change with the times. We have seen it before, Kodak, blockbuster, whatever. You say this as though all LLM/otherwise automated traffic is for the purposes of fulfilling a request made by a user 100% of the time which is just flatly on-its-face untrue. Companies mak…

Tech bros just respect money. Making money is very easy in the short term if you don't show ethics. Venture capitalism and the whole growth/indie hacking is focused around making money and making it fast.

Its a clear road for disaster. I am honestly surprised by how great Hackernews is, to that comparison where most people are sharing it for the love of the craft as an example. And for that hackernews holds a special place in my heart. (Slightly exaggerating to give it a thematic ending I suppose)

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#257
post #156

Earlier quoted context omitted.

> companies who want AI to recommend their products need to turn this off before it starts hurting them financially Content marketing, gamified SEO, and obtrusive ads significantly hurt the quality of Google search. For all its flaws, LLMs don’t feel this gamified yet. It’s disappointing that this is probably where we’re headed. But I hope OpenAI and Anthropic realize that this drop in search result quality might be…

This has already started with people using special tags also people making content just for llms.

I hope they realize Cloudflare opted them in to blocking LLMs.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#258
post #241

Earlier quoted context omitted.

Spam and DDOS are serious problems, it's not fair to suggest Cloudflare is just doing this to gatekeep the Internet for its own sake.

It's definitely not a DDOS when it's a single http request per year. I don't know if they do it on purpose but the fact is none of the big tech crawlers are limited.

This is most attributable to the fact that traffic is essentially anonymous so the source ip address is the best that a service can do if it's trying to protect an endpoint.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#259
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

> The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior

Plenty of site/service owners explicitly want Google, Meta and Apple bots (because they believe they have a symbiotic relationship with it) and don't want your bot because they view you as, most likely, parasitic.

Post reply on HN