Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

511–520 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#511

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

In theory, couldn't the LLM access the content on your browser and it's cache, rather than interacting with the website directly? Browser automation directly related to user activity (prefetch etc) seems qualitatively different to me. Similarly, refusing to download content or modifying content after it's already in my browser is also qualitatively different. That all seems fair-use-y. I'm not sure there's a technica…

> couldn't the LLM access the content on your browser

Yes, orbit, a now deprecated firefox extension by mozilla was doing that. This way you could also use it to summarise content that would not be available to a third party (eg sth in google docs).

You can still sort of do the same with the ai chatbot panel in firefox, sort of, but ctrl+A>right click>AI chatbot>summarise.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#512
post #86

In unrelated news, Fedora (the Linux distro) has been taken down by a DDoS today which I understand is AI-scraping related: https://pagure.io/fedora-infrastructure/issue/12703

The last comment there now reads:

"It was actually a caching issue on our end. ;) I just fixed it a few min ago..."

Lets not go on a witch hunt and blame everything on AI scrapers.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#513

Earlier quoted context omitted.

Sure. There's lots of things you could do, but you don't do them because they are wrong. Might does not make right.

How is it wrong to send my personal shopper? How is it wrong to have an agent act directly on my behalf? It's like saying a web browser that is customized in any way is wrong. If one configures their browser to eagerly load links so that their next click is instant, is that now wrong?

Here's a good rule of thumb: if you have to do it without other people knowing, because otherwise they wouldn't let you do it: chances are it's a bad thing to do.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#514
post #463

Earlier quoted context omitted.

Why bring up capitalism? I don't get it. What's stopping people from lying and cheating under any other system?

When lying and cheating doesn't get you ahead, there is no reason to do it.

If we look at any communist society, the only way to get ahead was lying and cheating. China was forced to adopt capitalist markets to deal with this, hence why modern China hardly resembles the USSR, Cuba, Venezuela, or Laos.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#515

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

It's because they own the content so they get to set the terms.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#516
post #470

Earlier quoted context omitted.

[flagged]

> High trust is prima facie incompatible with capitalism Quite compatible > If you want a high trust society, you don't want capitalism. There is nothing at all in capitalism that would prevent a high level of trust in society. > Capitalism is inherently low trust But that's not true. The thing about capitalism is that it's RESILENT to low trust. It does not require low levels of trust, but is capable of functioning…

https://theonion.com/this-war-will-destabilize-the-entire-mi...

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#517

Earlier quoted context omitted.

Too bad. Build a bigger store or publish this information so we don't need 10,000 personal shoppers. Was this not the whole point of having a website? Who distorted that simple idea into the garbage websites we have now?

Weird take. The store doesn't owe your personal shippers anything.

That's fair, but if there's enough of supply and demand for this to get traction (and online shopping is bug, and autonomous agents are sort of trending), this conflict of interest paired with a no-compromise "we don't own you anything" attitude is bound to escalate in an arms race. And YMMV but I don't like where that race may possibly end.

If store businesses at least partially relies on obscurity of information that can be solved through automated means (e.g. storefronts tend to push visitors towards products they don't want, and buyer agents are fighting that and looking for something buyers instructed them) just playing this cat and mouse game of blocking agents, finding workarounds, and repeating the cycle is only creating perverse technological contraptions that neither party is really interested in - but both are circumstantially forced to invest into.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#518

Earlier quoted context omitted.

http is neutral. it's up to the client to ignore robots.txt You can block IP's at the host level but there's pretty easy ways around that with proxy networks.

> http is neutral. Who misled you with that statement?

Http doesnt have emotions or thought last time I checked.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#519

Earlier quoted context omitted.

http is neutral. it's up to the client to ignore robots.txt You can block IP's at the host level but there's pretty easy ways around that with proxy networks.

> http is neutral. Who misled you with that statement?

IETF?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#520

Earlier quoted context omitted.

There is a significant distinction between 2 and 3 that you glossed over. In 1 and 2, you the human may be forced to prove that you are human via a captcha. You are present at the time of the request. Once you’ve performed the exchange, then the HTML is on your computer and so you can do what you want to it. In 3, although you do not specify, I assume you mean that a bot requests the page, as opposed to you visiting…

To me it's even simpler: 3 is a request made from another ip address that isn't directly yours. Why should an LLM request that acts exactly like a VPN request be treated differently from a VPN request?

Yeah, I also find the analogy about "agent on behalf of the user interacting with a website" weak, because it is not about "an agent", it is a 3rd party service that actually takes content from a website, processes it and serves it to the user (even with their own ads?). It is more akin to, let's say, a scammy website that copies content from other legit websites and serves their own ads, than software running on the user's computer.

There are legitimate reasons to do that, of course. Maybe I am trying to find info about some niche topic or how to do X, I ask an llm, the llm goes through some search results, a lot of which is search engine optimised crap, finds the relevant info and answers my question.

But if I wrote articles in a news site, I am supported by ads or subscriptions and see my visits plummel because people, who would usually google about topic X and then visit my website that I wrote about X, were now reading the google summary that appeared when googling about topic X, based on my article, maybe I would have less motivation to continue writing.

The only end result possible in such a scenario is that everything commercial of some quality being heavily paywalled, some tiny amount of free and open small web, and a huge amount of AI generated slop, because the value of an article in the open internet is now so low that only AI can produce it (economically, time-wise) efficiently enough.

Post reply on HN