Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

191–200 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#191
post #139

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

It's all about scale. The impact of your personal shopper is insignificant unless you manage to scale it up into a business where everyone has a personal shopper by default.

How is everyone having a personal shopper a problem of scale? I was going to shop myself, but I sent someone else to do it for me.

At this moment I am using Perplexity's Comet browser to take a spotify playlist and add all the tracks to my youtube music playlist. I love it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#192

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

If you as a human are well behaved, that is absolutely fine.

If you as a human spam the shit out of my website and waste my resources, I will block you.

If you as a human use an agent (or browser or extension or external program) that modifies network requests on your behalf, but doesn't act as a massive leech, you're still welcome.

If you as a human use an agent (or browser or extension or external program) that wrecks my website, I will block you and the agent you rode in on.

Nobody would mind if you had an LLM that intelligently knew what pages contain what (because it had a web crawler backed index that refreshes at a respectful rate, and identifies itself accurately as a robot and follows robots.txt), and even if it needed to make an instantaneous request for you at the time of a pertinent query, it still identified itself as a bot and was still respectful... there would be no problem.

The problem is that LLMs are run by stupid, greedy, evil people who don't give the slightest shit what resources they use up on the hosts they're sucking data from. They don't care what the URLs are, what the site owner wants to keep you away from. They download massive static files hundreds or thousands of times a day, not even doing a HEAD to see that the file hasn't changed in 12 years. They straight up ignore robots.txt and in fact use it as a template of what to go for first. It's like hearing an old man say "I need time to stand up because of this problem with my kneecaps" and thinking "right, I best go for his kneecaps because he's weak there"

There are plenty of open crawler datasets, they should be using those... but they don't, they think that doesn't differentiate them enough from others using "fresher" data, so they crawl even the smallest sites dozens of times a day in case those small sites got updated. Their badly written software is wrecking sites, and they don't care about the wreckage. Not their problem.

The people who run these agents, LLMs, whatever, have broken every rule of decency in crawling, and they're now deliberately evading checks, to try and run away from the repercussions of their actions. They are bad actors and need to be stopped. It's like the fuckwads who scorch the planet mining bitcoin; there's so much money flowing in the market for AI, that they feel they have to fuck over everyone else, as soon as possible, otherwise they won't get that big flow of money. They have zero ethics. They have to be stopped before their human behaviour destroys the entire internet.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#193

Earlier quoted context omitted.

It’s possible to violate all sorts of social norms. Societies that celebrate people that do so are on the far opposite end of the spectrum from high trust ones. They are rather unpleasant.

[flagged]

A place where you can lose you wallet and get it back with all the cash inside.

The horror!!

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#194
At work I'm considering blocking all the ip prefixes announced by ASNs owned by Microsoft and other companies known for their LLMs. At this point it seems like the only viable solutions.

LLM scrapers bots are starting to make up a lot of our egress traffic and that is starting to weight on our bills.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#195
post #139

Earlier quoted context omitted.

It's all about scale. The impact of your personal shopper is insignificant unless you manage to scale it up into a business where everyone has a personal shopper by default.

Well then. Seems like you would be a fool to not allow personal shoppers then. The point is the web is changing, and people use a different type of browser now. Ans that browser happens to be LLMs. Anybody complaining about the new browser has just not got it yet, or has and is trying to keep things the old way because they don’t know how or won’t change with the times. We have seen it before, Kodak, blockbuster, wha…

Some people use LLMs to search. Other people still prefer going to the actual websites. I'm not going to use an LLM to give me a list of the latest HN posts or NY Times articles, for example.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#197
post #20

Earlier quoted context omitted.

> I think most people would draw a distinction between the two, and would at least agree the latter is more acceptable than the former. No. I should be able to control which automated retrieval tools can scrape my site, regardless of who commands it. We can play cat and mouse all day, but I control the content and I will always win: I can just take it down when annoyed badly enough. Then nobody gets the content, and…

> Then nobody gets the content, and we can all thank upstanding companies like Perplexity for that collapse of trust. But they didn't take down the content, you did. When people running websites take down content because people use Firefox with ad-blockers, I don't blame Firefox either, I blame the website.

>But they didn't take down the content, you did.

That skips the part about one party's unique role in the abuse of trust.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#198

Not sure I would consider a user copy-pasting an URL being a bot. Should curl be considered a bot too? What's the difference?

> Should curl be considered a bot too? What's the difference?

Perplexity definitely does:

    $ curl -sI https://www.perplexity.ai | head -1
    HTTP/2 403

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#199
post #92

Earlier quoted context omitted.

I think the intelligent conclusion would be that the people you are looking at have more nuanced beliefs than you initially thought. Talking about broken brains is often just mediocre projecting

>I think the intelligent conclusion would be that the people you are looking at have more nuanced beliefs than you initially thought. You don't seem to reject my claim that for many, principles took a backseat to "does this help or hurt evil corporations". If that's what passes as "nuance" to you, then sure. >Talking about broken brains is often just mediocre projecting To be clear, that part is metaphorical/hyperbol…

People never agreed DOSing a site to take copyright material was acceptable. Many people did not have a problem with taking copyright material in a respectful way that didn't kill the resource.

LLMs are killing the resource. This isn't a corporation vs person issue. No issue with an llm having my content but big issue with my server being down because llms are hammering the same page over and over.

Post reply on HN