Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

541–550 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#542
post #443

Earlier quoted context omitted.

[flagged]

Whoa, please don't post like this. We end up banning accounts that do. https://news.ycombinator.com/newsguidelines.html

Aw, alright. I thought it was a funny way to make the point and I figured the yo momma structure was traditional enough to not be taken as a proper insult. Heard tho.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#543
It's time to stop blocking crawlers and using captchas and start building web sites that are intentionally AI-friendly by design. Even before the modern LLMs, anti-scraper measures apparently were primarily befitting Google whose scrapers were the most common exception.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#544
post #484

Earlier quoted context omitted.

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

These are more like a store putting up a billboard or catalog and asking people to turn off their meta AI glasses nearby because the store doesn't want AI translating it on your behalf as a tourist.

It is not because the store does not expend any resources on the singular instance of the glasses capturing the content of the billboard. Web requests cost money.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#545
Previously it was all sniper and sneaker bots scanning websites for product availability and attempting purchases continuously to snipe when it comes back online.

Now, it's a gazillion of AI crawlers and python crawlers, MCP servers that offer the same feature to anyone "building (personal workflow) automation" incl. bypass of various, standard protection mechanisms.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#546

Earlier quoted context omitted.

It seems that a 403 makes you sad though.

iproyal.com makes me smile again

And Cloudflare makes you cry. See, it's not neutral. Glad you learned something today. The more one learns everyday, the less stupid you become.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#547

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

> why would the LLM accessing the website on my behalf be in a different legal category as my Firefox web browser accessing the website on my behalf?

is it just on your behalf? or is it on Perplexity's behalf? are they not archiving the pages to train on?

it's the difference between using Google Chrome vs. Chrome beaming full page snapshots to train Gemini on.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#548
post #492

Earlier quoted context omitted.

Repeat after me - intentional discrimination of computer programs over humans is a good and praise worthy thing. We can and should make execution of computer programs harder and harder, even disproportionately so, if that makes lives of humans better and easier. LLM programs does not have human rights.

"if that makes lives of humans better" is doing a lot of heavy lifting, and remains to be explained. Computer programs don't take actions, people do. If I use a web browser, or scrape some site to make an LLM, that's me doing it, not the program. And I have human rights. If you think training LLMs should be illegal, just say that. If you think LLM companies are putting an undue strain on computer networks and they sh…

I was being unclear, but that's on me since it was intentional. But to clarify my stance - I'm against accidental or intentional equalization of programs with individual humans and using that as a foundation for all kinds of negative (imo) corporate behavior later.

For example - humans can learn, programs can't. The "learning" cop out for LLM-corpos shouldn't be accepted by anyone, let alone by law. Humans have a fair use carve out of the copyright laws, not because it's something axiomatic, it's because some humans with empathy have forced others to allow all humans a leeway in legally using other's IP works. Just because such law exist for humans, doesn't mean that random computer programs should be applicable to it. Scraping web for LLMs should not be considered "fair use" because a) it is clearly not (commercialized later) and b) programs aren't humans and don't have equal rights.

And the list goes on. Now, I do get that train has long left the station and we are all collectively living in the anecdote about stealing a bicycle and asking god for forgiveness. But that doesn't mean I agree with this state. I'm just shouting my displeasure towards that passing train cause I'm weird like that. It's like with climate change - we are doing nothing that matters, no one discusses what really matters and I just accepted that nothing will really change. Doesn't mean I like the situation.

PS: tl;dr - LLMs clearly should be legal, it's just simple code is all. LLM corporations who steal IP content without compensation to the authors should be illegal, but of course they won't ever be.

PPS: there is a huge, gigantic gap between a single person scraping a few thousand pages for a personal use, maybe even some small local commercial use (though that's a grey area already) and a billion dollar megacorp, intent on destroying everything of value for humans in the internet for profit.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#549

Earlier quoted context omitted.

Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…

As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?

technical limitations / data poisoning measures

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#550
post #39

Earlier quoted context omitted.

Unless I am misunderstanding you, you are talking about something different than the article. The article is talking about web-crawling. You are talking about local / personal LLM usage. No one has any problems with local / personal LLM usage. It's when Perplexity uses web crawlers that an issue arises.

You probably need a computer that costs $250,000 or more to run the kind of LLM that Perplexity uses, but with batching it costs pennies to have the same LLM fetch a page for you, summarize the content, and tell you what is on it. And the power usage similarly, running the LLM for a single user will cost you a huge amount of money relative to the power it takes in a cloud environment with many users. Perplexity's "we…

does Perplexity store crawled pages for training?
Post reply on HN