Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

631–640 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#631
post #105

Seems a win. CF being internet police is a problem too but someone credible publicly shaming a company for shady scraping is good. Even if it just creates conversation Somehow this needs to go back to search era where all players at least attempt to behave. This scrapping Ddos stuff and I don’t care if it kills your site (while “borrowing” content) is unethical bullshit

Shaming doea not work in the era of "no shame".

Any better workable ideas that do work?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#632

Earlier quoted context omitted.

Yes, this is the crux of the matter. The "social contract" that has been established over the last 25+ years is that site owners don't mind their site being crawled reasonably provided that the indexing that results from it links back to their content. So when AltaVista/Yahoo/Google do it and then score and list your website, interspersing that with a few ads, then it's a sensible quid pro quo for everyone. LLM AI ou…

Anything but expanding copyright laws. Tbh, a pay per citation with an opt in database to add your info (think music streaming style monetization) would be reasonable to me. Not that I think it's a good scheme for music but I think it's fitting for web crawling. Though it does inevitably lead to enshitification. Pick your poison I guess.

The reason it works for music is because the people behind the databases have a team of lawyers that will come after you for violating copyright/performance legislation if you don’t pay your dues.

The argument that LLM outfits are using is that they are just exercising “fair use” / education rights to do an end run around copyright law. Without strengthening the rules on that I’m not sure I see how the database + team of lawyers approach would work.

But with that, sure, that’s an approach that seems to have legs in other contexts.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#633

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

I like the terminology "crawler" vs. "fetcher" to distinguish between mass scraping and something more targeted as a user agent. I've been working on AI agent detection recently (see https://stytch.com/blog/introducing-is-agent/ ) and I think there's genuine value in website owners being able to identify AI agents to e.g. nudge them towards scoped access flows instead of fully impersonating a user with no controls. O…

If Perplexity has millions of users, there’s no distinction between “mass fetching” and “mass crawling” — the snapshots of web pages will still be stored in Perplexity’s own crawl index.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#634
post #629

What if their “crawler” is just cheap human labor in some country with very low wages? Would that be allowed, because these are not robots?

What if they have significant robotic body parts? Or what if they make heavy use of automation processes and they barely click a button to index a page (so they just maniacally click all day long)?

What if robots.txt should refer to the ultimate beneficiaries... one which in this case would be the AI product that uses that content... to serve another ultimate beneficiary, a human user.

The problem here is obviously the higher prices for hosting the content, and less revenue for those that serve ads, have product placement on their sites, etc.

As long as robots.txt is about ethics/money and is enforced by morality, it doesn't matter who it refers to anyway.

Public-shaming enforcement might work in some cases though, but I doubt it will be that useful. We're talking about companies that have trained their AIs on IPs, and tried their best to later hide it. Does shame affect robots, or companies for that matter?

Cloudflare would very much like to be the middleman for monetary transactions between AI services and site owners (https://blog.cloudflare.com/introducing-pay-per-crawl/), but at the moment they don't have a law to hold their back, so articles like these are the best they got.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#635

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

I don't really see the issue.

The web admin should be able to block usages 1, 2 or 3 at their discretion. It's their website.

Similarly the user is free to try to engage via 1, 2, 3, or refuse to interact with the website entirely.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#637

I've built and run a personal search engine, that can do pretty much what perplexity does from a basic standpoint. Testing with friends it gets about 50/50 preference for their queries vs Perplexity. The engine can go and download pages for research. BUT, if it hits a captcha, or is otherwise blocked, then it bails out and moves on. It pisses me off that these companies are backed by billions in VC and they think the…

This sounds fascinating! Are you able to elaborate on what is different about yours vs Perp'lexity's?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#638
post #335

This is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the L…

> The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If you are the source I think they could make plenty of sense. As an example, I run a website where I've spent a lot of time documenting the history of a somewhat niche activity. Much of this information isn't available online anywhere else. As it happens I'm happy to let b…

> I think it's a reasonable stance to not want other companies to profit from my hard work

For me, the dividing line is whether someone else's profit is at my expense. If I sell a book, and someone starts hawking cheaper photocopies of it, that takes away my future sales. It's at my expense, and I'm harmed.

But if someone takes my book's story and writes song lyrics derived from it, I might feel a little envy (perhaps I've always wanted to be a songwriter), but I don't think I'd harbor ill will. I might even hope for the song to be successful, as it would surely drive further sales of my book.

It's human nature to covet someone else's success, but the fact is there was nothing stopping me (except talent) from writing the song.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#639

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Note that a book author cannot publish a book and then refuse to let libraries buy copies and lend them out. This was litigated 100+ years ago.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#640

Earlier quoted context omitted.

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

But I can send my personal shopper and you'll be none the wiser.

But the store owner can ask the personal shopper to leave, if e.g. they find out that they work for a personal shopper service.
Post reply on HN