Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

661–670 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#661
post #585

Earlier quoted context omitted.

With all the crypto development how come we haven't got to HTTP/1.1 402 Payment Required WWW-price: 0.0000001 BTC, 0.000001 ETH, 0.00001 DOGE > You are less likely to participate in discussion you (or AI on your behalf) paid instead. Many sites would probably like it better.

It's not a development problem, it's an adoption problem. Publishers are desperate to sell us on a $20+/month subscription, they don't want to offer convenient affordable access to single articles.

$20/month would be nice if it wasn't a tier with less ads. I want no ads, and full-text rss feeds (because I want to use my clients to read). It's like how Netflix refuses to build a basic search and filter, or Spotify refuses to an actual library manager. They don't want you in control of your consumption.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#662

Earlier quoted context omitted.

That is a problem, but is not related to my comment. The person I'm replying to is acting as if consent is a relevant aspect of the public web, I am saying it isn't. That is not the same as saying "you can do whatever you want to a public server". It is just that what you are allowed to do is not related to the arbitrary whim of the server operator.

Consent is also expressed through technical conventions. I, the website owner, express my intention through - for example - robots.txt. if you write a bot that specifically ignores it, you are violating consent. Likewise, I may prevent certain user-agents to visit my site. If you - say, an AI megacorp - are intentionally spoofing the user-agent to appear as a user, you are also violating consent.

I don't know how to make this any clearer. You - website owner - your consent does not matter. You are publishing information on the internet. I do not think you have a right to decide who is allowed to read it or not, or how they use what they read. You have exactly two legal rights: the right not to be DoS'd/hacked, and the right not to have your copyright infringed. Neither of those rights have anything to do with what you "consent" to. It is a bad thing to connect consent - the arbitrary, capricious whim of a website operator - to access of a public resource. Consent is for people, for your body, for your relationships. It is not a magic spell to give you arbitrary control over people.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#663
post #642

We (humanity) need to invent a simple GPLv3 style license “You can derive any data on the data you see here, any derived data you sell or share should mention this place as a source and is subject to the same copyright as the source”. This will imply scraped datasets should become public and the law enforcement bodies will be able to work in an established framework to fight copyright and license crimes. Just blockin…

Because scrapers would certainly comply with that

/s

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#664
post #146

It's ironic Perplexity itself blocks crawlers: $ curl -sI https://www.perplexity.ai | head -1 HTTP/2 403 Edit: trying to fake a browser user agent with curl also doesn't work, they're using a more sophisticated method to detect crawlers.

Try this then: https://github.com/lwthiker/curl-impersonate

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#665
post #95

AI companies continuing to have problems with the concept of "consent" is increasingly alarming god help us if they ever manage to build anything more than shitty chatbots

Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?

Yes. Asking for consent is built into the HTTP protocol. The issue at hand here is that Perplexity scrapers lie about who they are by providing a false user agent. Thus consent was given on a false pretense.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#666
post #537

Earlier quoted context omitted.

yea but those are not open sites, try imposing that on an open site you'd want to actually attract human traffic to

see, literally, reddit requiring teenagers to open their mouth and roll their heads around to enter.

I heard that it could be easily bypass through realistic 3D human game model with basic mouth open and head tilt animation, even gmod can do such thing.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#667

Earlier quoted context omitted.

Consent is also expressed through technical conventions. I, the website owner, express my intention through - for example - robots.txt. if you write a bot that specifically ignores it, you are violating consent. Likewise, I may prevent certain user-agents to visit my site. If you - say, an AI megacorp - are intentionally spoofing the user-agent to appear as a user, you are also violating consent.

I don't know how to make this any clearer. You - website owner - your consent does not matter . You are publishing information on the internet. I do not think you have a right to decide who is allowed to read it or not, or how they use what they read. You have exactly two legal rights: the right not to be DoS'd/hacked, and the right not to have your copyright infringed. Neither of those rights have anything to do wit…

You can repeat it, but we fundamentally disagree, it's not a matter of understanding.

Fundamentally it's not true that the moment I publish something on the internet, I lose control of who can consume my intellectual property. Licensing, for example, is a way we regulate the way that code or prose can be consumed even if public.

Also expressing my consent is not in any way a way to control others, is a way to control my ideas, my writing, my [whatever] and people are not automatically entitled to it because it's published on the internet.

So overall I understand your position, but I so much disagree with it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#668
post #316

Earlier quoted context omitted.

Yeah I'm not so sure about that. If Perplexity are visiting that page on your behalf to give you some information and aren't doing anything else with it, and just throw away that data afterwards, then you may have a point. As a site owner, I feel it's still my decision what I do and don't let you do, because you're visiting a page that I own and serve. But if, as I suspect, Perplexity are visiting that page and then…

Perplexity can then just ask the user to copy/paste the page content. That should be legal , it's what the user wants. The cases are equivalent

I can’t copy/paste the content of a book or a movie or music, that’s piracy.

But when a trillion dollar industry does it, its okay?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#669
post #646

Maybe we can just configure webservers to block anyone who requests robots.txt, regular browsers don't do it, but robots do to get list of urls to crawl (while ignoring rules). Just create simple PHP/CGI script that adds client IP addres to iptables once /robots.txt is accessed.

One way to easily bypass is to let external services fetching robots.txt (archive.org, GitHub actions, etc...) to cache it and either expose through separate apis/webhook/manual download to the actual scrape server.

robots txt file size is usually small and would not alert external services.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#670
post #463

Earlier quoted context omitted.

When lying and cheating doesn't get you ahead, there is no reason to do it.

If we look at any communist society, the only way to get ahead was lying and cheating. China was forced to adopt capitalist markets to deal with this, hence why modern China hardly resembles the USSR, Cuba, Venezuela, or Laos.

Communist with a capital C.

I've never seen a stateless, classless, moneyless society. It may be impossible.

Post reply on HN