Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

771–780 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#771

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

1. Sometimes you should prove that you are human first.

I think the line is drawn at "on my behalf". The silent agreement of the web is that humans are served content via a browser, and robots are obeying rules. All we need to support this status quo is to perform data processing by ML models on a client's side, in the browser, the same way we rip out ads.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#772
Cloudflare's test was to setup a dummy domain that had never been indexed, and had blocks in the robots.txt and the firewall.

Then when they asked perplexity it came up with details about the 'exact' content (according to Cloudflare) but their attached screenshot shows the opposite, it shows some generic guesses about the domain ownership and some dynamic ads based on the domain name.

If Perplexity was stealthily visiting the dummy site they would have seen it, as the site was not indexed and no one else was visiting the site. Instead it appears they made assertions about general traffic, not their dummy site.

Its not very convincing.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#773

Earlier quoted context omitted.

Consent is also expressed through technical conventions. I, the website owner, express my intention through - for example - robots.txt. if you write a bot that specifically ignores it, you are violating consent. Likewise, I may prevent certain user-agents to visit my site. If you - say, an AI megacorp - are intentionally spoofing the user-agent to appear as a user, you are also violating consent.

I don't know how to make this any clearer. You - website owner - your consent does not matter . You are publishing information on the internet. I do not think you have a right to decide who is allowed to read it or not, or how they use what they read. You have exactly two legal rights: the right not to be DoS'd/hacked, and the right not to have your copyright infringed. Neither of those rights have anything to do wit…

You realize that consent in this case is just what we refer to as "authorization", right? And it is absolutely within any website operator's rights to only authorize their site for certain users and for certain purposes.

Websites are not "public resources"; site operators just mostly choose to allow the general public to access them. There's no legal requirement that they do so.

If you want anti-discrimination laws that apply to businesses to also cover bots, that is well outside of current law. A site operator can absolutely morally and legally decide they do not allow non-human visitors, just like a store can prohibit pets.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#774

Earlier quoted context omitted.

That's exactly why I said conciliation court. None of what you've outlined is required nor is it expensive. But, for each case, the defendant is still required to show up. I've successfully used conciliation court against large corporations in the past which is why I question it here. And while this should be able to be handled via legislation it won't be. Beyond that a workaround could force that to happen.

> conciliation court Sorry, I had never heard that term before. You would still have to show standing though. How would you try to prove that their violating your TOS cost you money?

Maybe you could say the increase in traffic increased your hosting costs by a penny or whatever.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#775
post #443

Earlier quoted context omitted.

Whoa, please don't post like this. We end up banning accounts that do. https://news.ycombinator.com/newsguidelines.html

Aw, alright. I thought it was a funny way to make the point and I figured the yo momma structure was traditional enough to not be taken as a proper insult. Heard tho.

Thanks for this. Now that you explain your intent, I see the joke. Unfortunately, it's too easy for the intent not to come across in these forsaken little text blobs that we're all limited to here. A lot of it boils down to the absence of voice tone and body language.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#776

Earlier quoted context omitted.

Hi, website operator here. I don't want my content to be accessible to you through Perplexity. I want my work to be freely available to any person who wants it. Feel free to transform my material as you see fit. Hell, do it with LLMs! I don't care. The LLM isn't the problem, it's what companies like Perplexity are doing with the LLM. Do not create commercial products that regurgitate my work as if it was your own. It…

no one cares about your shitty blog lol

You can't post like this here.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#777
post #289

C'mon CF. What are you doing? You are literally breaking the internet with your police behaviour. Starts to look like the Great Firewall.

Not affiliated with CF in any way. Respectfully disagree. Calling out bad actors is in the public interest.

It's not in my interest that a tech company from the US decides what a bad actor is.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#778
post #721
post #620

Earlier quoted context omitted.

> I made a stateful Internet implementation in Python earlier for proof-of-concept Is there a repo or some other form of public access? I'd like to see this.

it's not in a shareable state; is unsafe as-is. can share general idea and sample "webpage" files, though. the server ("lodge") passes JSON to the client from what are called .branch files. the client receives JSON, parses it, then builds the UI and state representation from the JSON, then stored in that client's memory (self.current_doc and self.page_state in python client). branches can invoke waterwheel (.ww) file…

I was right to ask, this seems extremely cool. Hit me up via mail [in bio] if you ever end up polishing it enough to share.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#779

Earlier quoted context omitted.

If magazines and newspapers were once able to be funded by native ads, so can websites. The spying industry doesn't want you to know this, but ads work without spying too - just look at all the IRL billboards still around.

I never said anything about spying. Magazines and newspapers were able to by funded by native ads because you couldn't auto-remove ads from their printed media and nobody could clone their content and give it away for free.

You can't remove ads that are part of a site's native HTML either - well, not easily, not without an AI determining what is an ad based on the content itself. The few ads I see despite uBlock are like that - something the website author themself included, and not by pulling it in from a different domain.

And those ads don't spy. They tend to be a jpg that functions as a link. That's why I mentioned spying.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#780
post #159
post #141

Earlier quoted context omitted.

We need to wake up and understand that all the information already uploaded is more or less a free web material, once taken through the lens of ML-somethings. With all the second, and third-order effects such as the fact that this changes completely the whole motivation, and consequence of open-source perhaps. It is also only a matter of time scrapers once again get through walls by twitter, reddit and alike. This is…

Reddit sold their data already. Twitter made thier own AI.

Precisely my point, and there is little if any evidence, there is anyone among the big players who puts peoples' rights before else by respecting licensing agreements before scrapping for training.

Indeed, Reddit sold their data the other thay GPT2 was announced, and it was very apparent why everyone closed their APIs in 2021-2023. Wonder what Aaron would've said about it.

Now we have walled gardens of information where people are allowed to plant, but never own the blossom.

Post reply on HN