Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

781–790 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#781

Earlier quoted context omitted.

If ads were more respectful I wouldn’t have to remove them. Alas they can’t help themselves and so I do. When ads were far less invasive, I had a lot more tolerance. Now they want my data, they want to play audio, video, hijack the content, page etc. Advertising scum can not be trusted to forever take more and more and more.

I also have ad-blockers for the same reason. However, if you don't support the people or companies producing the media you consume then don't be surprised when they go out of business.

> don't be surprised when they go out of business.

I’m ok with this. I support the media I truely want to see, and that media offers alternatives that are not ads.

For instance, I pay for YouTube premium. That said, many will not pay.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#783

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

> If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it.

Because the website has every right to block you or refuse access to you if you do that, just like an establishment has the right to refuse you access if you try to enter without a shirt, if you're denying them access to revenue that they predicated your access on.

Similarly, if you're using a user-agent the website doesn't like, they have the right to block you, or take action against that user-agent to prevent it from existing if they can't reliable detect it to block it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#784

Earlier quoted context omitted.

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

Who cares what Hacker News wants? You’re not obliged to participate in discussion. Am I supposed to spend money on Amazon.com when I visit the website just because Amazon wants me to?

It was a corollary example

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#785

Earlier quoted context omitted.

That's exactly why I said conciliation court. None of what you've outlined is required nor is it expensive. But, for each case, the defendant is still required to show up. I've successfully used conciliation court against large corporations in the past which is why I question it here. And while this should be able to be handled via legislation it won't be. Beyond that a workaround could force that to happen.

> conciliation court Sorry, I had never heard that term before. You would still have to show standing though. How would you try to prove that their violating your TOS cost you money?

Is it not viable to produce a work of art and say that this is free for humans, but not for bots and cannot be used for training and said violation cost X?

Again, I can't copy and distribute a game Microsoft rents to me. But if I do I can be found held accountable for a ridiculous amount of money. If it's my work of art the terms can dictate who doesn't need to pay and who does. If an LLM is consuming my work of art and now distributing it within their user base how is that not the same?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#786

Their test seems flawed: > We created multiple brand-new domains, similar to testexample.com and secretexample.com. These domains were newly purchased and had not yet been indexed by any search engine nor made publicly accessible in any discoverable way. We implemented a robots.txt file with directives to stop any respectful bots from accessing any part of a website: > We conducted an experiment by querying Perplexit…

robots.txt isn't even designed to stop recursive fetches. It is designed to ask nicely recursive fetchers not to recursively fetch. It comes from a time where site operators wanted their sites to be scraped by search engines, but not things like edit links and admin panels.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#787

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Think of it like tge telephone game.

Do you -really- want that much abstracrion?

Theres a bunch of nerds and capitalists about to rediscover GIGO

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#789

Earlier quoted context omitted.

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

Who cares what Hacker News wants? You’re not obliged to participate in discussion. Am I supposed to spend money on Amazon.com when I visit the website just because Amazon wants me to?

> You’re not obliged to participate in discussion.

Are website owners obligated to serve content to AI agents and/or LLM scrapers?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#790

Earlier quoted context omitted.

> conciliation court Sorry, I had never heard that term before. You would still have to show standing though. How would you try to prove that their violating your TOS cost you money?

Is it not viable to produce a work of art and say that this is free for humans, but not for bots and cannot be used for training and said violation cost X? Again, I can't copy and distribute a game Microsoft rents to me. But if I do I can be found held accountable for a ridiculous amount of money. If it's my work of art the terms can dictate who doesn't need to pay and who does. If an LLM is consuming my work of art…

These are arguments you would tell the judge. And the judge would almost certainly tell you 'this is the wrong venue for that. You are in small claims. I need an itemized list of monetary damages you have suffered before I can make a judgement.'
Post reply on HN