I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
How about I open a proxy, replace all ads with my ads, redirect the content to you and we share the ad revenue?
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
11–20 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#12I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? It’s a different UI, sure, but there should be no discrimination towards it as there should be no discrimination towards, say, Links terminal browser, or some exotic Firefox derivative.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#13If you put info on the web, it should be available to everyone or everything with access.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#14Sorry CF, give up. the courts are on our sides here
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#15I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
If the LLM were running this sort of thing at the user's explicit request this would be fine. The problem is training. Every AI startup on the planet right now is aggressively crawling everything that will let them crawl. The server isn't seeing occasional summaries from interested users, but thousands upon thousands of bots repeatedly requesting every link they can find as fast as they can.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#16I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
If the LLM were running this sort of thing at the user's explicit request this would be fine. The problem is training. Every AI startup on the planet right now is aggressively crawling everything that will let them crawl. The server isn't seeing occasional summaries from interested users, but thousands upon thousands of bots repeatedly requesting every link they can find as fast as they can.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#17If you put info on the web, it should be available to everyone or everything with access.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#18I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
If I put time and effort into a website and it's content, I should expect no compensation despite bearing all costs.
Is that something everyone would agree with?
The internet should be entirely behind paywalls, besides content that is already provided ad free.
Is that something everyone would agree with?
I think the problem you need to be thinking about is "How can the internet work if no one wants to pay anything for anything?"
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#19Cloudflare screaming into the void desperate to insert themselves as a middleman, in a market ( that they will never succeed in creating) where they extort scrapers for access to websites they cover. Sorry CF, give up. the courts are on our sides here
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#20>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…
> I think most people would draw a distinction between the two, and would at least agree the latter is more acceptable than the former. No. I should be able to control which automated retrieval tools can scrape my site, regardless of who commands it. We can play cat and mouse all day, but I control the content and I will always win: I can just take it down when annoyed badly enough. Then nobody gets the content, and…
But they didn't take down the content, you did. When people running websites take down content because people use Firefox with ad-blockers, I don't blame Firefox either, I blame the website.