Earlier quoted context omitted.
Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.
So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
411–420 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#412Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#413Earlier quoted context omitted.
> We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. I agree, but your idea below that is overly complicated. You can't micro-transact the whole internet. That idea feels like those episodes of Star Trek DS9 that take place on Feregenar - where you have to pay admission…
> You can't micro-transact the whole internet. I agree that end-users cannot handle micro transactions across the whole internet. That said, I would like to point out that most of the internet is blanketed in ads and ads involve tons of tiny quick auctions and micro transactions that occur on each page load. It is totally possible for a system to evolve involving tons of tiny transactions across page loads.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#414Earlier quoted context omitted.
Can't agree more, cloudflare is destroying the internet. We've entered the equivalent of when having McAffe antivirus was worse than having an actual virus because it slowed down your computer to much. These user hostile solutions have taken us back to dialup era page loading speeds for many sites, it's absurd that anyone thinks this is a service worth paying for.
So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#415[flagged]
Now you're just lying.
I checked several of my Cloudflare sites and none have it enabled by default:
"No robots.txt file found. Consider enabling Cloudflare managed robots.txt or generate one for your website"
"A robots.txt was found and is not managed by Cloudflare"
"Instruct AI bot traffic with robots.txt" disabled
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#416Earlier quoted context omitted.
The article did not test if the issue was specific to robots.txt or if it can not find other files. There is a difference between doing a poor summarization of data, and failing to even be able to get the data to summarize in the first place.
> specific to robots.txt > poor summarization of data I'm not really addressing the issue raised in the article. I am noting that the LLM, when asked, is either lying to the user or making a statement that it does not know to be true (that there is no robots.txt). This is way beyond poor summarization.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#417>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…
In theory retrieving a page on behalf of a user would be acceptable, but these are AI companies who have disregarded all norms surrounding copyright, etc. It would be stupid of them not to also save contents of the page and use it for future AI training or further crawling
Crawling is legal. Training is presumably legal. Long may the little guys do both.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#418Earlier quoted context omitted.
So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.
No they're supposed to allow scraping and information aggregation. That's the essence of the web: it's all text, crawlable, machine-readable (sort of) and parseable. Feel free to block ddos'es.
Also after starting the crawl, you can read about Aaron Swartz while waiting.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#419Earlier quoted context omitted.
So server owners are just supposed to bend over and take all the abuse they get from shitty bots and DDOS attacks and do nothing? That seems pretty unreasonable.
No, they're supposed to rally together and fight for better laws and enforcement of those laws. Which is, arguably, exactly what they've done just in a way that you and I don't like.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#420> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…