> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
Spam and DDOS are serious problems, it's not fair to suggest Cloudflare is just doing this to gatekeep the Internet for its own sake.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
241–250 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#242Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#243Their test seems flawed: > We created multiple brand-new domains, similar to testexample.com and secretexample.com. These domains were newly purchased and had not yet been indexed by any search engine nor made publicly accessible in any discoverable way. We implemented a robots.txt file with directives to stop any respectful bots from accessing any part of a website: > We conducted an experiment by querying Perplexit…
> > We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains. This response was unexpected, as we had taken all necessary precautions to prevent this data from being retrievable by their crawlers. Right, I'm confused why CloudFlare is confused. You a…
Right, and the domain was configured to disallow crawlers, but Perplexity crawled it anyway. I am really struggling to see how this is hard to understand. If you mean to say "I don't think there is anything wrong with ignoring robots.txt" then just say that. Don't pretend they didn't make it clear what they're objecting to, because they spell it out repeatedly.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#244No amount of robots.txt or walled-gardening is going to be sufficient to impede generative AI improvement: common crawl and other data dumps are sufficiently large, not to mention easier to acquire and process, that the backlash against AI companies crawling folks' web pages is meaningless.
Cloudflare and other companies are leveraging outrage to acquire more users, which is fine... users want to feel like AI companies aren't going to get their data.
The faster that AI companies are excluded from categories of data, the faster they will shift to categories from which they're not excluded.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#245> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
So you just came here to bitch about Cloudflare? It's wild to even comment on this thread if this does not make sense to you.
They're building a search index. Every AI is going to struggle at being a tool to find websites & business listings without a search index.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#246> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources.
And the thing is... people already pay for internet. They pay their ISP. So people are perfectly happy to pay for resources that they consume on the Internet, and they already have an infrastructure for doing so.
I feel like the answer is that all web requests should come with a price tag, and the ISP that is delivering the data is responsible for paying that price tag and then charging the downstream user.
It's also easy to ratelimit. The ISP will just count the price tag as 'bytes'. So your price could be 100 MB or whatever (independent of how large the response is), and if your internet is 100 mbps, the ISP will stall out the request for 8 seconds, and then make it. If the user aborts the request before the page loads, the ISP won't send the request to the server and no resources are consumed.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#247Earlier quoted context omitted.
> Otherwise there is literally no reason for them to make any of it available on the open web This is the hypothesis I always personally find fascinating in light of the army of semi-anonymous Wikipedia volunteers continuously gathering and curating information without pay. If it became functionally impossible to upsell a little information for more paid information, I'm sure some people would stop creating informati…
Any information that requires something approximating a full-time job worth of effort to produce will necessarily go away, barring the small number of independently wealthy creators. Existing subject-matter experts who blog for fun may or may not stick around, depending on what part of it is “fun” for them. While some must derive satisfaction from increasing the total sum of human knowledge, others are probably blogg…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#248I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Definitely don't agree. I don't think you should be shown the content, if for example:
1. You're in a country the site owner doesn't want to do business in.
2. You've installed an ad blocker or other tool that the site owner doesn't want you to use.
3. The site owner has otherwise identified you as someone they don't want visiting their site.
You are welcome to try to fool them into giving you the content but it's not your right to get it.