Earlier quoted context omitted.
> Otherwise there is literally no reason for them to make any of it available on the open web This is the hypothesis I always personally find fascinating in light of the army of semi-anonymous Wikipedia volunteers continuously gathering and curating information without pay. If it became functionally impossible to upsell a little information for more paid information, I'm sure some people would stop creating informati…
Any information that requires something approximating a full-time job worth of effort to produce will necessarily go away, barring the small number of independently wealthy creators. Existing subject-matter experts who blog for fun may or may not stick around, depending on what part of it is “fun” for them. While some must derive satisfaction from increasing the total sum of human knowledge, others are probably blogg…
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
471–480 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#472>Perplexity spokesperson Jesse Dwyer dismissed Cloudflare’s blog post as a “sales pitch,” adding in an email to TechCrunch that the screenshots in the post “show that no content was accessed.” In a follow-up email, Dwyer claimed the bot named in the Cloudflare blog “isn’t even ours.”
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#473Earlier quoted context omitted.
Networking is so cheap, unless ISPs drastically inflate their price, users won’t care. The average American allegedly* downloads 650-700GB/month, or >20GB/day. 10MB is more than enough for a webpage (honestly, 1MB is usually enough), so that means on average, ISPs serve over 2000 webpages worth of data per day. And the average internet plan is allegedly** $73/month, or That’s cheap enough, wrapped in a monthly bill,…
Wait, so the ISPs do from taking $73/user home today to taking $0/user home tomorrow under this plan?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#474Earlier quoted context omitted.
Why bring up capitalism? I don't get it. What's stopping people from lying and cheating under any other system?
When lying and cheating doesn't get you ahead, there is no reason to do it.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#475Earlier quoted context omitted.
My first reaction: This solution would basically kill what little remaining fun there is to be had browsing the Internet and all but assure no new sites/smaller players will ever see traffic. Curious to hear other perspectives here. Maybe I’m over reacting/misunderstanding.
If site operators can’t afford the costs of keeping sites up in the face of AI scraping, the new/smaller sites are gone anyway.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#476Earlier quoted context omitted.
> We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. I agree, but your idea below that is overly complicated. You can't micro-transact the whole internet. That idea feels like those episodes of Star Trek DS9 that take place on Feregenar - where you have to pay admission…
> You can't micro-transact the whole internet. I agree that end-users cannot handle micro transactions across the whole internet. That said, I would like to point out that most of the internet is blanketed in ads and ads involve tons of tiny quick auctions and micro transactions that occur on each page load. It is totally possible for a system to evolve involving tons of tiny transactions across page loads.
The lengths Meta and the like go to in order to maximize clickthroughs...
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#477I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#478This is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the L…
> The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If you are the source I think they could make plenty of sense. As an example, I run a website where I've spent a lot of time documenting the history of a somewhat niche activity. Much of this information isn't available online anywhere else. As it happens I'm happy to let b…
Imagine someone at another company reads your site, and it informs a strategic decision they make at the company to make money around the niche activity you're talking about. And they make lots of money they wouldn't have otherwise. That's totally legal and totally ethical as well.
The reality is, if you do hard work and make the results public, well you've made them public. People and corporations are free to profit off the facts you've made public, and they should be. There are certain limited copyright protections (they can't sell large swathes of your words verbatim), but that's all.
So the idea that you don't want companies to profit from your hard work is unreasonable, if you make it public. If you don't want that to happen, don't make anything public.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#479Earlier quoted context omitted.
But I can send my personal shopper and you'll be none the wiser.
It’s possible to violate all sorts of social norms. Societies that celebrate people that do so are on the far opposite end of the spectrum from high trust ones. They are rather unpleasant.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#480Earlier quoted context omitted.
Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?
I am not told I cannot access. And, yes, I would, because I'd be breaking the law otherwise.
No you wouldn't be. Even if someone tells you not to visit your site, you have every legal right to continue visiting it, at least in the US.
Under common interpretation of the CFAA, there needs to be a formal mechanism of authorized access. E.g. you could be charged if you hacked into a password-protected area of someone's site. But if you're merely told "hey bro don't visit my site", that's not going to reach the required legal threshold.
Which is why crawlers aren't breaking the law. If you want to restrict authorization, you need to actually implement that as a mechanism by creating logins, restricting content to logged-in users, and not giving logins to crawlers.