Question for those in this thread who are okay with this: If I have endpoints that are computationally expensive server-side, what mechanism do you propose I could use to avoid being overwhelmed? The web will be a much worse place if such services are all forced behind captchas or logins.
How do you make the money you need to finance these computationally expensive server-side endpoints?
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
501–510 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#502I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Therefore artifices like "you don't have the right to view this website without ads" or "you can't use your phone, computer, or LLM to download or process this outside of my terms because copyright" become possible, institutionalizable, enforceable, and eventually unbypassable by technology.
If we reverted back to the Constitutional purpose of copyright (to Progress the Science and Useful Arts) then things might be more free. That's probably not happening in my lifetime or yours.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#503Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#504Earlier quoted context omitted.
Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?
Repeat after me - intentional discrimination of computer programs over humans is a good and praise worthy thing. We can and should make execution of computer programs harder and harder, even disproportionately so, if that makes lives of humans better and easier. LLM programs does not have human rights.
Computer programs don't take actions, people do. If I use a web browser, or scrape some site to make an LLM, that's me doing it, not the program. And I have human rights.
If you think training LLMs should be illegal, just say that. If you think LLM companies are putting an undue strain on computer networks and they should be forced to pay for it, say that. But don't act like it's a virtue to try and capriciously gatekeep access to a public resource.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#505Earlier quoted context omitted.
> No, he (Matthew) opted everyone in by default Now you're just lying. I checked several of my Cloudflare sites and none have it enabled by default: "No robots.txt file found. Consider enabling Cloudflare managed robots.txt or generate one for your website" "A robots.txt was found and is not managed by Cloudflare" "Instruct AI bot traffic with robots.txt" disabled
I think lying is a bit strong, I think they're potentially incorrect at worst. The Cloudflare blog post where they announced this a few weeks ago stated "Cloudflare, Inc. (NYSE: NET), the leading connectivity cloud company, today announced it is now the first Internet infrastructure provider to block AI crawlers accessing content without permission or compensation, by default." [1] I was also a bit confused by this w…
I understand that you're trying to be generous, but the claim that "Matthew opted everyone in by default" is flat out incorrect.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#506Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#507I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Unless I am misunderstanding you, you are talking about something different than the article. The article is talking about web-crawling. You are talking about local / personal LLM usage. No one has any problems with local / personal LLM usage. It's when Perplexity uses web crawlers that an issue arises.
It is confusing.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#508Earlier quoted context omitted.
> The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior The Big Tech bots provide proven value to most sites. They have also through the years proven themselves to respect robots.txt, including crawl speed directives. If you manage…
> The Big Tech bots provide proven value to most sites. They provide valeu for their companies. If you get some value from them it's just a side effect.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#509Earlier quoted context omitted.
Googlebot respects robots.txt. And Google doesn't use the fetched data from users of Chrome to supplement their search index (as a2128 is speculating that Perplexity might do when they fetch pages on the user's behalf).
Yes, but there's no way to say "allow indexing for search, but not for AI use", right?
There is an user agent for search that you can control in robots.txt.
user-agent: Googlebot
There is another user agent for AI training. user-agent: Google-ExtendedRe: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#510Earlier quoted context omitted.
Why bring up capitalism? I don't get it. What's stopping people from lying and cheating under any other system?
When lying and cheating doesn't get you ahead, there is no reason to do it.