Earlier quoted context omitted.
> Otherwise there is literally no reason for them to make any of it available on the open web This is the hypothesis I always personally find fascinating in light of the army of semi-anonymous Wikipedia volunteers continuously gathering and curating information without pay. If it became functionally impossible to upsell a little information for more paid information, I'm sure some people would stop creating informati…
Any information that requires something approximating a full-time job worth of effort to produce will necessarily go away, barring the small number of independently wealthy creators. Existing subject-matter experts who blog for fun may or may not stick around, depending on what part of it is “fun” for them. While some must derive satisfaction from increasing the total sum of human knowledge, others are probably blogg…
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
651–660 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#652Earlier quoted context omitted.
Ofttimes people are sufficiently anti-ad that this point won't resonate well. I'm personally mostly in that camp in that with relatively few exceptions money seems to make the parts of the web I care about worse (it's hard to replace passion, and wading through SEO-optimized AI drivel to find a good site is a lot of work). Giving them concrete examples of sites which would go away can help make your point. E.g., Shel…
But even your example gets worse with AI potentially - the "upsell" of his blog isn't paid posts but more subscribers so there will be thankful readers, a few donators, people talking about it. If the only interface becomes an AI summary of his work without credit, it's much more likely he stops writing as it'll seem like he's just screaming into the void
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#653I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
The problem is not about personal use. It's about big corporations scrapping billions of pages to make money.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#654Earlier quoted context omitted.
If I were DOSing your blog, you'd ask me to stop. I run server ops for multiple online communities that are being severely negatively impacted and DOSed by these AI scrapers, and we have very few ways to stop them.
That is a problem, but is not related to my comment. The person I'm replying to is acting as if consent is a relevant aspect of the public web, I am saying it isn't. That is not the same as saying "you can do whatever you want to a public server". It is just that what you are allowed to do is not related to the arbitrary whim of the server operator.
Likewise, I may prevent certain user-agents to visit my site. If you - say, an AI megacorp - are intentionally spoofing the user-agent to appear as a user, you are also violating consent.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#655Earlier quoted context omitted.
To me it's even simpler: 3 is a request made from another ip address that isn't directly yours. Why should an LLM request that acts exactly like a VPN request be treated differently from a VPN request?
Yeah, I also find the analogy about "agent on behalf of the user interacting with a website" weak, because it is not about "an agent", it is a 3rd party service that actually takes content from a website, processes it and serves it to the user (even with their own ads?). It is more akin to, let's say, a scammy website that copies content from other legit websites and serves their own ads, than software running on the…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#656Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#657I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
The problem in your logic is that all points starts wit "I". You're not the only stakeholder in any of those interactions. There's you, a mediator (search or LLM), and the website owner. The website owner (or its users) basically do all the work and provide all the value. They produce the content and carry the costs and risks. The pre-LLM "deal" was that at least some traffic was sent their way, which helps with reac…
They are already paying, it is the way they are paying that causes the mess. When you buy a product, some fraction of the price is the ad budget that gets then distributed to websites showing ads. Therefore there is also nothing wrong with blocking ads, they have already been paid for, whether you look at them or not. The ad budget will end up somewhere as long as not everyone is blocking all ads, only the distribution will get skewed. Which admittedly might be a problem for websites that have a user base that is disproportionally likely to use ad blockers.
Paying for content directly has the problem that you can only pay for a selected few websites before the amount you have to pay becomes unreasonable. If you read one article on a hundred different websites, you can not realistically pay for a hundred subscriptions that are all priced as if you spent all your time on a single website. Nobody has yet succeeded in creating a web wide payment method that only charges you for the content that you actually consume and is frictionless enough to actually work, i.e. does not force you to make a conscious payment decisions for a few cents or maybe even only fractions of a cent for every link you click and is not a privacy nightmare collecting all the links you click for billing purposes.
Also if you directly pay for content, you will pay twice - you will pay for the subscription and you will still pay into the ad budget with all the stuff you buy.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#658Earlier quoted context omitted.
On a more human level, I think it's bleak that someone who makes a blog just to share stuff for fun is going to have most of his traffic be scrapers that distill, distort, and reheat whatever he's writing before serving it to potential readers.
I don't think it's bleak, just the opposite. If someone writes valuable stuff on a blog almost nobody finds, that's a tragedy. If LLM's can process the information and provide it to people in conversations where it will be most helpful, where they never would have found it otherwise, then that's amazing! If all you're trying to do is help people with the information you've discovered, why do you care if it's delivere…
This is why I care if my ideas are presented to others by an LLM (that maybe cites me in some % of cases) or directly to a human. There is already a difference between a human visiting my space (acknowledging it as such) to read and learn information and being a footnote reference that may or may not be read or opened, without an immediate understanding of which information comes from me.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#659Earlier quoted context omitted.
Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.
Foo news wants you to visit the site, look at the main page, watch the ads, click on them and buy the products advertised by third parties which will give money to Foo news in exchange for this service. And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads. They claim that since they are free to not buy an advertised product, why would they…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#660I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
The problem is not about personal use. It's about big corporations scrapping billions of pages to make money.
Circa 2008 I worked for a startup that would scrape Google Books and a variety of other sources for public domain content to then print via Amazon’s Print-on-Demand services. Google, of course, didn’t like this and introduced a Captcha not very long after we started scraping.
So we hired a team of underemployed / unemployed English majors during the height of the recession, paid them $10 per hour to type in Captchas all day long and we downloaded their full corpus anyways.
If Google can’t win, you won’t either.