Earlier quoted context omitted.
But users depend on major sites like google [insert service] still and will prioritize their usage accordingly like limited minutes and texts back in the day, right?
Networking is so cheap, unless ISPs drastically inflate their price, users won’t care. The average American allegedly* downloads 650-700GB/month, or >20GB/day. 10MB is more than enough for a webpage (honestly, 1MB is usually enough), so that means on average, ISPs serve over 2000 webpages worth of data per day. And the average internet plan is allegedly** $73/month, or That’s cheap enough, wrapped in a monthly bill,…
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
371–380 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#372Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#373Earlier quoted context omitted.
If there's an article you want to read, and the ToS says that in between reading each paragraph, you must switch to their YouTube channel and look at their ads about cat food for 5 minutes, are your going to do that?
Hacker News has collectively answered this question by consistently voting up the archive.is links in the comments of every paywalled article posted here.
I never really minded magazine ads or even television ads. They might have tried to make me associate boobs with a brand of soda but they didn't data mine my life and track me everywhere. I'd much rather have old fashioned manipulation than pervasive and dangerous surveillance capitalism.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#374Earlier quoted context omitted.
I think it's an issue of scale. The next step in your progression here might be: If / when people have personal research bots that go and look for answers across a number of sites, requesting many pages much faster than humans do - what's the tipping point? Is personal web crawling ok? What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to gather information regularly to…
Doesn't o3 sort of already do this? Whenever I ask it something, it makes it look like it simultaneously opens 3-8 pages (something a human can't do). Seems like a reasonable stance would be something like "Following the no crawl directive is especially necessary when navigating websites faster than humans can." > What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to ga…
ChatGPT probably uses a cache though. Theoretically, the average load on the original sites could be far less than users accessing them directly.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#375Earlier quoted context omitted.
That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.
You ignore the issue of motivation. Most web content exists because someone wants to make money on it. If the content creator can't do that, they will stop producing content. These AI web crawlers (Google, Perplexity, etc) are self-cannibalizing robots. They eat the goose that laid the golden egg for breakfast, and lose money doing it most of the time. If something isn't done to incentivize content creators again eve…
While I don’t blame creators for using ads now, I don’t think they’re a long-term solution. Ads are already blocked when people visit the site with ad blockers, which are becoming more popular. Obvious sponsored content may be blocked with the ads, and non-obvious sponsored content turns these “creators” into “shills” who are inauthentic and untrustworthy. Even without Google summaries, ad revenue may decrease over time as advertisers realize they aren’t effective or want more profit; even if it doesn’t, it’s my personal opinion that society should decrease the overall amount of ads.
Not everyone creates only for money, the best only create for enough money to sustain themselves. A long-term solution is to expand art funding (e.g. creators apply for grants with their ideas and, if accepted, get paid a fixed rate to execute them) or UBI. Then media can be redistributed, remixed, etc. without impacting creators’ finances.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#376> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
> the "perplexity bots" arent crawling websites, they fetch URLs that the users explicitly asked. This shouldnt count as something that needs robots.txt access. It's not a robot randomly crawling, it's the user asking for a specific page and basically a shortcut for copy/pasting the content You say "shouldn't" here, but why? There seems to be a fundamental conflict between two groups who each assert they have "rights…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#377> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
Spam and DDOS are serious problems, it's not fair to suggest Cloudflare is just doing this to gatekeep the Internet for its own sake.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#378> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
> This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. Am I misunderstanding something. I (the site owner) pay Cloudflare to do this. It is my fault this happens, not Cloudflare's.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#379I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
1. I actually disagree. I think teasers should be free but websites should charge micropayments for their content. Here is how it can be done seamlessly, without individuals making decisions to pay every minute: https://qbix.com/ecosystem 2. This also intersects with copyright law. Ingesting content to your servers en masse through automation and transforming it there is not the same as giving people a tool (like Saf…
I expect a lot more paywalls for valuable content. General information is commoditized and offered in aggregated form through models. But when an AI is fetching information for you from a website, the publisher is still paying the cost of producing that content and hosting that content. The AI models are increasing the cost of hosting the content and then they are also removing the value of producing the content since you are just essentially offering value to the AI model. The user never sees your site.
I know Ads are unpopular here, but the truth is that is how publishers were compensated for your attention. When an AI model views the information that a publisher produces, then modifies it from its published form, and removes all ad content. Then you now have increased costs for producers, reduced compensation in producing content (since they are not getting ad traffic), and the content isn't even delivered in the original form.
The end result is that publishers now have to paywall their content.
Maybe an interesting middle-ground is if the AI Model companies compensated for content that they access similar to how Spotify compensates for plays of music. So if an AI model uses information from your site, they pay that publisher a fraction of a cent. People pay the AI models, and the AI models distribute that to the producers of content that feed and add value to the models.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#380> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…