The web will be a much worse place if such services are all forced behind captchas or logins.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
401–410 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#402> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#403AI companies continuing to have problems with the concept of "consent" is increasingly alarming god help us if they ever manage to build anything more than shitty chatbots
Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#404Why single out Perplexity? Pretty much no crawler out there fetches robots.txt. robots.txt is not a blocking mechanism; it's a hint to indicate which parts of a site might be of interest to indexing. People started using robots.txt to lie and declare things like no part of their site is interesting, and so of course that gets ignored.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#405Earlier quoted context omitted.
Wouldn't this lead to pirated page clones where customer pays less for same-ish content, and less, all the way down to essentially free? Because I as an user would be glad to have "free sites only" filter, and then just steal content :)) But it's an interesting idea and thought experiment.
That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#406what machine learning algorithms are they using? time to deploy them onto our websites
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#407> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#408I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Websites should be able to request payment. Who cares if it is a human or an agent of a human if it is paying for the request?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#409> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
I crawl 3000 RSS feeds once a week. Let me tell you! Cloudflare sucks. What business is it of theirs to block something that is meant to be accessed by everyone. Like an RSS feeds. FU Cloudflare.
If they want the RSS feeds to be accessible then they should configure it to allow those requests.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#410Earlier quoted context omitted.
That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.
You ignore the issue of motivation. Most web content exists because someone wants to make money on it. If the content creator can't do that, they will stop producing content. These AI web crawlers (Google, Perplexity, etc) are self-cannibalizing robots. They eat the goose that laid the golden egg for breakfast, and lose money doing it most of the time. If something isn't done to incentivize content creators again eve…
If only the one in ten thousand with something to share are left standing to share it, no manufactured content, that's a fine thing.