Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

271–280 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#271
post #246
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

Hah still remember the old “solving the internet with hate” idea from Zed Shaw in the glory days of Ruby on Rails.

https://weblog.masukomi.org/2018/03/25/zed-shaws-utu-saving-...

I do believe we will end there eventually, with the emerging tech like Brazil’s and India’s payment architectures it should be a possibility in the coming decades

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#273

Earlier quoted context omitted.

[flagged]

[flagged]

Isn't that the system that we are already living in?

Democracy in its american form or even at many others show almost complete paralysis of the entire system basically if bad actors infiltrate it (Looking at ya donald)

It is honestly a little sad since conservatives usually think of their society as this high trust society and they were the ones who primarily voted and are being taken advantaged of by the few untrustworthy individuals.

Politics is a cult/religion and you can't prove me otherwise.

I vote because I vote for lesser evil not for greater good. I do think that frankly, both the parties or just most parties in every nation are just so short of reality but I created a discord server of 100 people and I can see how I can't manage 100 people and so maybe I expect so much from the govt.

I used to focus so much on history and politics but its bloody mess and there is no good or bad. Now I just feel like going into the woods and into the darks living alone, maybe coding.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#274
post #199
post #92

Earlier quoted context omitted.

>I think the intelligent conclusion would be that the people you are looking at have more nuanced beliefs than you initially thought. You don't seem to reject my claim that for many, principles took a backseat to "does this help or hurt evil corporations". If that's what passes as "nuance" to you, then sure. >Talking about broken brains is often just mediocre projecting To be clear, that part is metaphorical/hyperbol…

People never agreed DOSing a site to take copyright material was acceptable. Many people did not have a problem with taking copyright material in a respectful way that didn't kill the resource. LLMs are killing the resource. This isn't a corporation vs person issue. No issue with an llm having my content but big issue with my server being down because llms are hammering the same page over and over.

>People never agreed DOSing a site to take copyright material was acceptable. Many people did not have a problem with taking copyright material in a respectful way that didn't kill the resource.

Has it be shown that perplexity engages in "DOSing"? I've heard of anecdotes of AI bots gone amuck, and maybe that's what's happening here, but cloudflare hasn't really shown that. All they did was set up a robots.txt and shown that perplexity bypassed it. There's probably archivers out there that's using youtube-dl to hit download from youtube at 1+Gbit/s, tens of times more than a typical viewer is downloading. Does that mean it's fair game to point to a random instance of someone using youtube-dl and characterizing that as "DOSing"?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#276
post #92

Earlier quoted context omitted.

I think the intelligent conclusion would be that the people you are looking at have more nuanced beliefs than you initially thought. Talking about broken brains is often just mediocre projecting

>I think the intelligent conclusion would be that the people you are looking at have more nuanced beliefs than you initially thought. You don't seem to reject my claim that for many, principles took a backseat to "does this help or hurt evil corporations". If that's what passes as "nuance" to you, then sure. >Talking about broken brains is often just mediocre projecting To be clear, that part is metaphorical/hyperbol…

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#278
post #246
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

Wouldn't this lead to pirated page clones where customer pays less for same-ish content, and less, all the way down to essentially free?

Because I as an user would be glad to have "free sites only" filter, and then just steal content :))

But it's an interesting idea and thought experiment.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#279

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

There is a significant distinction between 2 and 3 that you glossed over. In 1 and 2, you the human may be forced to prove that you are human via a captcha. You are present at the time of the request. Once you’ve performed the exchange, then the HTML is on your computer and so you can do what you want to it. In 3, although you do not specify, I assume you mean that a bot requests the page, as opposed to you visiting…

To me it's even simpler: 3 is a request made from another ip address that isn't directly yours. Why should an LLM request that acts exactly like a VPN request be treated differently from a VPN request?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#280
Perplexity claims that you can “use the following robots.txt tags to manage how their sites and content interact with Perplexity.” https://docs.perplexity.ai/guides/bots

Their fetcher (not crawler) has user agent Perplexity-User. Since the fetching is user-requested, it ignores robots.txt . In the article, it discusses how blocking the “Perplexity-User” user agent doesn’t actually work, and how perplexity uses an anonymous user agent to avoid being blocked.

Post reply on HN