Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

761–770 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#761

Earlier quoted context omitted.

LLMs are not "someone", LLMs are something, and they don't "read content", they by definition acquire and reuse that content (for example, by summarizing it), as part of their product. So here the consent is indeed about what can be done with the data. In general, it's absolutely the norm that public websites (I.e., unauthenticated) restrict even who can access the data. The simplest example that comes to mind is geo…

LLMs are indeed not "someone". They are programs, like web browsers, acting on user instruction. The user is a person. I am only talking about people - I never said that an LLM does anything of its own volition. > The simplest example that comes to mind is geoblocking. Do you think it is alright to geoblock people, for arbitrary reasons? It is one thing when GDPR imposes a legal obligation on you for serving content…

> like web browsers, acting on user instruction.

No, they are not like browsers. The browser access my content in a transparent way. An LLM reuses the information and acts as an opaque intermediary which - maybe - will at most add a reference to my content.

> I never said that an LLM does anything of its own volition

It doesn't matter why it does what it does, it matters what it does. Your previous comment stressed the idea that it's possible to regulate _what can be done_ with my intellectual property (licensing), but not who can access it, once made it public. What I am saying is that this is exactly the case for LLMs, who _use_ my intellectual property, they are not a tool to _access_ it (like a browser).

> Do you think it is alright to geoblock people, for arbitrary reasons?

Yes. Why wouldn't it be? And if you believe it's not, where do you draw the line? Once you share a picture with your partner, everyone has the right to see it? Or if you share it with your group of friends? Or if you share it on a private social media profile (where you have acquaintances)? When does the audience turn from "a restricted group" to "everyone"? Or why would it be different with my blog? If I want my blog accessible only from my country, I can absolutely do that and there is nothing wrong with it at all. Nobody is entitled to my intellectual property. Obviously I am playing devil's advocate, but this was to say that the fact that something is public, doesn't mean it's unrestricted. And don't get me started on "the spirit of the internet". I can't imagine something breaking that spirit more than LLMs acting as interface between people and the other people on the internet. That spirit is gone, and belongs to a time when the internet was tiny. When OpenAI and company will respect the "spirit of the internet", maybe I will think about doing the same.

> As far as you, the service owner, are concerned, they are simply fetching the content for the user. It is none of your business what the user and the AI company go on to do with "your content".

No, as far as I am concerned the program can take my information, summarize, change, distort, misinterpret it and then present it back to its user. This can happen with or without the user ever knowing that the information can from me. Considering this equal to the user accessing the information is something I simply will not concede and is a fundamental disagreement between us, from which many other disagreements stems.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#762
post #535

Earlier quoted context omitted.

On a more human level, I think it's bleak that someone who makes a blog just to share stuff for fun is going to have most of his traffic be scrapers that distill, distort, and reheat whatever he's writing before serving it to potential readers.

I don't think it's bleak, just the opposite. If someone writes valuable stuff on a blog almost nobody finds, that's a tragedy. If LLM's can process the information and provide it to people in conversations where it will be most helpful, where they never would have found it otherwise, then that's amazing! If all you're trying to do is help people with the information you've discovered, why do you care if it's delivere…

In the grand scheme of things, I guess it's good to have an impact, even an indirect one, but come on, we're talking about human beings here.

Even if someone were to do it out of sheer passion without a care for financial gains, I'm sure they'd still appreciate basic validation and recognition. That's like the cheapest form of payment you could give for someone's work.

I don't understand why "actually, you're egotistical if you dare to desire recognition for stuff you put love and effort to" is such a common argument in those discussions. People are treated like machines that should swallow their pride and sense of self for the greater good, while on the other end, there is a (not saying YOU in particular did it) push to humanize LLMs.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#763
post #206
post #152

Earlier quoted context omitted.

They hate AI it seems. I don’t see them offering any AI products or embracing it in any way. Seems like they’ll get left behind in the AI race.

If they managed to enforce the pay-per-scrape, that would be a huge revenue, bigger than AdSense

Huge revenue sure, but at what cost? Models should be able to be trained on anything a human can read and see without paying.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#764
post #415
post #123

[flagged]

> No, he (Matthew) opted everyone in by default Now you're just lying. I checked several of my Cloudflare sites and none have it enabled by default: "No robots.txt file found. Consider enabling Cloudflare managed robots.txt or generate one for your website" "A robots.txt was found and is not managed by Cloudflare" "Instruct AI bot traffic with robots.txt" disabled

I recently created a new Cloudflare account for a project I’m working on and moved two domains into it, and the settings were both on by default without asking me about it at all. The original press release specifically mentioned enabling it by default.

> Cloudflare, Inc. (NYSE: NET), the leading connectivity cloud company, today announced it is now the first Internet infrastructure provider to block AI crawlers accessing content without permission or compensation, *by default*.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#765

Earlier quoted context omitted.

I know that ads are based on impressions as I told you before, but my money still has to end up somewhere even if I am using an ad blocker. So where does it end up if not as ad revenue on some websites? You must not confuse the people paying for the ads and in turn for the ad revenue of websites by buying stuff with the people deciding how that money gets distributed among all the websites by looking at ads. We can e…

You’re describing socialism (wealth redistribution to be exact). At this point, just make that money a tax and give it to the publishers directly. Cut out the middlemen.

Well, what is the difference, the ad budget fraction of the price is like a tax. I think given a choice most people would prefer to get their stuff a bit cheaper and not contribute to the ad budget. But we pay it and then the companies hand the money out to various parties to display ads creating the possibility of running a business on ad revenue. And in many cases I can ignore ads, I can not look at billboards, I can switch to a different channel during the commercial break, I can flip over the ad pages in newspapers and magazines but they still get paid. Only on the internet have we decided to only pay for ads when somebody actually looks at them. I just asked for the same thing on the internet, pay for the inclusion on the website, whether someone actually sees it or not. Not sure how that is socialism and wealth redistribute.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#766
post #679

Earlier quoted context omitted.

Because scrapers would certainly comply with that /s

More like have easier to assess legality status.

A cost sink that has no upside? How on earth would scrapers stop themselves saying yes?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#767

Earlier quoted context omitted.

I never said anything about spying. Magazines and newspapers were able to by funded by native ads because you couldn't auto-remove ads from their printed media and nobody could clone their content and give it away for free.

If ads were more respectful I wouldn’t have to remove them. Alas they can’t help themselves and so I do. When ads were far less invasive, I had a lot more tolerance. Now they want my data, they want to play audio, video, hijack the content, page etc. Advertising scum can not be trusted to forever take more and more and more.

I also have ad-blockers for the same reason. However, if you don't support the people or companies producing the media you consume then don't be surprised when they go out of business.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#768

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

We have a faceted search that creates billions of unique URLs by combinations of the facets. As such, we block all crawlers from it in robots.txt, which saves us AND them from a bunch of pointless indexing load. But a stealth bot has been crawling all these URLs for weeks. Thus wasting a shitload of our resources AND a shitload of their resources too. Whoever it is (and I now suspect it is Perplexity based on this Cl…

We have the same issue (billions of URLs). The newer bots that rotate IPs across thousands of IP ranges are killing us and there is no good way to block them short of captcha's or forcing logins, which we would really rather not inflict on our users.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#769

Earlier quoted context omitted.

You’re describing socialism (wealth redistribution to be exact). At this point, just make that money a tax and give it to the publishers directly. Cut out the middlemen.

Well, what is the difference, the ad budget fraction of the price is like a tax. I think given a choice most people would prefer to get their stuff a bit cheaper and not contribute to the ad budget. But we pay it and then the companies hand the money out to various parties to display ads creating the possibility of running a business on ad revenue. And in many cases I can ignore ads, I can not look at billboards, I c…

[deleted]
Post reply on HN