Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

751–760 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#751

Earlier quoted context omitted.

Well, I don't give a shit about the advertising goals of Apple or anyone else, that is why I block ads. And that is also completely irrelevant, the question was whether I am screwing over websites when I am using an ad blocker. I argue not, because as a consumer I still contribute to the ad budgets that become the ad revenue of the websites. What I am not doing when I block ads is influencing how the money gets distr…

LOL, you don’t. You really don’t. As I told you like four hours ago, ads are impression-based. Just because you bought something that helped them buy an ad doesn’t mean you did shit for my website. In fact, you did the opposite.

I know that ads are based on impressions as I told you before, but my money still has to end up somewhere even if I am using an ad blocker. So where does it end up if not as ad revenue on some websites? You must not confuse the people paying for the ads and in turn for the ad revenue of websites by buying stuff with the people deciding how that money gets distributed among all the websites by looking at ads.

We can even go one step further, if anyone is screwing over websites, then that is the ad industry by not paying for blocked ads. I buy an iPhone and Apple takes some additional money from me to spend on advertising. I did not ask for that but I am fine with it. Now I expect Apple to spend the money they took from me on ads in order to support websites. But if the guy that Apple wants to show the ad that I paid for does not want to see it and blocks it, then I want Apple to respect that and still pay the website. I know, not going to happen, but do not put the blame on people blocking ads.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#752

Earlier quoted context omitted.

Of course some people want that. And at the moment they can prevent it. But those methods may stop working. Will it then be alright to do it? Of course not, so why bother mentioning that they are able to prevent it now - just give a justification. Your license is probably not relevant. I can go to the cinema and watch a movie, then come on this website and describe the whole plot. That isn't copyright infringement. E…

> I can go to the cinema and watch a movie, then come on this website and describe the whole plot. That isn't copyright infringement. This is a false analogy. A correct one would be going to a 1000 movies and creating the 1001th movie with scenes cropped from these 1000 movies and assemble it as a new movie, and this is copyright infringement. I don't think any of the studios would applaud and support you for your cr…

> This is a false analogy.

I think you are describing something much more like stable diffusion. This article is about Perplexity, which is much closer to "watch a movie and tell me the plot" than it is like "take these 1000 movies and make a collage". The copyright points are different - stable diffusion are on much shakier ground than perplexity.

> Why does it have to be always about money?

Before I mentioned money I said "because it hurts my feelings". I'm sorry I can't give a more charitable interpretation, but I really do see this kind of objection as "I don't want you to have access to this web page because I don't like LLMs". This is not a principled objection, it is just "I don't like you, go away". I don't think this is a good principle to build the web on.

Obviously you can make your website private, if you want, and that would be a shame. But you can't have this kind of pick-and-choose "public when you feel like" option. By the way I did not mention, but I am ok with people using Anubis and the like as a compromise while the situation remains unjust. But the justification is very important.

> If companies can scrape my pages to sell my content as theirs, I can scrape theirs and unpaywall them.

This is probably not a gambit you want to make. You literally can do this, and they would probably like it if you did. You don't want to do that, because the output of LLMs is usually not that good.

In fact, LLM companies should probably be taxed, and the taxes used to fund real human AI-free creations. This will probably not happen, but I am used to disappointment.

> P.S.: Oh, try to claim that you can train a model with medical data

Medical data is not public, for good reasons.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#753

Earlier quoted context omitted.

Ok, that wasn't clear before since you just kept saying how you expressed your consent rather than why your consent should be taken into account. Licensing is much much more limited than you seem to be thinking of it. For instance, you said explicitly you want a way to control your ideas. The only thing this can mean is a way to control who gets to use your ideas, or what they get to use them for. So if I express a p…

LLMs are not "someone", LLMs are something, and they don't "read content", they by definition acquire and reuse that content (for example, by summarizing it), as part of their product. So here the consent is indeed about what can be done with the data. In general, it's absolutely the norm that public websites (I.e., unauthenticated) restrict even who can access the data. The simplest example that comes to mind is geo…

LLMs are indeed not "someone". They are programs, like web browsers, acting on user instruction. The user is a person. I am only talking about people - I never said that an LLM does anything of its own volition.

> The simplest example that comes to mind is geoblocking.

Do you think it is alright to geoblock people, for arbitrary reasons? It is one thing when GDPR imposes a legal obligation on you for serving content in a particular way. Note that that actually doesn't prevent you from seeing the content, it just prevents you from being served by that server. The distinction is important - circumventing a geoblock is something I think should be legally protected.

> They are not humans, they are not consumers, they don't simply fetch the content and present it to the users

They simply fetch the content, run it through a software, and present it to the user. As far as you, the service owner, are concerned, they are simply fetching the content for the user. It is none of your business what the user and the AI company go on to do with "your content".

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#754

Earlier quoted context omitted.

LOL, you don’t. You really don’t. As I told you like four hours ago, ads are impression-based. Just because you bought something that helped them buy an ad doesn’t mean you did shit for my website. In fact, you did the opposite.

I know that ads are based on impressions as I told you before, but my money still has to end up somewhere even if I am using an ad blocker. So where does it end up if not as ad revenue on some websites? You must not confuse the people paying for the ads and in turn for the ad revenue of websites by buying stuff with the people deciding how that money gets distributed among all the websites by looking at ads. We can e…

You’re describing socialism (wealth redistribution to be exact). At this point, just make that money a tax and give it to the publishers directly. Cut out the middlemen.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#755

This is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the L…

Hi, website operator here. I don't want my content to be accessible to you through Perplexity. I want my work to be freely available to any person who wants it. Feel free to transform my material as you see fit. Hell, do it with LLMs! I don't care. The LLM isn't the problem, it's what companies like Perplexity are doing with the LLM. Do not create commercial products that regurgitate my work as if it was your own. It…

no one cares about your shitty blog lol

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#757
post #705

Earlier quoted context omitted.

> And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads. I think this is a pretty different scenario. Here the user and the news website are talking directly to each other, but then the user is making a choice around what to do with the content the news website send to them. With AI agents, there is a company inserting themselves between th…

I understand; but as an excercise to better understand this problem I'll keep doing devil's advocate and I'll raise with: What if my executive assistant reading the news website and giving me a digest? Would the website owners rather prefer me doing my reading directly?

Yes. Because they want to own your attention and that only works if they are interfacing directly to you.

I remember that Samsung was at one time offering to play non-skippable full-screen apps on their newest 8K OLED TVs and their argument was precisely that these ads will reach those rich people who normally pay extra to avoid getting spammed with ads. Or going with your executive assistant example, there are situations where it makes sense to bribe them to get access to you and/or your data. E.g. "evil maid attack".

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#758

Earlier quoted context omitted.

> It works to get real humans like myself to stop visiting your site If we talk about Anubis, it's pretty invisible. You wait a couple of seconds in the first visit, and don't get challenged for a couple of weeks, at least. With more tuning some of the sites using Anubis work perfectly well without ever seeing Anubis' wall while stopping AI crawlers. > And to be clear, what you are advocating for is DRM. Yes. It's pr…

> If we talk about Anubis, it's pretty invisible. You wait a couple of seconds in the first visit, and don't get challenged for a couple of weeks, at least. With more tuning some of the sites using Anubis work perfectly well without ever seeing Anubis' wall while stopping AI crawlers. It's not invisible, the sites using it don't work perfectly well for all users and it doesn't stop AI crawlers.

I've never seen problems with Anubis.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#759

Earlier quoted context omitted.

Maybe we should just institutionalize and explicitly legalize the Internet Archive and Archive Team. Then, I can download a complete and halfway current crawl of domain X from the IA and that way, no additional costs are incurred for domain X. But of course, most website publishers would hate that. Because they don't want people to access their content, they want people to look at the ads that pay them. That's why to…

https://commoncrawl.org/ >Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.

The problem is that many websites and domains are missing from it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#760
post #585

Earlier quoted context omitted.

With all the crypto development how come we haven't got to HTTP/1.1 402 Payment Required WWW-price: 0.0000001 BTC, 0.000001 ETH, 0.00001 DOGE > You are less likely to participate in discussion you (or AI on your behalf) paid instead. Many sites would probably like it better.

If people were forced to pay for websites by the http request people would demand that websites stop loading a ton of externally hosted JS, stop filling sites with ads, and would demand that websites actually have content worth the price. There are so many links I click on these days that are such trash I'd be demanding refunds constantly.

>There are so many links I click on these days that are such trash

That is why AI "summarization" becomes a necessary intermediate layer. You'd not see nor trash nor ads, and thus the payment instead of being exposed to the ads. AI saves the Internet :)

Post reply on HN