Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

741–750 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#741
post #585

Earlier quoted context omitted.

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

With all the crypto development how come we haven't got to HTTP/1.1 402 Payment Required WWW-price: 0.0000001 BTC, 0.000001 ETH, 0.00001 DOGE > You are less likely to participate in discussion you (or AI on your behalf) paid instead. Many sites would probably like it better.

If people were forced to pay for websites by the http request people would demand that websites stop loading a ton of externally hosted JS, stop filling sites with ads, and would demand that websites actually have content worth the price.

There are so many links I click on these days that are such trash I'd be demanding refunds constantly.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#742

Earlier quoted context omitted.

Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…

As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?

Nothing, and that's why I expect they all do it.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#743

Earlier quoted context omitted.

Thanks for sharing your experience. A little off-topic but I'd like to start hosting some personal content, guides/tutorials, etc. Do you still see authentic human traffic on your domains, is it easy to discern? I feel like I missed the bus on running a blog pre-AI.

I intentionally doesn't keep detailed analytics on my homepage server and my digital garden, because I respect my users and don't want to push unnecessary Javascript on them. The blog platform I use (Mataroa) keeps rudimentary analytics (essentially page hit counters, nothing more) on index, RSS and per post. Both my blog homepage and posts see mostly human traffic. Sometimes bots crawl the site and they appear as sp…

> I intentionally doesn't keep detailed analytics on my homepage server and my digital garden, because I respect my users and don't want to push unnecessary Javascript on them.

Absolutely, I'm in agreement here. I want to run a JS-free blog, just plain old static HTML. I plan to use GoAccess to parse the access logs but that's it. I think I would find it encouraging to see real human traffic.

> I don't write these for others, first. Both my blog (more meta) and digital garden (more technical) are written for myself primarily, and left open.

That is a great way to view it, thank you.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#744

Earlier quoted context omitted.

The problem in your logic is that all points starts wit "I". You're not the only stakeholder in any of those interactions. There's you, a mediator (search or LLM), and the website owner. The website owner (or its users) basically do all the work and provide all the value. They produce the content and carry the costs and risks. The pre-LLM "deal" was that at least some traffic was sent their way, which helps with reac…

But the entire reason that the web is so frustrating is that visitors don't want to pay for anything. They are already paying, it is the way they are paying that causes the mess. When you buy a product, some fraction of the price is the ad budget that gets then distributed to websites showing ads. Therefore there is also nothing wrong with blocking ads, they have already been paid for, whether you look at them or not…

I feel like this could work if the payment was handled by your ISP. Content provider tells the ISP how much their content costs that there subscribers pay, and the ISP pays them. I already pay my ISP. The real problem is that it's kinda too late for this kind of change. And also the ISP would need to prevent their users from running up a bill that the ISP would be responsible for and without tracking them that's not possible.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#745

Earlier quoted context omitted.

I am neither complaining nor trying them what to do with their money, that looks like a complete deflection to me. If I am buying Apple products, am I contributing to their ad budget? If so, where does that money end up? Is it likely that some of it will end up as ad revenue on some website? What difference does it make whether or not I block ads? Or the other way around, if I am visiting websites and look at Apple a…

Maybe in the cosmic sense you are, in that they have a giant pile of money, and you contributed a few pennies to it, but this is not how accounting works. Your transaction and their ad budget are separate things. Also, advertising does other things than tell you to buy something, and it doesn’t always take the form of banner ads. Apple, for example, does a ton of brand awareness advertising. Affiliate marketing often…

Well, I don't give a shit about the advertising goals of Apple or anyone else, that is why I block ads. And that is also completely irrelevant, the question was whether I am screwing over websites when I am using an ad blocker. I argue not, because as a consumer I still contribute to the ad budgets that become the ad revenue of the websites. What I am not doing when I block ads is influencing how the money gets distributed among all the websites, I can live with that. And if the money is not consumer money, so what? What do I have to do with companies distributing VC money among websites?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#746

Earlier quoted context omitted.

Maybe in the cosmic sense you are, in that they have a giant pile of money, and you contributed a few pennies to it, but this is not how accounting works. Your transaction and their ad budget are separate things. Also, advertising does other things than tell you to buy something, and it doesn’t always take the form of banner ads. Apple, for example, does a ton of brand awareness advertising. Affiliate marketing often…

Well, I don't give a shit about the advertising goals of Apple or anyone else, that is why I block ads. And that is also completely irrelevant, the question was whether I am screwing over websites when I am using an ad blocker. I argue not, because as a consumer I still contribute to the ad budgets that become the ad revenue of the websites. What I am not doing when I block ads is influencing how the money gets distr…

LOL, you don’t. You really don’t. As I told you like four hours ago, ads are impression-based. Just because you bought something that helped them buy an ad doesn’t mean you did shit for my website.

In fact, you did the opposite.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#747

Earlier quoted context omitted.

this is such a wild comment -- there are countless products where regardless of purchase -- the user is still served advertisements. i have no idea what reality, or timeline, this comment belongs in. broadcast television, paid streaming entertainment is just straight up the most glaringly obvious example of a paid service overflowing with advertisements. paid radio broadcasts (xm/Sirius). operating systems (windows s…

> pay for services in full directly Those are hybrid subscriptions/subsidies. Not paid in full. If you are being exposed to ads in something you paid for, you are almost certainly being charged less money. Companies can compete on cost by introducing ads, and it's why the cheaper you go, the more ad infested it gets. Pure ad-free things tend to be much more expensive then their ad subsidized counterparts. Ad subsidiz…

this seems like semantics and corporate hand-waving -- that's not what is conveyed to the user in what i have observed as the context of paid services and the promises asserted around what a purchase gets a customer.

in the subsidized example, xm/Sirius is marketed to users as an "ad-free paid radio broadcast"; the marketing literally attempts to leverage the notion of it being ad-free as a consequence of your purchase (power) in order to highlight its supposed competitive edge and usefulness, and provide the user an incentive to spend money, except for the fact that the marketing is false. you still get served promotions and ads, just less "conventional" ads.

i go to a football game and im literally inundated with ads -- the whole game has time stoppage dedicated to serving ads. i guess my season ticket purchase with the hopes of seeing football in person is.. apparently not spending enough money?

i see this as attempting to move the goalposts and gaslight users on their purchase expectations, as a way to offload the responsibility and accountability back onto the user -- "you don't pay enough, you only think that you pay enough, so we are still going to serve you ads because .

why then is there any expectation of a service being ad-free upon purchasing?

who the hell actually enjoys sitting through 1.5 hours of advertisements and play stoppage?

over time users have been conditioned to just tolerate it, and over time, the advertising reclaims ground it previously gave up one inch at a time in the same way people are price-gouged in those stadiums -- they don't have much alternative, but apparently the problem is the user should fork up more money for tickets so as to align their expectations with reality? while they're getting strong-armed at the concession stand via proximity and circumstance and lack of competition, no less.

are you really trying to tell me the problem there is, they need to make... more money? and THEN and only THEN we can have ad-free, paid for entertainment otherwise known as american football? is this really about user expectations, or is this about companies wanting their cake and eating it, too?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#748

Earlier quoted context omitted.

> So far, AI has had the opposite effect on my site. I've now been featured on both Hackaday and Adafruit's blog. Both features were clearly AI-generated. Both posts coincided with an influx of emails from folks interested in my work. This may be missing some context, but it seems as though you're saying that you made something with AI and it led to traction. That's great! Seems off the point that blocking LLM servic…

> This may be missing some context, but it seems as though you're saying that you made something with AI and it led to traction. That's great! Seems off the point that blocking LLM service will lead to less exposure over time though. Hah, I can see how you would have read it that way. Quite the opposite. I don't use AI tools for my writing. Hackaday and Adafruit have both featured my posts, and their posts were prett…

@ryukoposting - i am the founder of hackaday, but do not run the site now, and i am also the managing director of adafruit and editor of the adafruit blog. the adafruit does not use generative text, etc. unless clearly indicated https://www.adafruit.com/editorialstandards ... appreciate a correction to your post, hard to combat misinformation with ai, but you can email and i can prove i am human if ya want... pt at adafruit dot com

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#749

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

The simple answer to #3 is advertising, including telemetry, tracking and other forms of web-based surveillance. These usually rely on certain browser "features" and/or default settings. The goal is not to make the content usable. The goal is to get the traffic. When advertising alone is the "business model", e.g., not the value of the "content", then even Cloudflare is going to try to protect it (the advertising, no…

(IMHO) The correct way to limit "abuse", e.g., by "bots", is to rate limit. But as other commenters point out, Cloudflare routinely (and knowingly) blocks humans sending only a single GET request, e.g., with Javascript disabled. Needless to say, this does not exceed any reasonable rate limit. It is not "abuse". By Cloudflare's own admission, and as demonstrated by the case of Perplexity AI, this "bot protection" does not stop "bots".

It does stop any humans not using popular advertising-sponsored web browsers.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#750

Earlier quoted context omitted.

I intentionally doesn't keep detailed analytics on my homepage server and my digital garden, because I respect my users and don't want to push unnecessary Javascript on them. The blog platform I use (Mataroa) keeps rudimentary analytics (essentially page hit counters, nothing more) on index, RSS and per post. Both my blog homepage and posts see mostly human traffic. Sometimes bots crawl the site and they appear as sp…

> I intentionally doesn't keep detailed analytics on my homepage server and my digital garden, because I respect my users and don't want to push unnecessary Javascript on them. Absolutely, I'm in agreement here. I want to run a JS-free blog, just plain old static HTML. I plan to use GoAccess to parse the access logs but that's it. I think I would find it encouraging to see real human traffic. > I don't write these fo…

> That is a great way to view it, thank you.

You're welcome. I'm glad it helped.

> I want to run a JS-free blog, just plain old static HTML.

If you want to start fast until you find a template you want to work with, I can recommend Mataroa [0]. The blog have almost no JS (it binds a couple of keys for navigation, that's it), and it's $10/year. When you feel right in your self-hosted solution, you can move there. It's all Markdown at the end of the day.

> I plan to use GoAccess to parse the access logs but that's it.

That's the only thing I use, too. Nothing else.

If you want to look at what I do, how I do, and reach out to me, the rabbit hole starts from my profile, here.

Wish you all the best, and you may find bliss and joy you never dreamed of!

[0]: https://www.mataroa.blog

Post reply on HN