Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

731–740 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#731

Earlier quoted context omitted.

If I buy an iPhone, does some fraction of the price contribute to Apple's ad budget? If so, where does that money end up? What would change if I did not block Apple ads?

It’s up to them how they spend their money, not you. You can complain if they somehow damaged your product, they got your money unfairly, or were somehow doing something bad with your data, but at some point it is their money to spend how they see fit. They earned it, and they might spend it on advertising. If I buy stuff at a grocery store, I can’t get a random bagger fired just because I feel like it. At some point…

I am neither complaining nor trying them what to do with their money, that looks like a complete deflection to me.

If I am buying Apple products, am I contributing to their ad budget? If so, where does that money end up? Is it likely that some of it will end up as ad revenue on some website? What difference does it make whether or not I block ads? Or the other way around, if I am visiting websites and look at Apple ads but do not buy Apple products, am I contributing to the ad revenue of the websites?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#732

Earlier quoted context omitted.

> So far, AI has had the opposite effect on my site. I've now been featured on both Hackaday and Adafruit's blog. Both features were clearly AI-generated. Both posts coincided with an influx of emails from folks interested in my work. This may be missing some context, but it seems as though you're saying that you made something with AI and it led to traction. That's great! Seems off the point that blocking LLM servic…

> This may be missing some context, but it seems as though you're saying that you made something with AI and it led to traction. That's great! Seems off the point that blocking LLM service will lead to less exposure over time though. Hah, I can see how you would have read it that way. Quite the opposite. I don't use AI tools for my writing. Hackaday and Adafruit have both featured my posts, and their posts were prett…

That still sounds like a great deal. Less work for those post authors, and you benefitting from being cited in some way (maybe they used Perplexity or similar, and didn't even visit your site themselves).

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#733

Earlier quoted context omitted.

Netflix CAN "stop you from pointing a camera at your TV and distributing it" because of copyright law.

Which is also how AI scrapers should be solved. Papering over the issue with technological "solutions" only hurts real users.

in the UK at least, that has been recognized as fair use.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#734

Earlier quoted context omitted.

Because lawyers are expensive and big tech companies have lots of them. Because it takes a ton of time and effort to sue someone. Because you need to show standing, which means you need to be able to demonstrate you lost something of value by their actions. Because the power imbalance is heavily weighted towards a corporation. Because the way to deal with such things should be legislation and not court decisions. And…

That's exactly why I said conciliation court. None of what you've outlined is required nor is it expensive. But, for each case, the defendant is still required to show up. I've successfully used conciliation court against large corporations in the past which is why I question it here. And while this should be able to be handled via legislation it won't be. Beyond that a workaround could force that to happen.

> conciliation court

Sorry, I had never heard that term before. You would still have to show standing though. How would you try to prove that their violating your TOS cost you money?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#735

Earlier quoted context omitted.

It’s up to them how they spend their money, not you. You can complain if they somehow damaged your product, they got your money unfairly, or were somehow doing something bad with your data, but at some point it is their money to spend how they see fit. They earned it, and they might spend it on advertising. If I buy stuff at a grocery store, I can’t get a random bagger fired just because I feel like it. At some point…

I am neither complaining nor trying them what to do with their money, that looks like a complete deflection to me. If I am buying Apple products, am I contributing to their ad budget? If so, where does that money end up? Is it likely that some of it will end up as ad revenue on some website? What difference does it make whether or not I block ads? Or the other way around, if I am visiting websites and look at Apple a…

Maybe in the cosmic sense you are, in that they have a giant pile of money, and you contributed a few pennies to it, but this is not how accounting works. Your transaction and their ad budget are separate things.

Also, advertising does other things than tell you to buy something, and it doesn’t always take the form of banner ads. Apple, for example, does a ton of brand awareness advertising. Affiliate marketing often targets direct transactions. Maybe your goal is to simply start a relationship that might someday lead to a really big purchase.

Often, in the era of SaaS, people advertise to existing customers. Apple does this—they have a TV service and a music service and a cloud service.

There are plenty of reasons for them to advertise after you bought the original product.

But your original point was that customers bought the ads. Maybe they didn’t! Maybe they were given funding by a VC firm and the company decided it wanted to build an audience. Maybe they want to advocate for a political issue.

I think the biggest problem with your argument is that it has tunnel vision and sees advertising as this one dimensional thing, when in reality it takes many forms. Plenty of those forms are bad, but it is not as simple as “I bought a product, now I never want to see an Apple ad ever again.” Many businesses (Amazon, eBay) make most of their money off of customers they’ve already advertised to that they advertise to again and again.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#736
post #335

This is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the L…

> The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If you are the source I think they could make plenty of sense. As an example, I run a website where I've spent a lot of time documenting the history of a somewhat niche activity. Much of this information isn't available online anywhere else. As it happens I'm happy to let b…

Perhaps the better way forward here (for all) is for some kind of central content archive that bots can then pull from (Internet Archive?). But then there'll be questions of how up-to-date the archive is compared to the source, and if the source is showing the same thing that it's allowing to be archived.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#737
post #679

Earlier quoted context omitted.

Because scrapers would certainly comply with that /s

More like have easier to assess legality status.

How so? The legal status without a license is already "All Rights Reserved".

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#738

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

> Some stores do not welcome Instacart or Postmates shoppers

First time hearing this. Almost every single grocery store either supports Instacart, or has partnership with a similar service.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#739

Earlier quoted context omitted.

Foo news wants you to visit the site, look at the main page, watch the ads, click on them and buy the products advertised by third parties which will give money to Foo news in exchange for this service. And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads. They claim that since they are free to not buy an advertised product, why would they…

It's not ads. We have ads in paper magazines and newspapers and no one went around with scissors to remove them. It's obnoxious ads, designed to violently grabs your attention and trackers (malware). It's like a newspapers giving your address to a whole crew of salemens that intrudes on your property at 3am and looking at you sleeping and installing cameras in your bathroom. All so that they can jump at you in the st…

> We have ads in paper magazines and newspapers and no one went around with scissors to remove them.

I never saw people bother with scissors but I've seen people pulling the ads out of the newspaper countless times.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#740
post #3

>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…

Sounds like an ad for Perplexity. They do end up looking bad out of Cloudflare's report, who are the "good guys" in this story - btw Cloudflare's been very pushy lately with their we'll save the web, content independence day marketspeak. But deep in the back of my head, Cloudflare's goodwill elevates Perplexity cunning habilities (assuming they're the culprit since no real evidence, only heresay is in the OP), both c…

Sounds like ad for cloudflare. Didn’t they announce a month ago they will protect websites from llm content sweep? And now they realize they cannot deliver on that promise. We did it correctly but these guys are doing it illegal way! That’ll be 14.99 per month btw..
Post reply on HN