Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

571–580 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#571
post #479

Earlier quoted context omitted.

[flagged]

[flagged]

this is such a wild comment -- there are countless products where regardless of purchase -- the user is still served advertisements. i have no idea what reality, or timeline, this comment belongs in.

broadcast television, paid streaming entertainment is just straight up the most glaringly obvious example of a paid service overflowing with advertisements.

paid radio broadcasts (xm/Sirius).

operating systems (windows serves you ads any chance it gets).

monthly subscriptions to gyms where youre constantly hit with ads, marketing, and promotions be it at the gym or via push notification (you got opted into and therefore have to opt out of intentionally after the service is paid).

mobile phones, especially prepaid come LOADED with ads and bloatware.

i mean the list goes on -- you cannot be serious.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#572

Earlier quoted context omitted.

Sure. There's lots of things you could do, but you don't do them because they are wrong. Might does not make right.

How is it wrong to send my personal shopper? How is it wrong to have an agent act directly on my behalf? It's like saying a web browser that is customized in any way is wrong. If one configures their browser to eagerly load links so that their next click is instant, is that now wrong?

if you send your personal shopper to a store, and the business is... closed for business, or refusing you entry, and you just... go in anyway.

that's called breaking and entering, and generally frowned upon -- by-passing the "closed sign".

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#573
post #554

Earlier quoted context omitted.

As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?

> What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? What prevents anyone else? robots.txt is a request, not an access policy.

This honor system mostly worked at scale because interests align, which seems to be no longer the case.

Does information no longer wants to be free now? Maybe internet, just like social media was just a social experiment at the end, albeit a successful one. Thanks GenAI.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#574

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

It's a tough issue indeed. One thing that comes to my mind is: If a human tries to answer a question via the web, he will browse one site after the other. If that human asks an LLM, it will ping 25 sites in parallel. Scale this up to all of humanity, and it should be expected that internet traffic will rise 25x - just from humans manually asking questions every now and then - we are not even talking about AI companie…

If a human tries to answer a question via the web, he will browse one site after the other.

Not me, I often open multiple tabs and windows at once to compare and contrast the results.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#575

Earlier quoted context omitted.

Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…

As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?

It's your server. You're free to do whatever you want. You can serve different versions of the page depending on the UserAgent (has been done many times before).

You can put up a paywall depending on UserAgent or OS (has been done).

In short, it's a 2-way street: the client on the other end of the TCP pipe makes a request, and your server fulfills the request as it sees fit.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#576

Earlier quoted context omitted.

I have mixed feelings on this. Many websites (especially the bigger ones) are just businesses. They pay people to produce content, hopefully make enough ad revenue to make a profit, and repeat. Anything that reproduces their content and steals their views has a direct effect on their income and their ability to stay in business. Maybe IA should have a way for websites to register to collect payment for lost views or…

If magazines and newspapers were once able to be funded by native ads, so can websites. The spying industry doesn't want you to know this, but ads work without spying too - just look at all the IRL billboards still around.

I never said anything about spying.

Magazines and newspapers were able to by funded by native ads because you couldn't auto-remove ads from their printed media and nobody could clone their content and give it away for free.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#577

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

To stretch the analogy to the breaking point: If you send 10,000 personal shoppers all at once to the same store just to check prices, the store's going to be rightfully annoyed that they aren't making sales because legit buyers can't get in.

That is not the breaking point at all of the analogy-- that literally happens to my custom CMS/wiki/image host I built for my niche, kpopping.com. We are constantly attacked by crawlers. Meanwhile google rewards wordpress slop that buys backlinks with #1 pageranks for years. Welcome to the internet.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#578

Earlier quoted context omitted.

I have mixed feelings on this. Many websites (especially the bigger ones) are just businesses. They pay people to produce content, hopefully make enough ad revenue to make a profit, and repeat. Anything that reproduces their content and steals their views has a direct effect on their income and their ability to stay in business. Maybe IA should have a way for websites to register to collect payment for lost views or…

If magazines and newspapers were once able to be funded by native ads, so can websites. The spying industry doesn't want you to know this, but ads work without spying too - just look at all the IRL billboards still around.

Thanks for pointing this out! This is too often ignored!

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#579
post #479

Earlier quoted context omitted.

It’s possible to violate all sorts of social norms. Societies that celebrate people that do so are on the far opposite end of the spectrum from high trust ones. They are rather unpleasant.

[flagged]

[flagged]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#580

Earlier quoted context omitted.

The HTTP spec draws such a distinction, albeit implicitly, in the form (and name) of its concept of "user agent."

Over time it degraded into declaring compatibility with a bunch of different browser engines and doesn't reflect the actual agent anymore. And very likely Perplexity is in fact using a Chrome-compatible engine to render the page.

user agent = which bullshit css hacks and js polyfills will be needed
Post reply on HN