Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

701–710 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#701
post #118

Earlier quoted context omitted.

I think it’s basically impossible to prevent AI crawlers. It is like video game cheating, at the extreme they could literally point a camera at the screen and have it do image processing, and talk to the computer through the USB port emulating, a mouse and keyboard outside the machine. They don’t do that, of course, because it is much easier to do it all in software, but that is the ultimate circumvention of any atte…

I don’t subscribe to technological inevitabilism. Cloudflare banning bad actors has at least made scraping more expensive, and changes the economics of it - more sophisticated deception is necessarily more expensive. If the cost is high enough to force entry, scrapers might be willing to pay for access. But I can imagine more extreme measures. e.g. old web of trust style request signing[0]. I don’t see any easy way f…

A web of trust is not going to plug the analog hole gp already mentioned.

Meanwhile its going to fuck over real users.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#702

Earlier quoted context omitted.

> Cloudflare banning bad actors has at least made scraping more expensive, and changes the economics of it - more sophisticated deception is necessarily more expensive. If the cost is high enough to force entry, scrapers might be willing to pay for access. I think this might actually point at the end state. Scraping bots will eventually get good enough to emulate a person well enough to be indistinguishable (are we t…

Netflix CAN "stop you from pointing a camera at your TV and distributing it" because of copyright law.

Which is also how AI scrapers should be solved. Papering over the issue with technological "solutions" only hurts real users.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#703
post #37

Earlier quoted context omitted.

Ads are a problematic business model, and I think your point there is kind of interesting. But AI companies disintermediating content creators from their users is NOT the web I want to replace it with. Let’s imagine you have a content creator that runs a paid newsletter. They put in lots of effort to make well-researched and compelling content. They give some of it away to entice interested parties to their site, whe…

> Otherwise there is literally no reason for them to make any of it available on the open web This is the hypothesis I always personally find fascinating in light of the army of semi-anonymous Wikipedia volunteers continuously gathering and curating information without pay. If it became functionally impossible to upsell a little information for more paid information, I'm sure some people would stop creating informati…

> Do people (generally) put things online to get money or because they want it online?

IME it's mostly because someone else put something "wrong" online first.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#704
post #223

Earlier quoted context omitted.

> Otherwise there is literally no reason for them to make any of it available on the open web This is the hypothesis I always personally find fascinating in light of the army of semi-anonymous Wikipedia volunteers continuously gathering and curating information without pay. If it became functionally impossible to upsell a little information for more paid information, I'm sure some people would stop creating informati…

Any information that requires something approximating a full-time job worth of effort to produce will necessarily go away, barring the small number of independently wealthy creators. Existing subject-matter experts who blog for fun may or may not stick around, depending on what part of it is “fun” for them. While some must derive satisfaction from increasing the total sum of human knowledge, others are probably blogg…

> Any information that requires something approximating a full-time job worth of effort to produce will necessarily go away

Many people put more effort into their hobbies than into their "full time" job.

Some of it will go away but perhaps without the expectation that you can earn money more people will share freely.

> While some must derive satisfaction from increasing the total sum of human knowledge, others are probably blogging to engage with readers or build their own personal brand, neither of which is served by AI scrapers.

We don't have to make all business models that someone might want possible though.

> Wikipedia is an interesting case. I still don’t entirely understand why it works, though I think it’s telling that 24 years later no one has replicated their success.

Actually this model is quite common. There are tons of sources of free information curated by volunteers - most are just too niece to get to the scale of Wikipedia.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#705

Earlier quoted context omitted.

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

Foo news wants you to visit the site, look at the main page, watch the ads, click on them and buy the products advertised by third parties which will give money to Foo news in exchange for this service. And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads. They claim that since they are free to not buy an advertised product, why would they…

> And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads.

I think this is a pretty different scenario. Here the user and the news website are talking directly to each other, but then the user is making a choice around what to do with the content the news website send to them. With AI agents, there is a company inserting themselves between the user and the news website and acting as a middleman.

It seems reasonable to me that the news website might say they only want to deal with users and not middlemen.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#706
post #37

Earlier quoted context omitted.

Ads are a problematic business model, and I think your point there is kind of interesting. But AI companies disintermediating content creators from their users is NOT the web I want to replace it with. Let’s imagine you have a content creator that runs a paid newsletter. They put in lots of effort to make well-researched and compelling content. They give some of it away to entice interested parties to their site, whe…

Maybe, on a social level, we all win by letting AI ruin the attention economy: The internet is filled with spam. But if you talk to one specific human, your chance of getting a useful answer rises massively. So in a way, a flood of written AI slop is making direct human connections more valuable. Instead of having 1000+ anonymous subscribers for your newsletter, you'll have a few weekly calls with 5 friends each.

I'm not sure what experiences you are basing this optimism on but I'm happy for you.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#707
so cloudflare blocked the agent from accessing the site. then when it couldn't access the robots.txt because it was blocked they punished it for using intelligent work around to access a website with no known history. perplexity is running a browser that follows the instruction of the user. if the user could manually do it then the agent is simply a tool to do the manual thing. this is a battle about websites and advertisers pissed that their analytics show and impressions... let's not pretend cloudflare is protecting anyone

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#708

Earlier quoted context omitted.

I have never created a website that I would not mind being fully crawled and indexed into another dataset that was divorced from the source (other than such divorcement makes it much harder to check pedigree, which is an academic concern, not a data-content concern: if people want to trust information from sources they can't know and they can't verify I can't fix that for them). In fact, the "old web" people sometime…

There's an important distinction that we are glossing over I think. In the times of the "old web", people were putting things online to interact with a (large) online audience. If people found your content interesting, they'd keep coming back and some of them would email you, there'd be discussions on forums, IRC chatrooms, mailing lists, etc. Communities were built around interesting topics, and websites that starte…

I agree but that cancer isn't limited to the internet or even originated from it. An until society as a whole is ready to deal with it the only thing we can do is form our own subculture that rejects this new normal. Instead of caring about what gets scraped or otherwise used by mega corporations for profit, care about finding more exchanges with real humans. Or in other words: be part creating the world you want to see and ignore those that choose not to participate.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#709
post #639

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Note that a book author cannot publish a book and then refuse to let libraries buy copies and lend them out. This was litigated 100+ years ago.

The difference is that libraries aren't all about concentrating wealth for themselves.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#710
post #705

Earlier quoted context omitted.

Foo news wants you to visit the site, look at the main page, watch the ads, click on them and buy the products advertised by third parties which will give money to Foo news in exchange for this service. And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads. They claim that since they are free to not buy an advertised product, why would they…

> And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads. I think this is a pretty different scenario. Here the user and the news website are talking directly to each other, but then the user is making a choice around what to do with the content the news website send to them. With AI agents, there is a company inserting themselves between th…

I understand; but as an excercise to better understand this problem I'll keep doing devil's advocate and I'll raise with:

What if my executive assistant reading the news website and giving me a digest?

Would the website owners rather prefer me doing my reading directly?

Post reply on HN