Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

691–700 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#691

Earlier quoted context omitted.

> as LLM adoption increases you'll merely find that your site has fewer and fewer visits overall, so your content will only be utilized by you and a vanishingly small group of other persons. So far, AI has had the opposite effect on my site. I've now been featured on both Hackaday and Adafruit's blog. Both features were clearly AI-generated. Both posts coincided with an influx of emails from folks interested in my wo…

> So far, AI has had the opposite effect on my site. I've now been featured on both Hackaday and Adafruit's blog. Both features were clearly AI-generated. Both posts coincided with an influx of emails from folks interested in my work. This may be missing some context, but it seems as though you're saying that you made something with AI and it led to traction. That's great! Seems off the point that blocking LLM servic…

> This may be missing some context, but it seems as though you're saying that you made something with AI and it led to traction. That's great! Seems off the point that blocking LLM service will lead to less exposure over time though.

Hah, I can see how you would have read it that way. Quite the opposite. I don't use AI tools for my writing. Hackaday and Adafruit have both featured my posts, and their posts were pretty clearly AI-generated.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#692

Earlier quoted context omitted.

Just the Silicon Valley ethos extended to it's logical conclusions. These companies take advantage of public space, utilities and goodwill at industrial scale to "move fast and break things" and then everyone else has to deal with the ensuing consequences. Like how cities are awash in those fucking electric scooters now. Mind you I'm not saying electric scooters are a bad idea, I have one and I quite enjoy it. I'm sa…

My city impounded them and made them pay a fee to get them back. Now they have to pay a fee every year to be able to operate. Win/win.

Do those fees actually improve anything for the citizens who now have to deal with vehicles abandoned on sidewalks everywhere or does it just buy the major a nicer yacht?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#694

Earlier quoted context omitted.

That's a very sad and lonely way to live.

I don't think we're talking about the same thing.

Obviously. You should heed the advice of other posters who told you to look up the meaning of the word.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#695

Why single out Perplexity? Pretty much no crawler out there fetches robots.txt. robots.txt is not a blocking mechanism; it's a hint to indicate which parts of a site might be of interest to indexing. People started using robots.txt to lie and declare things like no part of their site is interesting, and so of course that gets ignored.

This is objectively wrong. Take it straight from the source: https://www.rfc-editor.org/rfc/rfc9309.html

[dead]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#696

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

But the store owner can ask the personal shopper to leave, if e.g. they find out that they work for a personal shopper service.

What the article is advocating for is hiring bouncers that strip all shoppers so they can do just that.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#698
post #6

If you put info on the web, it should be available to everyone or everything with access.

I agree, but that does not mean that you should use excessive requests and unnecessary scraping and overloading the servers to access them. The files should be mirrored. Some may be better copied in other ways, e.g. a git repository can be cloned and mirrored in that way, and should not need to crawl the web pages to do so.

[dead]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#699
Adhering to robots.txt is merely a courtesy.

Much like a trolley drop off at your local shopping center car park. Some users will adhere to it and drop their trolleys in after their done. Others will not and will leave it wherever.

Your machine might access a page via a browser that is human readable. My machine might read it via software and present the content to me in some other form of my choosing. Neither is wrong. Just different.

Don't like it? Then don't post your website on the internet...

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#700
post #37

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Ads are a problematic business model, and I think your point there is kind of interesting. But AI companies disintermediating content creators from their users is NOT the web I want to replace it with. Let’s imagine you have a content creator that runs a paid newsletter. They put in lots of effort to make well-researched and compelling content. They give some of it away to entice interested parties to their site, whe…

A agree with your first line but the rest sounds like a similar argument to the ridiculous damages video game companies used to claim due to piracy when most of those pirates never would have bought the game in the first place.

Ultimately the root issue is that copyright is inherently flawed because it tries to increase available useful information by restricting availability. We'd be better off by not pretending that information is scarce and looking for alternative to fund its creation.

Post reply on HN