Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

551–560 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#551

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Is a perplexity visit not cached and shared between users who perform a similar search? I don't know much about perplexity, but I'd be surprised if scraped results weren't used to serve multiple searches and users. By bypassing the no-crawl directive, that is a violation of the website's expressed request. I think it is different if individual users chooses to bypass certain things on a website, but for a company to choose to do it is another story.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#553
post #535

Earlier quoted context omitted.

> but I think it's a reasonable stance to not want other companies to profit from my hard work Imagine someone at another company reads your site, and it informs a strategic decision they make at the company to make money around the niche activity you're talking about. And they make lots of money they wouldn't have otherwise. That's totally legal and totally ethical as well. The reality is, if you do hard work and ma…

On a more human level, I think it's bleak that someone who makes a blog just to share stuff for fun is going to have most of his traffic be scrapers that distill, distort, and reheat whatever he's writing before serving it to potential readers.

I don't think it's bleak, just the opposite.

If someone writes valuable stuff on a blog almost nobody finds, that's a tragedy.

If LLM's can process the information and provide it to people in conversations where it will be most helpful, where they never would have found it otherwise, then that's amazing!

If all you're trying to do is help people with the information you've discovered, why do you care if it's delivered via your own site or via LLM? You just want it out there helping people.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#554

Earlier quoted context omitted.

Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…

As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?

> What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle?

What prevents anyone else? robots.txt is a request, not an access policy.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#557

Earlier quoted context omitted.

To stretch the analogy to the breaking point: If you send 10,000 personal shoppers all at once to the same store just to check prices, the store's going to be rightfully annoyed that they aren't making sales because legit buyers can't get in.

Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…

Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion.

When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#558
post #494

Earlier quoted context omitted.

That's somewhat antisocial, but perfectly legal in the US. It's called PayPal Honey, for example, and has been running for 13 years now.

Since when does PayPal Honey replace ads on websites? > PayPal Honey is a browser extension that automatically finds and applies coupon codes at checkout with a single click.

They overwrite ad attributions, affiliate links, clickthrough attributions with their own.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#559
post #3

>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…

Sounds like an ad for Perplexity.

They do end up looking bad out of Cloudflare's report, who are the "good guys" in this story - btw Cloudflare's been very pushy lately with their we'll save the web, content independence day marketspeak. But deep in the back of my head, Cloudflare's goodwill elevates Perplexity cunning habilities (assuming they're the culprit since no real evidence, only heresay is in the OP), both companies look like titans fighting, which ends up being positive for Perplexity, at least in the inflated perception of their firepower... if that makes any sense.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#560
I kind love this fast escalation. Clearly the web can benefit from people to start thinking for locally or narrowly instead of "global audiences". By locally I don't necessarily mean geographically local, just socially local. Build your audience then invite them into private(r) spaces. The (old) open web will be filled with machines built for machines.

We learned to dislike "bubbles" in the past decades but bubbles make sense and are natural, obviously if you're not alone in it.

When it becomes awfully busy with machines and machine content humans will learn to reconnect.

Post reply on HN