Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

451–460 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#451
post #190

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. > If you want to gatekeep your content, use authentication. Are there no limits on what you use the content for? I can start my own search engine that just scrapes Google results?

I think OP based this on an old case about what you can do with data from Facebook vs LinkedIn based on if you need to be logged in to get it. Not relevant when you talk about scraping in this case I think. P is clearly in the wrong here.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#452
post #84

Earlier quoted context omitted.

In theory retrieving a page on behalf of a user would be acceptable, but these are AI companies who have disregarded all norms surrounding copyright, etc. It would be stupid of them not to also save contents of the page and use it for future AI training or further crawling

If you allow Googlebot to crawl your website and train Gemini, but you don't allow smaller AI companies to do the same thing, then you're contributing to Google's hegemony. Given that AI is likely to be an increasingly important part of society in the future, that kind of discrimination is anti-social. I don't want a future where everything is run by Google even more than it currently is. Crawling is legal. Training…

Googlebot respects robots.txt. And Google doesn't use the fetched data from users of Chrome to supplement their search index (as a2128 is speculating that Perplexity might do when they fetch pages on the user's behalf).

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#453
post #299

Earlier quoted context omitted.

Here's how perplexity works: 1) It takes your query, and given the complexity might expand it to several search queries using an LLM. ("rephrasing") 2) It runs queries against a web search index (I think it was using Bing or Brave at first, but they probably have their own by now), and uses an LLM to decide which are the best/most relevant documents. It starts writing a summary while it dives into sources (see next).…

What’s wrong with it downloading documents when the user asks it to? My browser also downloads whole documents and sometimes even prefetches documents I haven’t even clicked on yet. Toss in a adblocker or reader mode and my browser also strips all the ads. Why is it okay for me to ask my browser to do this but I can’t ask my LLM to do the same?

When Google sends people to a review website, 30% of users might have an adblocker, but 70% don't. And even those with adblockers might click an affiliate link if they found the review particularly helpful.

When ChatGPT reads a review website, though? Zero ad clicks, zero affiliate links.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#454
post #107

Earlier quoted context omitted.

Maybe we should just institutionalize and explicitly legalize the Internet Archive and Archive Team. Then, I can download a complete and halfway current crawl of domain X from the IA and that way, no additional costs are incurred for domain X. But of course, most website publishers would hate that. Because they don't want people to access their content, they want people to look at the ads that pay them. That's why to…

Or websites can monetize their data via paid apis and downloadable archives. That's what makes Reddit the most valuable data trove for regular users.

I don't think Reddit pays the people who voluntarily write Reddit content. Valuable to Reddit, I guess.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#455

"Stealth" crawlers are always going to win the game. There are ways to build scrapers using browser automation tools [0,1] that makes detection virtually impossible. You can still captcha, but the person building the automation tools can add human-in-the-loop workflows to process these during normal business hours (i.e., when a call center is staffed). I've seen some raster-level scraping techniques used in game dev…

> "Stealth" crawlers are always going to win the game. no, because we'll end up with remote attestation needed to access any site of value

Yes, because there's always the option for a camera pointed at the screen and a robot arm moving the mouse. AI is hoping to solve much harder problems.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#456

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

Question from a non-web-developer. In case 3, would it be technically possible for Perplexity's website to fetch the URL in question using javascript in the user's browser, and then send it to the server for LLM processing, rather than have the server fetch it? Or do cross-site restrictions prevent javascript from doing that?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#457

Earlier quoted context omitted.

The way to prevent people from downloading your pages and using them is to take them off the public internet. There are laws to prevent people from violating your copyright or from preventing access to your service (by excessive traffic). But there is (thankfully) no magical right that stops people from reading your content and describing it.

Many site operators want people to access their content, but prevent AI companies from scraping their sites for training data. People who think like that made tools like Anubis, and it works. I also want to keep this distinction on the sites I own. I also use licenses to signal that this site is not good to use for AI training, because it's CC BY-NC-SA-2.0. So, I license my content appropriately (No derivative, Non-c…

I guess that's a question that might be answered by the NYT vs OpenAI lawsuit at least on the enforceability of copyright claims if you're a corporation like NYT.

If you don't have the funds to sue an AI corp, I'd probably think of a plan B. Maybe poison the data for unauthenticated users. Or embrace the inevitability. Or see the bright side of getting embedded in models as if you're leaving your mark.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#458
post #75
post #54

Earlier quoted context omitted.

A/ i love this distinction. B/ my brother used to use "fetcher" as a non-swear for "fucker"

Did you tell him to stop trying to make fetcher happen?

Very funny. Now let's hear Paul Allen's joke.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#459
post #139

Earlier quoted context omitted.

It's all about scale. The impact of your personal shopper is insignificant unless you manage to scale it up into a business where everyone has a personal shopper by default.

How is everyone having a personal shopper a problem of scale? I was going to shop myself, but I sent someone else to do it for me. At this moment I am using Perplexity's Comet browser to take a spotify playlist and add all the tracks to my youtube music playlist. I love it.

Let's look at the opposite benefit to a store if a mom that would need to bring her 3 kids to the store vs that mom having a personal shopper. In this case, the personal shopper is "better" for the store as far as physical space. However, I'm sure the store would still rather have the mom and 3 kids physically in the store so that the kids can nag mom into buying unneeded items that are placed specifically to attract those kids' attention.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#460
post #3

>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…

The HTTP spec draws such a distinction, albeit implicitly, in the form (and name) of its concept of "user agent."

Over time it degraded into declaring compatibility with a bunch of different browser engines and doesn't reflect the actual agent anymore.

And very likely Perplexity is in fact using a Chrome-compatible engine to render the page.

Post reply on HN