Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

201–210 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#201
As others have mentioned the problem is that of scale. Perhaps there needs to be a rate limit (times they ping a site) set within robots.txt that a site bot can come but only X times per hour etc. At least we move from a binary scrape or no scrape to a spectrum then.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#202
post #190

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. > If you want to gatekeep your content, use authentication. Are there no limits on what you use the content for? I can start my own search engine that just scrapes Google results?

Yes, I believe that's basically what https://serpapi.com/ is doing.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#203

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

How about I open a proxy, replace all ads with my ads, redirect the content to you and we share the ad revenue?

That's the Brave browser.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#204
> The Internet as we have known it for the past three decades is rapidly changing, but one thing remains constant: it is built on trust.

I think we've been using different internets. The one I use doesn't seem to be built on trust at all. It seems to be constantly syphoning data from my machine to feed the data vampires who are, apparently, additing to (I assume, blood-soaked) cookies

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#205

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

To stretch the analogy to the breaking point: If you send 10,000 personal shoppers all at once to the same store just to check prices, the store's going to be rightfully annoyed that they aren't making sales because legit buyers can't get in.

Your comment and the above comment of course show different cases.

An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways.

But the sort of non-explicit just-in-case crawling that Perplexity might do for a general question where it crawls 4-6 sources isn't as easy to defend. "Are polar bears always white?" -- Now it's making requests I wouldn't have necessarily made, and it could even been seen as a sort of amplification attack.

That said, TFA's example is where they register secretexample.com and then ask Perplexity "what is secretexample.com about?" and Perplexity sends a request to answer the question, so that's an example of the first case, not the second.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#206
post #152

I am sorry, Cloudafre is the internet police now?

They hate AI it seems. I don’t see them offering any AI products or embracing it in any way. Seems like they’ll get left behind in the AI race.

If they managed to enforce the pay-per-scrape, that would be a huge revenue, bigger than AdSense

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#207
post #3

>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…

The HTTP spec draws such a distinction, albeit implicitly, in the form (and name) of its concept of "user agent."

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#208

Earlier quoted context omitted.

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

But I can send my personal shopper and you'll be none the wiser.

Sure. There's lots of things you could do, but you don't do them because they are wrong.

Might does not make right.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#209

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process.

> Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control.

How does one follow the other? It's my web server and I can gatekeep access to my content however I want (eg Cloudflare). How is that an "abuse" of internet protocols?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#210
post #95

AI companies continuing to have problems with the concept of "consent" is increasingly alarming god help us if they ever manage to build anything more than shitty chatbots

They're certainly pouring billions of dollars into trying to build something more. Or at least that's what they're telling the public and investors.
Post reply on HN