Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

321–330 of 555 posts

Re: Perplexity AI is lying about their user agent

#321

Earlier quoted context omitted.

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…

I'm not sure what you mean exactly. If Perplexity is actually doing something with your article in-band (e.g. downloading it, processing it, and present that processed article to the user) then they're just breaking the law. I've never used that tool (and don't plan to) so I don't know. If they just embed the content in an iframe or something then there's no issue (but then there's no need or point in scraping). If t…

Sure it is, but which of the many small websites are going to be able to fight them legally? Most companies would go broke before getting a ruling.

Reality is, the law doesn't matter if you're big enough. As long as they're not stealing content from the big ones, they're going to be fine.

Re: Perplexity AI is lying about their user agent

#322
post #113
post #3

I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.

I agree with that, but I also think that they should at least identify themselves instead of using a generic user agent.

I’d rather share less information than more to any site I visit. Why does a user want to share that info?

Re: Perplexity AI is lying about their user agent

#323
post #183

Earlier quoted context omitted.

What will happen if: Website owners decide to stop publishing because it’s not rewarded by a real human visit anymore? Then perplexity and the like won’t have new information to train their models on and no sites to answer the questions. I think there is a real content dilemma here at work. The incentives of Google and website owners were more or less aligned. This is not the case with perplexity.

What is a "visit"? TFA demonstrates that they got a hit on their site, that's how they got the logs. Is it necessary to load the JavaScript for it to count as a visit? What if I access the site with noscript? Or is it only a visit if I see all your recommended content? I usually block those recommendations so that I don't get distracted from the article I actually came to read—is my visit a less legitimate visit than…

A visit is a human reader.

At the very least they get exposed to your website name.

Notice your product/service if you get lucky.

Become a customer at a later visit.

We are talking about cutting the first step off so that everything which may come afterwards is cut off as well.

Re: Perplexity AI is lying about their user agent

#324
post #277

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

Personally I think AI is a major win for accessibility and we should not be preventing people to access information in the way that is best suited for them. Accessibility can mean everything from a blind person wanting to interacting with a website using voice, to someone recovering from a surgery and wanting something to reduce unnecessary popups and clicks on a website to get to the information they need. Accessibi…

> The way I see it, AI is not a robot and doesn't need to look at robots.txt

I don't think you are seeing it very clearly then. Your secretary can also be a robot. What do you think an AI is if not a robot??

It doesn't "need" to look at robots.txt because nothing does.

Re: Perplexity AI is lying about their user agent

#325
post #211
post #183

Earlier quoted context omitted.

What will happen if: Website owners decide to stop publishing because it’s not rewarded by a real human visit anymore? Then perplexity and the like won’t have new information to train their models on and no sites to answer the questions. I think there is a real content dilemma here at work. The incentives of Google and website owners were more or less aligned. This is not the case with perplexity.

How would an LLM training on your writing reduce your reward? I guess if you're doing it for a living sure, but most content I consume online is created without incentive (social media, blogs, stack overflow). I write a fair amount and have been for a few years. I like to play with ideas. If an llm learned from my writing and it helped me propagate my ideas, I'd be happy. I lose on social status imaginary internet po…

Speaking as an SO contributor, I'm perfectly fine with having an LLM read my answers and produce output based on them. What I'm not okay with is said LLM being closed-weight so that its creator can profit off it. When I posted my answers on SO, I did so under CC-BY-SA, and I don't think it's unreasonable for me to expect any derivatives to abide by both the letter and the spirit of this arrangement.

Re: Perplexity AI is lying about their user agent

#326

Earlier quoted context omitted.

Website terms for unauthenticated users and robots.txt have zero legal standing, so it doesn’t matter how much hand-wringing people like the OP do. It would be irresponsible as a business owner to hamstring themselves.

Then they should just say that outright instead of pretending they right thing.

They're not lying, you just misunderstood their docs [0].

> To provide the best search experience, we need to collect data. We use web crawlers to gather information from the internet and index it for our search engine.

> You can identify our web crawler by its user agent

To anyone who's familiar with web crawling and indexing, these paragraphs have an obvious meaning: Perplexity has a search engine which needs a crawler which crawls the internet. That crawler can be identified by the User-Agent PerplexityBot and will respect robots.txt.

Separately, if you give Perplexity a specific URL then it will go fetch the contents of that URL with a one-off request. That one-off request does not respect robots.txt any more than curl does, and that's 100% normal and ethical. The one-off request handler isn't PerplexityBot, it's a separate part of the application that's probably just a regular Chrome browser that issues the request.

[0] https://docs.perplexity.ai/docs/perplexitybot

Re: Perplexity AI is lying about their user agent

#327
post #212

Earlier quoted context omitted.

I cannot imagine how viewing/scraping a public website could ever be illegal, wrong, immoral etc. I just don't see the argument for it.

AI hysteria has made everyone lose their minds over normal things.

I guess people just LOVE twisting themselves in knots over some "ethical scandals" or whatnot. Maybe there's a statement on American puritanism hiding somewhere here...

Re: Perplexity AI is lying about their user agent

#328

Earlier quoted context omitted.

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…

> But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content (via Google, for example). Perplexity has source references. I find myself visiting the source references. Especially to validate the LLM output. And to learn more about the subject. Perplexity uses a Google search API to generate the reference li…

>> In a world where users ask Perplexity these Help questions about my SaaS, Perplexity may answer them and I would lose all the insights because I never get any traffic.

Alternative take: Perplexity is protecting users' privacy by not exposing them to be turned into "insights" by the SaaS.

My general impression is that the subset of complaints discussed in this thread and in the article, boils down to a simple conflict of interest: information supplier wants to exploit the visitor through advertising, upsells, and other time/sanity-wasting things; for that, they need to have the visitor on their site. Meanwhile, the visitors want just the information without the surveillance, advertising and other attention economy dark/abuse patterns.

The content is the bait, and ad-blockers, Google's instant results, and Perplexity, are pulling that bait off the hook for the fish to eat. No surprise fishermen are unhappy. But, as a fish, I find it hard to sympathize.

Re: Perplexity AI is lying about their user agent

#329
post #310

Earlier quoted context omitted.

Sure. So just return an HTTP 4XX response to requests you don't like. What's the problem?

Or, I return whatever content I want, within the bounds of the law, based on whatever parameters I decide. What's your problem with that? Again, connect to my server or don't. But don't tell me what type of response I'm obligated to provide you. If I think a given request is from an LLM training module, I don't have any legal obligation whatsoever to return my original content. Or a 400-series response. If I want to…

But nobody is arguing for that. Instead, what the server owners want is to mandate the clients connecting to them to provide enough information to reliably reject such connections.

Re: Perplexity AI is lying about their user agent

#330

Earlier quoted context omitted.

That line of thinking makes no sense. If the "content" had no value, why would google go through the effort of scraping it and presenting it to the user?

It's not that it has no value, it's that there is no established way (other than ad revenue) to charge users for that content. The fact that google is able to monetize ad revenue at least as well as, and probably better than, almost any other entity on the internet, means that big-G is perfectly positioned to cut out the creator -- until the content goes stale, anyway.

> until the content goes stale, anyway

This will be quite interesting in the future. One can usually tell if a blog post is stale, or whether it’s still relevant to the subject it’s presenting. But with LLMs they’ll just aggregate and regurgitate as if it was a timeless fact.

Post reply on HN