Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

471–480 of 555 posts

Re: Perplexity AI is lying about their user agent

#471

Earlier quoted context omitted.

You are not. Perplexity should always respect robots.txt, even for summarization requests. If I say that I don't want Perplexity crawling my site, I mean at all , and I explicitly would not want them "summarizing" my page. The response from Perplexity to such a request should be "The owner of this page/site does not permit Perplexity to process any data from this site." Period. LLMs can't summarize in any case: https…

> Perplexity should always respect robots.txt, even for summarization requests. If I say that I don't want Perplexity crawling my site, I mean at all Issuing a single HTTP request is definitionally not crawling, and the robots.txt spec is specifically for crawlers, which this is not. If you want a specific tool to exclude you from their web request feature you have to talk to them about it. The web was designed to ma…

You are definitionally incorrect. From Wikipedia:

> robots.txt is the filename used for implementing the Robots Exclusion Protocol, a standard used by websites to indicate to visiting web crawlers and other web robots which portions of the website they are allowed to visit.

From robotstxt.org/orig.html (the original proposed specification), there is a bit about "recursive" behaviour, but the last paragraph indicates "which parts of their server should not be accessed".

> WWW Robots (also called wanderers or spiders) are programs that traverse many pages in the World Wide Web by recursively retrieving linked pages. For more information see the robots page.

> In 1993 and 1994 there have been occasions where robots have visited WWW servers where they weren't welcome for various reasons. Sometimes these reasons were robot specific, e.g. certain robots swamped servers with rapid-fire requests, or retrieved the same files repeatedly. In other situations robots traversed parts of WWW servers that weren't suitable, e.g. very deep virtual trees, duplicated information, temporary information, or cgi-scripts with side-effects (such as voting).

> These incidents indicated the need for established mechanisms for WWW servers to indicate to robots which parts of their server should not be accessed. This standard addresses this need with an operational solution.

The draft RFC at robotstxt.org/norobots-rfc.txt, the definition is a little more strict about "recursive", but indicates that heuristics used and/or time spacing do not make it less a robot.

On robotstxt.org/faq/what.html, there is a paragraph:

> Normal Web browsers are not robots, because they are operated by a human, and don't automatically retrieve referenced documents (other than inline images).

One might argue that the misbehaviour of Perplexity on this matter is "at the instruction" of a human, but as Perplexity does not present itself as a web browser, but a data processing entity, it’s clearly not a web browser.

Here's what would be permitted unequivocally, even on a site that blocks bad actors like Perplexity: a browser extension that used Perplexity's LLM to pretend to summarize but actually shorten the content (https://ea.rna.nl/2024/05/27/when-chatgpt-summarises-it-actu...) when you visit the page as long as that summary were not saved in Perplexity's data.

Re: Perplexity AI is lying about their user agent

#472

Earlier quoted context omitted.

Should there be a difference in treatment between a user going on a website and manually copying the content over to a bot to process vs giving the bot the URL so it does the fetching as well? I've done both (mainly to get summaries or translations) and I know which I generally prefer.

Ideally no, but there are established norms and unwritten rules. Plus, a mechanism was built to communicate the limits. These norms were working for decades. The fences were reasonable because the demands were reasonable and both sides understood why they are there and respected these borders. This peace has been broken, norms are thrown away and people who did this cheered for what they did. Now, the people are figh…

Well, what'll happen for the most part is not users being mad, but a general migration to fenceless areas. Prompts will be for "content similar to X" and the bots will merely use what it has access to, rendering the fences moot. And there will always be authors who don't mind their content being monitized or utilized by AI.

Re: Perplexity AI is lying about their user agent

#473

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

[deleted]

Re: Perplexity AI is lying about their user agent

#474

Earlier quoted context omitted.

> Perplexity should always respect robots.txt, even for summarization requests. If I say that I don't want Perplexity crawling my site, I mean at all Issuing a single HTTP request is definitionally not crawling, and the robots.txt spec is specifically for crawlers, which this is not. If you want a specific tool to exclude you from their web request feature you have to talk to them about it. The web was designed to ma…

You are definitionally incorrect. From Wikipedia: > robots.txt is the filename used for implementing the Robots Exclusion Protocol, a standard used by websites to indicate to visiting web crawlers and other web robots which portions of the website they are allowed to visit. From robotstxt.org/orig.html (the original proposed specification), there is a bit about "recursive" behaviour, but the last paragraph indicates…

Every paragraph that you've included up there just reinforces my point.

The recursive behavior isn't incidental, it's literally part of the definition of a crawler. You can't just skip past that and pretend that the people who specifically included the word recursive (or the phrase "many pages") didn't really mean it.

The first paragraph of the two about access controls is the context for what "should not be accessed" means. It refers to "very deep virtual trees, duplicated information, temporary information, or cgi-scripts with side-effects (such as voting)", which are pages that should not be indexed by search engines but for the most part shouldn't be a problem for something like perplexity. As I said in my comment, it's about search engine crawlers and indexers.

I'm glad that you at least cherry-picked a paragraph from that second page, because I was starting to worry that you weren't even reading your sources to check if they support your argument. That said, that paragraph means very little in support of your argument (it just gives one example of what isn't a robot, which doesn't imply that everything else is) and you're deliberately ignoring that that page is also very specific about the recursive nature of the robots that are being protected against.

Again, this is the definition that you just cited, which can't possibly include a single request from Perplexity's server (emphasis added):

> WWW Robots (also called wanderers or spiders) are programs that traverse many pages in the World Wide Web by recursively retrieving linked pages.

The only way you can possibly apply that definition to the behavior in TFA is if you delete most of it and just end up with "programs ... that traverse ... the WWW", at which point you've also included normal web browsers in your new definition.

It honestly just feels like you really have a lot of beef with LLM tech, which is fair, but there are much better arguments to be made against LLMs than "Perplexity's ad hoc requests are made by a crawler and should respect robots.txt". Your sources do not back up what you claim—on the contrary, they support my claim in every respect—so you should either find better sources or try a different argument.

Re: Perplexity AI is lying about their user agent

#475

Earlier quoted context omitted.

It's not that it has no value, it's that there is no established way (other than ad revenue) to charge users for that content. The fact that google is able to monetize ad revenue at least as well as, and probably better than, almost any other entity on the internet, means that big-G is perfectly positioned to cut out the creator -- until the content goes stale, anyway.

> until the content goes stale, anyway This will be quite interesting in the future. One can usually tell if a blog post is stale, or whether it’s still relevant to the subject it’s presenting. But with LLMs they’ll just aggregate and regurgitate as if it was a timeless fact.

This is already a problem. Content farms have realised that adding "in $current_year" to their headlines helps traffic. It's frustrating when you start reading and realise the content is two years out of date.

Re: Perplexity AI is lying about their user agent

#476

Earlier quoted context omitted.

I'd argue it only demonstrates that it doesn't produce much value for the creator.

The Google summaries (before whatever LLM stuff they're doing now) are 2-3 sentences tops. The content on most of these websites is much, much longer than that for SEO reasons. It sucks that Google created the problem on both ends, but the content OP is referring to costs way more to produce than it adds value to the world because it has to be padded out to show up in search. Then Google comes along and extracts the…

It would be a lot better if Google just prioritised concise websites.

If Google preferred websites that cut the fluff, then website operators would have an incentive to make useful websites, and Google wouldn't have as much of an incentive to provide the answer in a snippet, and everyone wins.

I guess it's hard to rank website quality, so Google just prefers verbose websites.

Re: Perplexity AI is lying about their user agent

#477

Earlier quoted context omitted.

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…

> they are decreasing the probability that this user would come to by content (via Google, for example). Google has been providing summaries of stuff and hijacking traffic for ages. I kid you not, in the tourism sector this has been a HUGE issue, we have seen 50%+ decrease in views when they started doing it. We paid gazzilions to write quality content for tourists about the most different places just so Google could…

Google has been in trouble for doing so several times in the past and removed key features because of it. Examples: Viewing cached pages, linking directly to images, summarized news articles.

Re: Perplexity AI is lying about their user agent

#478

Earlier quoted context omitted.

In London, uber did not succeed. Uber drivers have to be licensed like minicab drivers.

Uber is widely used in London, so they succeeded. If they had waited decades for the regulatory landscape to even out they would have failed.

They succeeded commercially, but they didn't succeed in changing the regulatory landscape. I'm not sure what you mean by waiting for it to even out. They refused to comply, so they were banned, so they complied.

Re: Perplexity AI is lying about their user agent

#479

Earlier quoted context omitted.

… and I have absolutely no obligation to provide any particular response to any particular client. Parsing, rendering, and trusting that the payload is consistent from request to request is your problem . You can connect to my server, or not. I really don’t care. What you cannot do is dictate how my server responds to your request.

> What you cannot do is dictate how my server responds to your request. The client is under no obligation to be truthful in its communications with a server. Spoofing a User-Agent doesn't "dictate" anything. Your server dictates how it responds all on its own when it discriminates against some User-Agents.

With enough sophistication and bad intent, at some point being untruthful to a server falls under computer intrusion laws, eg using a password that is not yours. I don't believe spoofing user agent would be determinant for any such case though.

Even redistributing secret material you found on an accidentally open S3 bucket, without spoofing UA, could be considered intrusion if it was obvious the material was intended to be secret and you acted with bad intent.

Re: Perplexity AI is lying about their user agent

#480

Earlier quoted context omitted.

I cannot imagine how viewing/scraping a public website could ever be illegal, wrong, immoral etc. I just don't see the argument for it.

It's scraping content to then serve up that content to users who can now get that content from you (via a paid subscription service, or maybe ad-sponsored) instead of visiting the content creator and paying them (i.e., via ads on their website) It's the same reason I can't just take NYT archives or the Britannica and sell an app that gives people access to their content through my app. It totally undercuts content cr…

Serve it in a better way or wall it. The Internet is supposed to be free. If you don't want unauthorized eyes to see it, you have the ability to hide it behind logins.
Post reply on HN