Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

141–150 of 555 posts

Re: Perplexity AI is lying about their user agent

#141
post #25

Just the other day Perplexity CEO Aravind Srinivas was dunking on Google and OpenAI, and putting themselves on a superior moral position because they give citations while closed-book LLMs memorize the web information with large models and don't give credit. Funny they got caught not following robots.txt and hiding their identity. https://x.com/tsarnick/status/1801714601404547267

Nobody follows robots.txt, because every site's robots.txt forbids anybody that isn't google from looking at it.

Also, "hiding their identity" is what every single browser does since Mosaic changed its name.

Re: Perplexity AI is lying about their user agent

#143
post #77
post #31

Earlier quoted context omitted.

The lie is in their documentation - they claim to use the PerplexityBot string in their user-agent: https://docs.perplexity.ai/docs/perplexitybot .

It's not a lie. This is the agent string of the bot used for ingesting data for training the AI. In the blog post, this is not what is happening. It is merely feeding the webpage as context to the AI during inference. You are all confused here.

Website owners should be able to block this behavior as well — OpenAI has two different agents and doesn't obscure the agent when a user initiates a fetch.

Re: Perplexity AI is lying about their user agent

#145

Earlier quoted context omitted.

Agree - the first movers who scraped before changes to websites terms and robots files shouldn’t get an unfair advantage. That’s overall bad for society in terms of choice and competition

Website terms for unauthenticated users and robots.txt have zero legal standing, so it doesn’t matter how much hand-wringing people like the OP do. It would be irresponsible as a business owner to hamstring themselves.

Then they should just say that outright instead of pretending they right thing.

Re: Perplexity AI is lying about their user agent

#147
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

Right, it's ironic we spent 30 years fighting piracy and then suddenly corporations start doing it and now it's suddenly ok.

Re: Perplexity AI is lying about their user agent

#149
post #84

The author has misunderstood when the perplexity user agent applies. Web site owners shouldn’t dictate what browser users can access their site with - whether that’s chrome, firefox, or something totally different like perplexity. When retrieving a web page _for the user_ it’s appropriate to use a UA string that looks like a browser client. If perplexity is collecting training data in bulk without using their UA that…

Setting a correct user agent isn't required anyway, you just do it to not be an asshole. Robots.txt is an optional standard.

The article is just calling Perplexity out for some asshole behavior, it's not that complicated

It's clear they know they're engaging in poor behavior too, they could've documented some alternative UA for user-initiated requests instead of spoofing Chrome. Folks who trust them could've then blocked the training UA but allowed the alternative

Re: Perplexity AI is lying about their user agent

#150
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

I hate to argue this side of the fence, but when ai companies are taking the work of writers and artists en mass (replacing creative livelihoods with a machine trained on the artists stolen work) and achieving billion dollar valuations that’s actual stealing.

The key here is that creative content producers are being driven out of business through non consensual taking of their work.

Maybe it’s a new thing, but if it is, it’s worse than stealing.

Post reply on HN