Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

21–30 of 555 posts

Re: Perplexity AI is lying about their user agent

#21
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

>and thinking they can get away with it

Can they not? I think that remains to be seen.

Re: Perplexity AI is lying about their user agent

#22
post #8
post #4

Our bot traffic is up 10-fold since LLM Cambrian explosion.

Cambrian explosion implies that there’s a huge variety of different creatures out there, but I suspect those bots are all just wrappers around OpenAI/anthropic models. This is more like the rise of Cyanobacteria as a single early dominant lifeform

There are 112,391 language models on HuggingFace, most of them fine-tunes of a few base models, but still, a staggering number.

Re: Perplexity AI is lying about their user agent

#23
post #12

The only way out seems to be using obscene captcha.

Or detect the LLM and serve up an LLM rewritten version of the page. That way you feed it poisonous garbage.

I really like this idea. Someone needs to implement this. I'm not sure what the ideal poison would be. Randomly constructed sentences that follow the basic rules of grammar?

Re: Perplexity AI is lying about their user agent

#24

OpenAI scraped aggressively for years. Why should others put themselves behind an artificial moat? If you want to block access to a site, stop relying on arbitrary opt-in voluntary things like user agent or robots.txt. Make your site authenticated only, that’s literally the only answer here.

> OpenAI scraped aggressively for years. Why should others put themselves behind an artificial moat?

Not saying I agree/disagree with the whole "LLMs trained on scraped data is unethical", but this way of thinking seems dangerous.

If companies like Theranos can prop up their value by lying, does that make it ok for Theranos competitors to also lie, as another example?

Re: Perplexity AI is lying about their user agent

#25
Just the other day Perplexity CEO Aravind Srinivas was dunking on Google and OpenAI, and putting themselves on a superior moral position because they give citations while closed-book LLMs memorize the web information with large models and don't give credit.

Funny they got caught not following robots.txt and hiding their identity.

https://x.com/tsarnick/status/1801714601404547267

Re: Perplexity AI is lying about their user agent

#27
post #7

How about a trap URL in the Robots.txt file that triggers a 24 hour IP ban if you access it. If you don't want anyone innocent caught in the crossfire, you could make the triggering URL customized to their IP address.

Wouldn't help in this case, the post author banned the bot in the robots for, but then when asked the bot to fetch his web page explicitly by URL...

If a user has a bot directly acting on their behalf (not for training), I think that's fair use... And important to think twice before we block that, since it will be used for accessibility.

Re: Perplexity AI is lying about their user agent

#28
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

Can’t wait for OpenAI to settle with The New York Times. For a billion dollars no less.

Re: Perplexity AI is lying about their user agent

#30
Wow. The user agent they are using is so shady. But I am surprised they thought someone wouldn’t do just what the blog poster did to uncover the deception - that part is what surprises me most.

Other than being unethical, is this not illegal? Any IP experts in here?

Post reply on HN