Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

171–180 of 555 posts

Re: Perplexity AI is lying about their user agent

#171
> What is this post about https://rknight.me/blog/blocking-bots-with-nginx/

He is asking Perplexity to summarize a single page. This is simply automation for opening a browser, navigating to that URL, copying the content, pasteing the content into Perplexity.

This is not automated crawling or indexing. Since the person is driving the action. An automated crawler is driven into action by a bot.

Nor is this article added into the foundational model. It's simply in a person's session context.

If for some reason, the community deems this as automated crawling or indexing. One could write an extension to automate the process of copying the article content & pasting the content into an LLM/Rag like Perplexity.

Re: Perplexity AI is lying about their user agent

#172
post #122

Earlier quoted context omitted.

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

Aereo, Napster, Grokster, Grooveshark, Megaupload, and TVEyes: they all thought the same thing. Where are they now?

Heh, you're right, of course, but as someone who came of age on the internet around that era, it still seems strange to me that people these days are making the arguments the RIAA did. They were the big bad guys in my day.

Re: Perplexity AI is lying about their user agent

#173

Earlier quoted context omitted.

> Only reason OpenAI would do that would be to create a barrier for smaller entrants Only? No. Not even main. The main reason would be to halt discovery and setting a precedent that would fuel not only further litigation but also, potentially, legislation. That said, OpenAI should spin it as that master-of-the-universe take.

A billion dollar settlement is more than enough to fuel further litigation.

> billion dollar settlement is more than enough to fuel further litigation

The choice isn’t between a settlement and no settlement. It’s between settlement and fighting in court. Binding precedent and a public right increase the risks and costs to OpenAI, particularly if it looks like they’ll lose.

Re: Perplexity AI is lying about their user agent

#174
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

> piracy is not theft

Correct, but it is often a licensing breach (though sometimes depending upon the reading of some licenses, again these things are yet to be tested in any sort of court) and the companies doing it would be very quick to send a threatening legal letter if we used some of their output outside the stated licensing terms.

Re: Perplexity AI is lying about their user agent

#175

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy.

But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content (via Google, for example). This is unacceptable. A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content.

There are many benefits to having people visit your content on a property that you own. e.g., say you are a SaaS company and you have a bunch of Help docs. You can analyze traffic in this section of your website to get insights to improve your business: what are the top search queries from my users, this might indicate to me where they are struggling or what new features I could build. In a world where users ask Perplexity these Help questions about my SaaS, Perplexity may answer them and I would lose all the insights because I never get any traffic.

Re: Perplexity AI is lying about their user agent

#176
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

You wouldn't train an LLM on a car.

Re: Perplexity AI is lying about their user agent

#177

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

The companies will scrape and internalise the "customer asked for this" requests... and slowly turn the latter into the former, or just their own tool as the scraper.

No, easier to just ask a simple question: Does the company respect the access rules communicated via a web standard? No? In that case hard deny access to that company.

These companies don't need to be given an inch.

Re: Perplexity AI is lying about their user agent

#178

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

[deleted]

Re: Perplexity AI is lying about their user agent

#179
post #137
post #111

Earlier quoted context omitted.

> It’s retrieving the content then manipulating it. Perplexity isn’t a web browser. So a browser with an ad-blocker that's removing / manipulating elements on the page isn't a browser? What about reader mode?

How a user views a page isn't the same as a startup scraping the internet wholesale for financial gain.

But it's not scraping, it's retrieving the page on request from the user.

Re: Perplexity AI is lying about their user agent

#180
post #164

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

> The second concern, though, is can perplexity do a live web query to my website and present data from my website in a format that the user asks for? Arguing that we should ban this moves into very dangerous territory. This feels like the fundamental core component of what copyright allows you to forbid. > Everything from ad blockers to reader mode to screen readers do exactly the same thing that Perplexity is doing…

I'm curious to know where you draw the line for what constitutes legitimate manipulation by a person and when it becomes distribution.

I'm assuming that if I write code by hand for every part of the TCP/IP and HTTP stack I'm safe.

What if I use libraries written by other people for the TCP/IP and HTTP part?

What if I use a whole FOSS web browser?

What about a paid local web browser?

What if I run a script that I wrote on a cloud server?

What if I then allow other people to download and use that script on their own cloud servers?

What if I decide to offer that script as a service for free to friends and family, who can use my cloud server?

What if I offer it for free to the general public?

What if I start accepting money for that service, but I guarantee that only the one person who asked for the site sees the output?

Can you help me to understand where exactly I crossed the line?

Post reply on HN