Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

211–220 of 555 posts

Re: Perplexity AI is lying about their user agent

#211
post #183

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

What will happen if: Website owners decide to stop publishing because it’s not rewarded by a real human visit anymore? Then perplexity and the like won’t have new information to train their models on and no sites to answer the questions. I think there is a real content dilemma here at work. The incentives of Google and website owners were more or less aligned. This is not the case with perplexity.

How would an LLM training on your writing reduce your reward?

I guess if you're doing it for a living sure, but most content I consume online is created without incentive (social media, blogs, stack overflow).

I write a fair amount and have been for a few years. I like to play with ideas. If an llm learned from my writing and it helped me propagate my ideas, I'd be happy. I lose on social status imaginary internet points but I honestly don't care much for them.

The craziest one is the stack overflow contributors. They write answers for free to help people become better programmers but they're mad an llm will read their suggestions and answer questions that help people become better programmers. I guess they do it for the glory of having their handle next to the answer?

Re: Perplexity AI is lying about their user agent

#212

Earlier quoted context omitted.

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

I cannot imagine how viewing/scraping a public website could ever be illegal, wrong, immoral etc. I just don't see the argument for it.

AI hysteria has made everyone lose their minds over normal things.

Re: Perplexity AI is lying about their user agent

#213
post #28

Earlier quoted context omitted.

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

Can’t wait for OpenAI to settle with The New York Times. For a billion dollars no less.

Settling for a billion dollars would be insane. They'd immediately get sued by everyone who ever posted anything on the internet.

Re: Perplexity AI is lying about their user agent

#214
post #203

Earlier quoted context omitted.

Why is that the line and not a paid web browser? What about a paid web browser whose primary feature is a really powerful ad blocker?

Why would a paid web browser be the line? No one is distributing copies of anything to anyone then apart from the website that owns the content lawfully distributing a copy to the user. Also why is a paid web browser any different than a free one?

Paid is arguably different than free because the code that is actually asking for the data is owned by a company and licensed to the user, in much the same way as a cloud server licenses usage of their servers to the user. That said, I'll note that my argument is explicitly that the line doesn't exist, so I'm not saying a paid browser is the line.

I'm unfamiliar with the legal questions, but in 2024 I have a very hard time seeing an ethical distinction between running some proprietary code on my machine to complete a task and running some proprietary code on a cloud server to complete a task. In both cases it's just me asking someone else's code to fetch data for my use.

Re: Perplexity AI is lying about their user agent

#215

Earlier quoted context omitted.

Exactly. It's like when Uber started and flaunted the medallion taxi system of many cities. People said "These Uber people are idiots! They are going to get shut down! Don't they know the laws for taxis?" While a small number of cities did ban Uber (and even that generally only temporarily), in the end Uber basically won. I think a lot of people confuse what they want to happen versus what will happen.

In London, uber did not succeed. Uber drivers have to be licensed like minicab drivers.

Uber is widely used in London, so they succeeded.

If they had waited decades for the regulatory landscape to even out they would have failed.

Re: Perplexity AI is lying about their user agent

#216
Respecting robots.txt is something their training crawler should do, and I see no reason why their user agent (i.e. user asks it to retrieve a web page, it does) should, as it isn't a crawler (doesn't walk the graph).

As to "lying" about their user agents - this is 2024, the "User-Agent" header is considered a combination bug and privacy issue, all major browsers lie about being a browser that was popular many years ago, and recently the biggest browser(s?) standardized on sending one exact string from now on forever (which would obviously be a lie). This header is deprecated in every practical sense, and every user agent should send a legacy value saying "this is mozilla 5" just like Edge and Chrome and Firefox do (because at some point people figured out that if even one website exists that customizes by user agent but did not expect that new browsers would be released, nor was maintained since, then the internet would be broken unless they lie). So Perplexity doing the same is standard, and best, practice.

Re: Perplexity AI is lying about their user agent

#217

Earlier quoted context omitted.

> billion dollar settlement is more than enough to fuel further litigation The choice isn’t between a settlement and no settlement. It’s between settlement and fighting in court. Binding precedent and a public right increase the risks and costs to OpenAI, particularly if it looks like they’ll lose.

Right, but a billion dollars to a relative small fry in the publishing industry (even online only) like the ny times is chum in the water. The next six publishers are going to be looking for $100B and probably have the funds for better lawyers. At some point these are going to hit the courts, an NY Times probably makes sense as the plaintiff as opposed to one of the larger publishing houses.

> ny times is chum in the water

The Times has a lauder litigation team. Their finances are good and their revenue sources diverse. They’re not aching to strike a deal.

> NY Times probably makes sense as the plaintiff as opposed to one of the larger publishing houses

Why? Especially if this goes to a jury.

Re: Perplexity AI is lying about their user agent

#218
post #164

Earlier quoted context omitted.

> The second concern, though, is can perplexity do a live web query to my website and present data from my website in a format that the user asks for? Arguing that we should ban this moves into very dangerous territory. This feels like the fundamental core component of what copyright allows you to forbid. > Everything from ad blockers to reader mode to screen readers do exactly the same thing that Perplexity is doing…

I'm curious to know where you draw the line for what constitutes legitimate manipulation by a person and when it becomes distribution. I'm assuming that if I write code by hand for every part of the TCP/IP and HTTP stack I'm safe. What if I use libraries written by other people for the TCP/IP and HTTP part? What if I use a whole FOSS web browser? What about a paid local web browser? What if I run a script that I wrot…

Where exactly you crossed the line is a question for the courts. I am not a lawyer and will there for not help with the specifics.

However, please see the Aereo case [0] for a possibly analogous case. I am allowed to have a DVR. There is no law preventing me from accessing my DVR over a network. Or possibly even colocating it in a local data center. But Aereo definitely crossed a line. Also see Vidangel [1]. The fact that something is legal to do at home, does not mean that I can offer it as a cloud service.

[0] https://www.vox.com/2018/11/7/18073200/aereo

[1] https://en.m.wikipedia.org/wiki/Disney_v._VidAngel

Re: Perplexity AI is lying about their user agent

#219
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

> piracy is not theft

it was when Napster was doing it; but there's no entity like the RIAA to stop the AI bots

Re: Perplexity AI is lying about their user agent

#220

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

Citing the source doesn't bring you, the owner of the site, valuable data. When was your data accessed, who accessed it, from where, at what time, what device, etc. It brings data to the LLM's owner, and you get N O T H I N G. Could you change the way printed news magazines showed their content? No. Then, why is that a problem? Btw nobody clicks on sources. NOBODY.

> Btw nobody clicks on sources. NOBODY.

I always click on sources to verify what an LLM in this case says. I also hear the claim that a lot about people not reading sources (before LLM it was video content with references) but I always visited the sources. Is there a statistics or studies that actually support this claim? Or is it just a personal experience, of people (including me) enforcing it as generic behavior of all people?

Post reply on HN