Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

201–210 of 555 posts

Re: Perplexity AI is lying about their user agent

#201
post #197

Earlier quoted context omitted.

I'm curious to know where you draw the line for what constitutes legitimate manipulation by a person and when it becomes distribution. I'm assuming that if I write code by hand for every part of the TCP/IP and HTTP stack I'm safe. What if I use libraries written by other people for the TCP/IP and HTTP part? What if I use a whole FOSS web browser? What about a paid local web browser? What if I run a script that I wrot…

Obviously not legal advice and I doubt it's entirely settled law, but probably this step > What if I decide to offer that script as a service for free to friends and family, who can use my cloud server? You're allowed to make copies and adaptations in order to utilize the program (website), which probably covers a cloud server you yourself are controlling. You aren't allowed to do other things with those copies thoug…

Why is that the line and not a paid web browser? What about a paid web browser whose primary feature is a really powerful ad blocker?

Re: Perplexity AI is lying about their user agent

#202

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…

This is why media publishers went behind paywalls to get away from Google News

Re: Perplexity AI is lying about their user agent

#203
post #197

Earlier quoted context omitted.

Obviously not legal advice and I doubt it's entirely settled law, but probably this step > What if I decide to offer that script as a service for free to friends and family, who can use my cloud server? You're allowed to make copies and adaptations in order to utilize the program (website), which probably covers a cloud server you yourself are controlling. You aren't allowed to do other things with those copies thoug…

Why is that the line and not a paid web browser? What about a paid web browser whose primary feature is a really powerful ad blocker?

Why would a paid web browser be the line?

No one is distributing copies of anything to anyone then apart from the website that owns the content lawfully distributing a copy to the user.

Also why is a paid web browser any different than a free one?

Re: Perplexity AI is lying about their user agent

#204
post #3

I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.

More than this, I'd rather use a tool which lets me fake the user agent like I can in my browser.

Re: Perplexity AI is lying about their user agent

#205

Earlier quoted context omitted.

A billion dollar settlement is more than enough to fuel further litigation.

> billion dollar settlement is more than enough to fuel further litigation The choice isn’t between a settlement and no settlement. It’s between settlement and fighting in court. Binding precedent and a public right increase the risks and costs to OpenAI, particularly if it looks like they’ll lose.

Right, but a billion dollars to a relative small fry in the publishing industry (even online only) like the ny times is chum in the water.

The next six publishers are going to be looking for $100B and probably have the funds for better lawyers.

At some point these are going to hit the courts, an NY Times probably makes sense as the plaintiff as opposed to one of the larger publishing houses.

Re: Perplexity AI is lying about their user agent

#207

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

> I don't want to live in a world where website owners can use DRM to force me to display their website in exactly the way that their designers envisioned it.

I'm okay with this world, as a tradeoff. I'm not sure users should have _the right_ to reformat others' content.

Re: Perplexity AI is lying about their user agent

#208
post #41

Earlier quoted context omitted.

If using copyrighted material to train an LLM is theft, so is reading a book.

Is reading a book the same as photocopying it for sale? Which of the scenarios above is more similar to using it to train a LLM?

If I was forced to pick, LLMs are closer to reading than to photocopying.

But, and these are important, 1) quantity has a quality all of its own, and 2) if a human was employed to answer questions on the web, then someone asked them to quote all of e.g. Harry Potter, and this person did so, that's still copyright infringement.

Re: Perplexity AI is lying about their user agent

#209

Earlier quoted context omitted.

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…

> A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content. If I visit your site from Google with my browser configured to go straight to Reader Mode whenever possible, is my visit more useful to you than a summary and a link to your site provid…

Well for one thing you visiting his site and displaying it via reader mode doesn't remove his ability to sell paid licenses for his content to companies that would like to redistribute his content. Meanwhile having those companies do so for free without a license obviously does.

Re: Perplexity AI is lying about their user agent

#210
post #41

Earlier quoted context omitted.

If using copyrighted material to train an LLM is theft, so is reading a book.

So if I get access to the Perplexity AI source code (I borrow it from a friend), read all of it, and reproduce it at some level, then Perplexity will be:" sure, that's fine no harm, no IP theft, no copyright violation, because you read it so we're good"? No, they would sue me for everything I got, and then some. That's the weird thing about these companies, they are never afraid to use IP law to go after others, but…

If Perplexity’s source code is downloaded from a public web site or other repository, and you take the time to understand the code and produce your own novel implementation, then yes. Now, if you “get it from a friend”, illegally, _or_ you just redeploy the code, without creating a transformative work, then there’s a problem.

> Just pay the stupid license and if that makes your business unsustainable then it's not much a business is it?

In the persona of a business owner, why pay for something that you don’t legally, need to pay for? The question of how copyright applies to LLMs and other AI is still open. They’d be fools to buy licenses before it’s been decided.

More importantly, we’re potentially talking about the entire knowledge of humanity being used in training. There’s no-one on earth with that kind of money. Sure, you can just say that the business model doesn’t work, but we’re discussing new technologies that have real benefit to humanity, and it’s not just businesses that are training models this way.

Any decision which hinders businesses from developing models with this data will hinder independent researchers 10 fold, so it’s important that we’re careful about what precedent is set in the name of punishing greedy businessmen.

Post reply on HN