Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

441–450 of 555 posts

Re: Perplexity AI is lying about their user agent

#441
post #295

With all the ad blockers out there, which functionally demonetize content sites, why isn’t there an ad equivalent to robots.txt that says “don’t display this site if ads are blocked”? So many good comments from several points of view in this thread and the thing I can’t square is the same person championing ad blockers and condemning agents like Perplexity.

Because these are all voluntary standards. If you want your content to be discoverable and accessible, you don’t get to dictate how someone renders it. If you want to force monetization, adopt a different business model.

I don’t think you’re following my point (I probably explained it poorly).

People voluntarily agreed to follow the robots.txt model when they could have ignore it. To this day, a plurality of people seem to support that standard.

That doesn’t keep content from being discoverable or accessible. All sorts of ways to find web sites outside of sites that use crawlers — directories, web rings, social media, etc.

There could have been an ads.txt model, but people probably would have likely ignored it. Your response would seem to be the norm for defending ad blockers — you somehow have a right to the content and if they can’t force you to view their ad, that’s on them.

Why do people get to dictate who accesses a page but not how it’s accessed? That binary seems completely arbitrary.

Re: Perplexity AI is lying about their user agent

#442
post #279

Earlier quoted context omitted.

Paid is arguably different than free because the code that is actually asking for the data is owned by a company and licensed to the user, in much the same way as a cloud server licenses usage of their servers to the user. That said, I'll note that my argument is explicitly that the line doesn't exist , so I'm not saying a paid browser is the line. I'm unfamiliar with the legal questions, but in 2024 I have a very ha…

Great, so we agree that your previous comment asking I address "paid browsers" in particular was an unnecessary distraction. > I have a very hard time seeing an ethical distinction between running some proprietary code on my machine to complete a task and running some proprietary code on a cloud server to complete a task It's important to recognize that copyright is entirely artificial. Congress went "let's grant cre…

How is using perplexity any more so making a copy than your browser is making a copy? Unless you are distributing your website on thumb drives or floppy disks all distribution is achieved by making a copy. That's how networks work.

Your logic would also imply that viewing a website through a VPN not operated by yourself would require the VPN operator to have a redistribution license for all the content on the website which is not the case.

How do you think google is able to scrape whatever they like and redistribute summaries of the pages they have visited without consulting everyone who has ever made a website for a redistribution license.

That being said, Copyright is not enforced or interpreted consistently. It seems that individual cases can be decided based on what people ate for lunch on the day of the case, who the litigants are, and maybe the alignment of the planets.

Re: Perplexity AI is lying about their user agent

#443

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

You can poison all your images with Glaze and Nightshade. Then you don't have to stop them from using them - they have to stop themselves from using them or their image generator will be useless. I don't know if there's a comparable system for text. If there was, it would probably be noticeable to humans.

Re: Perplexity AI is lying about their user agent

#444
post #41

Earlier quoted context omitted.

If using copyrighted material to train an LLM is theft, so is reading a book.

So if I get access to the Perplexity AI source code (I borrow it from a friend), read all of it, and reproduce it at some level, then Perplexity will be:" sure, that's fine no harm, no IP theft, no copyright violation, because you read it so we're good"? No, they would sue me for everything I got, and then some. That's the weird thing about these companies, they are never afraid to use IP law to go after others, but…

Funny enough, their prompts leaked: https://www.reddit.com/r/perplexity_ai/s/kn6i20kMLH

And I’ve built a perplexity clone in about a day - it’s not that hard: search -> scrape results -> parse results —> summarize results -> summarize aggregate results into single summary.

I’m really not sure I even see their moat.

Re: Perplexity AI is lying about their user agent

#445

Earlier quoted context omitted.

I'd believe it if they were targeting entities that could fight back, like stock photo companies and disney, instead of some guy with an artstation account, or some guy with a blog. To me it sounds like these products can't exist without exploiting someone and they're too coward to ask for permission because they know the answer is going to be "no." Imagine how many things I could create if I just stole assets from o…

...which is a great argument for abolishing copyright:P

...which is a great argument for how unjust is a law that only protects those that can afford it.

Cheaper processes to protect smaller creators in cases like these is what is really needed.

Re: Perplexity AI is lying about their user agent

#446

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

To follow onto this:

If what Perplexity is doing is illegal, is it illegal to run an open-source LLM on your own machine, and have it do the same thing? If so, how are ad blockers or Reader Modes or screen readers legal?

And if it's legal to run an open-source LLM on your own machine, is it legal to run an open-source LLM on a rented server (e.g. because you need more GPUs)? And if that's legal, why is it illegal to run a closed-source LLM on servers? Could Perplexity simply release the model weights and keep doing what they're doing?

Re: Perplexity AI is lying about their user agent

#447
post #279

Earlier quoted context omitted.

Great, so we agree that your previous comment asking I address "paid browsers" in particular was an unnecessary distraction. > I have a very hard time seeing an ethical distinction between running some proprietary code on my machine to complete a task and running some proprietary code on a cloud server to complete a task It's important to recognize that copyright is entirely artificial. Congress went "let's grant cre…

How is using perplexity any more so making a copy than your browser is making a copy? Unless you are distributing your website on thumb drives or floppy disks all distribution is achieved by making a copy. That's how networks work. Your logic would also imply that viewing a website through a VPN not operated by yourself would require the VPN operator to have a redistribution license for all the content on the website…

TECHNICAL ANALYSIS

The key, as many here have missed, is authentication and authorization. You may have authorization to log in and view movies on Netflix. Not to rebroadcast them. Even the question of a VCR for personal use was debated in the past.

Distributing your own scripts and software to process data is not the same as distributing arbitrary data those scripts encountered on the internet for which you don’t have a license.

If someone wrote an article, your reader transforms it based on your authenticated request, and your user would have an authorized subscription.

But if that reader then sent the article down to a remote server to be processed for distribution to unlimited numbers of people, it would be “pirating” that information.

The problem is that much of the Web is not properly guarded against this. Xanadu had ideas about micropayments 30 years ago. Take a look at what I am building using the current web: https://qbix.com/ecosystem

LEGAL ANALYSIS

Much of the content published on the Web isn’t secured with subscriptions and micropayments, which is why the whole thing becomes a legal battle as silly as “exceeding authorized access” which landed someone like Aaron Swartz in jail.

In other words, it is the question of “piracy”, which has acquired a new character only in that the AI is trained on your data and transforms it before it republishes it.

There was also a lawsuit aboot scraping LinkedIn, which was settled as follows: https://natlawreview.com/article/hiq-and-linkedin-reach-prop...

Legally, you can grant access to people subject to a certain license (eg Creative Commons Share Alike) and then any derived content must have its weights opened. Similar to, say, Affero GPL license for derivative software.

Re: Perplexity AI is lying about their user agent

#448

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

If the user specifically asks for a file and asks a computer program to process it in a specific way, it should be permitted, regardless of user-agent spoofing (although user-agent spoofing should (normally) ideally only be done when the user specifically requests it; it should not do so automatically). However, this is better when using FOSS and/or local programs (or if the user is accessing them through a proxy, VP…

A user-agent requests the file using your credentials, eg a cookie or public key signature.

It is transforming the content for you, an authorized party.

That is not the same as then making derivative copies and distributing the information to others without paying. For example, if I bought a ticket to a show, taped it and then distributed it to everyone, disregarding that the show prohibited this.

If I shared my Netflix password with up to 5 others, at least I can argue that they are part of my “family” or something. But to unlimited numbers of people? Why would they pay for netflix, and how would the shows get made?

I am not necessarily endorsing government force enforcing copyright, which is why I have been building a solution to enforce it at the tech level: https://Qbix.com/ecosystem

Re: Perplexity AI is lying about their user agent

#449

Earlier quoted context omitted.

it's comparable exactly in the way 0.001% can be compared to 10^100 humans learning is the old-school digital copying. computers simply do it much faster, but it's the same basic phenomenon consider one teacher and one student. first there is one idea in one head but then the idea is in two heads. now add book technology1 the teacher writes the book once, a thousand students read it. the idea has gone from being in o…

> humans learning is the old-school digital copying. computers simply do it much faster, but it's the same basic phenomenon This is dangerous framing because it papers over the significant material differences between AI training and human learning and the outcomes they lead to. We all have a collective interest in the well-being of humanity, and human learning is the engine of our prosperity. Each individual has age…

you're reading what I say in the worst possible light

if anything, the parallel I draw between AI learning and humans learning is all the opposite of narrow and logical... in my intent, the analogy is loose and poetic, not mechanistic and exact.

AI are tools, if AI are enslaving is because there are human actors (I hope....) deciding to enslave other humans, not because of anything inherent to training (if AI; learning if humans)

but what I really think is that there are collections of rules (people "just doing their jobs") all collectively but disjointedly deciding that it makes the most sense to utilize AI technology to ensalve other humans because the data models indicate greater profit that way.

Re: Perplexity AI is lying about their user agent

#450

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

Let’s differentiate between:

1) a user-agent which makes an authenticated and authorized request for data, and delivers to the user

2) a user who then turns around and distributes the data or its derivatives to users in an unauthorized manner

A “dumber” example would be whether I can indefinitely cache and index most of information via the Google Places API, as long as my users request each item at least once. Can I duplicate all that map or streetview photo information that google paid cars to go around and photograph? Or how about the info that Google users entered as user-generated content?

THE REQUIREMENT TO OPEN SOURCE WEIGHTS

Legally, if I had a Creative Commons Share-Alike license on my data, and the LLM was trained on it and then served unlimited requests to others, without making the weights available…

…that would be almost exactly like if I had made my code available with Affero GPL license, someone would take my code but then incorporated it into a backend software hosting a social network or something, without making their own entire social network source code available. Technically this should be enforceable via a court order compelling the open sourcing to the public. (Alternatively, they’d have to pay damages in a class action lawsuit and stop using the tainted backend software or weights when serving all those people.)

TECHNICAL ANALYSIS

The key, as many here have missed, is authentication and authorization. You may have authorization to log in and view movies on Netflix. Not to rebroadcast them. Even the question of a VCR for personal use was debated in the past.

Distributing your scripts and software to process data is not the same as distributing arbitrary data the user agent found on the internet for which you don’t have a license.

If someone wrote an article, your reader transforms it based on your authenticated request, and your user would have an authorized subscription.

LEGAL ANALYSIS

Much of the content published on the Web isn’t secured with subscriptions and micropayments, which is why the whole thing becomes a legal battle as silly as “exceeding authorized access” which landed someone like Aaron Swartz in jail.

In other words, it is the question of “piracy”, which has acquired a new character only in that the AI is trained on your data and transforms it before it republishes it.

There was also a lawsuit aboot scraping LinkedIn, which was settled as follows: https://natlawreview.com/article/hiq-and-linkedin-reach-prop...

Legally, you can grant access to people subject to a certain license (eg Creative Commons Share Alike) and then any derived content must have its weights opened. Similar to, say, Affero GPL license for derivative software.

Post reply on HN