Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

391–400 of 555 posts

Re: Perplexity AI is lying about their user agent

#391
post #62
post #35

> Next up is some kind of GDPR request perhaps? GDPR doesn't preclude anyone from scraping you. In fact, scraping is not illegal in any context (LinkedIn keeps losing lawsuits). Using copyrighted data in training LLMs is a huge grey area, but probably not outright illegal and will take like a decade (if not more) before we'll have legislative clarity.

But per GDPR you could enforced your data fo be deleted. If enough people demand it the effort gets too high and costly

As far as I see, GDPR would not applicable here - GDPR is about control of "your data" as in "personal data about you as a private individual"[1], it is not about "your data" as in "content created or owned by you".

[1] GDPR Art 4.1 "‘personal data’ means any information relating to an identified or identifiable natural person (‘data subject’); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person;"

Re: Perplexity AI is lying about their user agent

#392
post #211

Earlier quoted context omitted.

How would an LLM training on your writing reduce your reward? I guess if you're doing it for a living sure, but most content I consume online is created without incentive (social media, blogs, stack overflow). I write a fair amount and have been for a few years. I like to play with ideas. If an llm learned from my writing and it helped me propagate my ideas, I'd be happy. I lose on social status imaginary internet po…

> I guess they do it for the glory of having their handle next to the answer? Yes, it's hardly surprising that people find upvotes and direct social rewards more exciting than being slurped somewhere into GPT-4's weights.

But they get to enjoy both the social proof on SO and GPT-4 existing.

It's not like they're getting validation from most readers anyway. People who vote and comment on answers are playing the SO social/karma game and will continue to do so whether GPT-4 exists or not. Conversely, people who'll find answers via an LLM instead of viewing it on SO are people who wouldn't bother logging in to SO, even if they had accounts on it in the first place.

People are complaining about losing the audience they never had.

Re: Perplexity AI is lying about their user agent

#393

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…

I don't know what the typical usage pattern is, but when I've used Perplexity, I generally do click the relevant links instead of just trusting Perplexity's summary. I've seen plenty of cases where Perplexity's summary says exactly the opposite of the source.

Re: Perplexity AI is lying about their user agent

#394
Lots of great arguments on this post, reasonable takes on all sides. At the end of the day though, an automated tool that identifies itself as such is “being a good citizen”, or better, “a good neighbor”. Regardless of the client or server’s notions of what constitutes bad behavior.

I haven’t heard the term “Netizen” in a while.

Re: Perplexity AI is lying about their user agent

#395
post #379

Earlier quoted context omitted.

In other words, content is bait, reward is a captured user whose attention - whose sanity, the finite amount of life - can be wasted or plain used against them. I'm more than happy to see all the websites with attention economy business models to shut down. Yes, that might be 90% of the Internet. That would be the 90% that is poisonous shit.

The attention economy will never die. Attention will only shift. From websites to aggregators like perplexity.

Perplexity isn't playing in the attention economy unless they upsell you, advertise to you, or put any other kind of bullshit between you and your goal. Attention economy is (as the name suggests) about monetizing attention; it does so through friction.

Re: Perplexity AI is lying about their user agent

#396

Earlier quoted context omitted.

> it's not scraping, it's retrieving the page on request from the user Search engines already tried it. It’s not retrieving on request because the user didn’t request the page, they requested a bot find specific content on any page.

But it's not what happened here. It WAS retrieving on request. > I went into Perplexity and asked "What's on this page rknight.me/PerplexityBot?". Immediately I could see the log and just like Lewis, the user agent didn't include their custom user agent

In this case you are 100% correct, but I think it’s reasonable to assume that the “read me this web page” use case constitutes a small minority of perplexity’s fetches. I find it useful because of the attribution - more so its references - which I almost always navigate to because its summaries are frequently crap.

Re: Perplexity AI is lying about their user agent

#397

Earlier quoted context omitted.

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

I'd believe it if they were targeting entities that could fight back, like stock photo companies and disney, instead of some guy with an artstation account, or some guy with a blog. To me it sounds like these products can't exist without exploiting someone and they're too coward to ask for permission because they know the answer is going to be "no." Imagine how many things I could create if I just stole assets from o…

...which is a great argument for abolishing copyright:P

Re: Perplexity AI is lying about their user agent

#398
"Not sure where we go from here. I don't want my posts slurped up by AI companies for free^[1] but what else can I do?"

Why not display a brief notice, like one sees on US government websites, that is impossible to miss. In this case the notice could be of the terms and conditions for using the website, in effect a brief copyright license that governs the use of material found on the website. The license could include a term prohibiting use of the material in machine learning and neural networks, including "training LLMs".

The idea is that even if these "AI" companies are complying with copyright law when using others' data for LLMs without permission, they would still be violating the license and this could be used to evade any fair use defense that the "AI" company intends to rely on.

https://www.authorsalliance.org/2023/02/23/fair-use-week-202...

Like using robots.txt, the contents of a user-agent header, if there is one, or using IP address, this costs nothing. Unlike robots.txt, User-Agent or IP addresss, it has potential legal enforceability.

That potential might be enough to deter some of these "AI" projects. You never know until you try.

Clearly, robots.txt, User-Agent header and IP address do not work.

Why would anyone aware of www history rely on the user-agent string as an accurate source of information?

As early as 1992, a year before the www went public, "user-agent spoofing" was expected.

https://raw.githubusercontent.com/alandipert/ncsa-mosaic/mas...

By 1998, webmasters who relied on user-agent strings were referred to as "ill-advised":

"Rather than using other methods of content-negotiation, some ill-advised webmasters have chosen to look at the User-Agent to decide whether the browser being used was capable of using certain features (frames, for example), and would serve up different content for browsers that identified themselves as ``Mozilla''."

"Consequently, Microsoft made their browser lie, and claim to be Mozilla, because that was the only way to let their users view many web pages in their full glory: Mozilla/2.0 (compatible; MSIE 3.02; Update a; AOL 3.0; Windows 95)"

https://www-archive.mozilla.org/build/user-agent-strings.htm...

https://webaim.org/blog/user-agent-string-history/

As for robots.txt, many sites do not even have one.

Re: Perplexity AI is lying about their user agent

#399

Earlier quoted context omitted.

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

> so far getting away with it (pending lawsuits notwithstanding). I know it feels like it's been longer, but it's not even been 2 years since ChatGPT was released. "So far" is in fact a very short amount of time in a world where important lawsuits like this can take 11 years to work their way through the courts [0]. [0] https://en.m.wikipedia.org/wiki/Oracle_v_Google

In 9 years time, robots will publish articles on the web, and they will put a humans.txt file at their root index to govern what humans are allowed to read the content.

Jokes aside, given how models become better, cheaper and smaller, RAG classification and filtering engines like Perplexity will become so ubiquitous that i don't see any way for a website owner to force anyone to visit the website anymore.

Re: Perplexity AI is lying about their user agent

#400
post #199

Earlier quoted context omitted.

But it's not scraping, it's retrieving the page on request from the user.

With no benefit provided to the creator — they're not directing users out, they're pulling data in.

They are directing users __in__ in some cases though, no? I’m a perplexity user, and their summaries are often way off which drives me to the references (attribution). The ratio of fetches to clickthroughs is what’s important now though; this new model (which we’ve not negotiated or really asked for) is driving that upward from 1, and not only are you paying more as a provider but your consumer is paying more ($ to perplexity and/or via ad backend) and you aren’t seeing any of it. And you pay those extra costs to indirectly finance the competitor who put you in this situation, who intends to drive that ratio as high as it can in order to get more money from more of your customers tomorrow. Yay.
Post reply on HN