Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

241–250 of 555 posts

Re: Perplexity AI is lying about their user agent

#241

Earlier quoted context omitted.

I think that's called a school

If you think going to school to get an education is the same thing as training an LLM then you are just so misguided. Normal people read books to gain an understanding of a concept, but do not retain the text verbatim in memory in perpetuity. This is not what training an LLM does.

LLMs don’t memorize everything they’re trained on verbatim, either. It’s all vectors behind the scenes, which is relatable to how the human brain works. It’s all just strong or weak connections in the brain.

The output is what matters. If what the LLM creates isn’t transformative, or public domain, it’s infringement. The training doesn’t produce a work in itself.

Besides that, how much original creative work do you really believe is out there? Pretty much all art (and a lot of science) is based on prior work. There are true breakthroughs, of course, but they’re few and far between.

Re: Perplexity AI is lying about their user agent

#242

Earlier quoted context omitted.

> they are decreasing the probability that this user would come to by content (via Google, for example). Google has been providing summaries of stuff and hijacking traffic for ages. I kid you not, in the tourism sector this has been a HUGE issue, we have seen 50%+ decrease in views when they started doing it. We paid gazzilions to write quality content for tourists about the most different places just so Google could…

> We paid gazzilions to write quality content for tourists about the most different places just so Google could put it on their homepage. It's just depressing It's a legitimate complaint, and it sucks for your business. But I think this demonstrates that the sort of quality content you were producing doesn't actually have much value.

I'd argue it only demonstrates that it doesn't produce much value for the creator.

Re: Perplexity AI is lying about their user agent

#243
post #183

Earlier quoted context omitted.

What will happen if: Website owners decide to stop publishing because it’s not rewarded by a real human visit anymore? Then perplexity and the like won’t have new information to train their models on and no sites to answer the questions. I think there is a real content dilemma here at work. The incentives of Google and website owners were more or less aligned. This is not the case with perplexity.

What is a "visit"? TFA demonstrates that they got a hit on their site, that's how they got the logs. Is it necessary to load the JavaScript for it to count as a visit? What if I access the site with noscript? Or is it only a visit if I see all your recommended content? I usually block those recommendations so that I don't get distracted from the article I actually came to read—is my visit a less legitimate visit than…

> TFA demonstrates that they got a hit on their site

Whats stopping perplexity caching this info say for 24 hours, and then redisplaying it to the next few hundred people who request it?

Re: Perplexity AI is lying about their user agent

#244

If you've ever tried to do any web scraping, you'll know why they lie about the User-Agent, and you'd do it too if you wanted your program to work properly. Discriminating based on User-Agent string is the unethical part.

I wouldn’t because I have ethics.

Here's my user agent on chrome:

>Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36

There are at least five lies here.

* It isn't made by Mozilla

* It doesn't use WebKit

* It doesn't use KHTML

* It isn't safari

* That isn't even my version of chrome, presumably it hides the minor/patch versions for privacy reasons.

Lying in your user agent in order to make the internet work is a practice that is almost as old as user agents. Your browser is almost certainly doing it right now to look at this comment.

Re: Perplexity AI is lying about their user agent

#245

Earlier quoted context omitted.

> they are decreasing the probability that this user would come to by content (via Google, for example). Google has been providing summaries of stuff and hijacking traffic for ages. I kid you not, in the tourism sector this has been a HUGE issue, we have seen 50%+ decrease in views when they started doing it. We paid gazzilions to write quality content for tourists about the most different places just so Google could…

> We paid gazzilions to write quality content for tourists about the most different places just so Google could put it on their homepage. It's just depressing It's a legitimate complaint, and it sucks for your business. But I think this demonstrates that the sort of quality content you were producing doesn't actually have much value.

That line of thinking makes no sense. If the "content" had no value, why would google go through the effort of scraping it and presenting it to the user?

Re: Perplexity AI is lying about their user agent

#246

Earlier quoted context omitted.

Citing the source doesn't bring you, the owner of the site, valuable data. When was your data accessed, who accessed it, from where, at what time, what device, etc. It brings data to the LLM's owner, and you get N O T H I N G. Could you change the way printed news magazines showed their content? No. Then, why is that a problem? Btw nobody clicks on sources. NOBODY.

> Btw nobody clicks on sources. NOBODY. I always click on sources to verify what an LLM in this case says. I also hear the claim that a lot about people not reading sources (before LLM it was video content with references) but I always visited the sources. Is there a statistics or studies that actually support this claim? Or is it just a personal experience, of people (including me) enforcing it as generic behavior o…

That's you, because you are a researcher or coder or someone who uses his brain much more than average, hence not an average joe. I ran a news site for 15 years and the stats showed that from 10000 views on an article, only a miniscule amount of clicks were made on the source links. Average people do not care where the info is coming from.

Also perplexity shows the videos on their site, you cannot go to youtube, you have to start it on their site, and then you have to click on the youtube player's logo in the lower right to get to the site.

Perplexity is getting greedy.

Re: Perplexity AI is lying about their user agent

#247

Earlier quoted context omitted.

You should be able to judge whether something is a copyright violation based on the resulting work. If a work was produced with or without computer assistance, why would that change whether it infringes?

As a normative claim, this is interesting, perhaps this should be the rule. As a descriptive claim, it isn't correct. Several lawsuits relating to sampling in hip-hop have hinged on whether the sounds in the recording were, in fact, sampled, or instead, recreated independently.

[deleted]

Re: Perplexity AI is lying about their user agent

#248
post #211
post #183

Earlier quoted context omitted.

What will happen if: Website owners decide to stop publishing because it’s not rewarded by a real human visit anymore? Then perplexity and the like won’t have new information to train their models on and no sites to answer the questions. I think there is a real content dilemma here at work. The incentives of Google and website owners were more or less aligned. This is not the case with perplexity.

How would an LLM training on your writing reduce your reward? I guess if you're doing it for a living sure, but most content I consume online is created without incentive (social media, blogs, stack overflow). I write a fair amount and have been for a few years. I like to play with ideas. If an llm learned from my writing and it helped me propagate my ideas, I'd be happy. I lose on social status imaginary internet po…

> The craziest one is the stack overflow contributors. They write answers for free to help people become better programmers.

In my experience they do it for points and kudos. Having people get your answers from LLMs instead of your answer on SO stops people from engaging with the gamification tools and so users get less points on the site.

Re: Perplexity AI is lying about their user agent

#249

Earlier quoted context omitted.

So if I get access to the Perplexity AI source code (I borrow it from a friend), read all of it, and reproduce it at some level, then Perplexity will be:" sure, that's fine no harm, no IP theft, no copyright violation, because you read it so we're good"? No, they would sue me for everything I got, and then some. That's the weird thing about these companies, they are never afraid to use IP law to go after others, but…

If Perplexity’s source code is downloaded from a public web site or other repository, and you take the time to understand the code and produce your own novel implementation, then yes. Now, if you “get it from a friend”, illegally, _or_ you just redeploy the code, without creating a transformative work, then there’s a problem. > Just pay the stupid license and if that makes your business unsustainable then it's not mu…

> They’d be fools to buy licenses before it’s been decided.

They are willingly ignoring licenses until someone sues them? That's still illegal and completely immoral. There is tons of data to train on. The entirety of Wikipedia, all of StackOverflow (at least previously), all of the BSD and MIT licenses source code on Github, the entire Gutenberg project. So much stuff, freely and legally available, yet their feel that they don't need to check licenses?

Re: Perplexity AI is lying about their user agent

#250
post #209

Earlier quoted context omitted.

Well for one thing you visiting his site and displaying it via reader mode doesn't remove his ability to sell paid licenses for his content to companies that would like to redistribute his content. Meanwhile having those companies do so for free without a license obviously does.

Should OP be allowed to demand a license for redistribution from Orion Browser [0]? They make money selling a browser with a built-in ad blocker. Is that substantially different than what Perplexity is doing here? [0] https://kagi.com/orion/

Orion browser presuming it does what does what it's name says it does doesn't redistribute anything... so presumably not.
Post reply on HN