Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

481–490 of 555 posts

Re: Perplexity AI is lying about their user agent

#481
post #113
post #3

I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.

I agree with that, but I also think that they should at least identify themselves instead of using a generic user agent.

I want all people in the world with a dirty arse to change their user agent so i can not serve my website to dirty arses.

Re: Perplexity AI is lying about their user agent

#482

Earlier quoted context omitted.

> humans learning is the old-school digital copying. computers simply do it much faster, but it's the same basic phenomenon This is dangerous framing because it papers over the significant material differences between AI training and human learning and the outcomes they lead to. We all have a collective interest in the well-being of humanity, and human learning is the engine of our prosperity. Each individual has age…

you're reading what I say in the worst possible light if anything, the parallel I draw between AI learning and humans learning is all the opposite of narrow and logical... in my intent, the analogy is loose and poetic, not mechanistic and exact. AI are tools, if AI are enslaving is because there are human actors (I hope....) deciding to enslave other humans, not because of anything inherent to training (if AI; learni…

Your response is fair and I hope you didn't take my message personally. I agree with you, AI is just a tool same as countless others that can be used for good or evil.

Re: Perplexity AI is lying about their user agent

#483

Earlier quoted context omitted.

It's scraping content to then serve up that content to users who can now get that content from you (via a paid subscription service, or maybe ad-sponsored) instead of visiting the content creator and paying them (i.e., via ads on their website) It's the same reason I can't just take NYT archives or the Britannica and sell an app that gives people access to their content through my app. It totally undercuts content cr…

One more point on this, lest some people think, "hey Kanye, or Taylor Swift, don't need any more money!" I 100% agree. But the problem with streaming is that is disproportionately rewards the biggest artists at the expense of the smaller ones. It's the small artist, barely making a living from their craft, who were most hurt by the switch from albums to streaming, not those making millions.

As a musician, Spotify is the best thing to happen to musicians. Imagine trying to distribute your shit via burned CDs you made yourself. The entitlement of thinking "I have a garage band and Spotify isn't paying me enough" is fucking ridiculous. 99.99% of bands have never made it. The ability to easily distribute your music worldwide is crazy. If people don't like it, you're either bad at marketing, or, more likely, your music is average at best. It's a big world.

Re: Perplexity AI is lying about their user agent

#484

Earlier quoted context omitted.

We would employ local guides all around the world to craft itinerary plans to visit places, give tips, tricks, recommend experiences and places (we made money by selling some of those through our website) and it was a success. Customers liked the in depth value of that content and it converted to buys (we sold experiences and other stuff, sort of like getyourguide). One day all of our content ended up on Google "what…

I totally get that it killed your traffic. If a thousand people a day typing in "what time is best to visit the Sagrada Familiar" stopped clicking on the link to your page because Google just told them "4 PM on Thursdays" at the top of the page, you lost a bunch of traffic. But why did you want the traffic? Was your revenue from ad impressions, or were you perhaps being paid by the city of Barcelona to provide useful…

Moreover, if it's the former, then good riddance. An ad-backed site is harming users a little on the margin for the marginal piece of information. Getting the same from a search engine is saving users from that harm.

Parent has the right question here: why did you want the traffic? Did you intend for anything good to happen to those people?. I'm going to guess not; there's hardly a scenario where people who complain about loss traffic and mean that traffic any good.

Re: Perplexity AI is lying about their user agent

#485
post #360

Earlier quoted context omitted.

Why are you ignoring his main argument?

I'm not. I'm asking why this flow is "distribution": * User types an address into Perplexity * Perplexity fetches the page, transforms it, and renders some part of it for the user But this flow is not: * User types an address into Orion Browser * Orion Browser fetches the page, transforms it, and renders some part of it for the user Regardless of the legal question (which I'm also skeptical of), I'm especially unconv…

The moral case is pretty obviously that Perplexity is preventing traffic from reaching the people who made the content.

Re: Perplexity AI is lying about their user agent

#486
post #355

Earlier quoted context omitted.

"Let us steal your content or you won't get any traffic" sounds extortionate

It is what it is. AI is increasingly being used to make lives easier. Those who choose to isolate from AI choose to isolate from the many using it.

We're burning long term value and the open web for shitty chat bots.

Re: Perplexity AI is lying about their user agent

#487

Earlier quoted context omitted.

The Google summaries (before whatever LLM stuff they're doing now) are 2-3 sentences tops. The content on most of these websites is much, much longer than that for SEO reasons. It sucks that Google created the problem on both ends, but the content OP is referring to costs way more to produce than it adds value to the world because it has to be padded out to show up in search. Then Google comes along and extracts the…

It would be a lot better if Google just prioritised concise websites. If Google preferred websites that cut the fluff, then website operators would have an incentive to make useful websites, and Google wouldn't have as much of an incentive to provide the answer in a snippet, and everyone wins. I guess it's hard to rank website quality, so Google just prefers verbose websites.

> Google wouldn't have as much of an incentive to provide the answer in a snippet, and everyone wins.

Google has at least two incentives to provide that answer, both of which wouldn't change. The bad one: they want to keep you on their page too, for usual bullshit attention economy reasons. The good one: users prefer the snippets too.

The user searching for information usually isn't there to marvel at beauty of random websites hiding that information in piles of noise surrounded by ads. They don't care about websites in the first place. They want an answer to the question, so they can get on with whatever it is they're doing. When Google can give them an answer, and this stops them from going from SERP to any website, then that's just few seconds or minutes of life that user doesn't have to waste. Lifespans are finite.

Re: Perplexity AI is lying about their user agent

#488

Earlier quoted context omitted.

It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…

I'm not sure what you mean exactly. If Perplexity is actually doing something with your article in-band (e.g. downloading it, processing it, and present that processed article to the user) then they're just breaking the law. I've never used that tool (and don't plan to) so I don't know. If they just embed the content in an iframe or something then there's no issue (but then there's no need or point in scraping). If t…

> If they're just scraping to train then I think you also imply there's no issue. If they're just copying your content (even if the prompt is "Hey Perplexity, summarise this article ") then that's vanilla infringement, whether they lie about their UA or not.

Except, it can't possibly be like that - that would kill the Internet as you know it. It makes sense to consider scrapping for purposes of training as infringement - I personally disagree, I'm totally on the side of AI companies on this one, but there's a reasonable argument there. But in terms of me requesting a summary, and the AI tool doing it server-side before sending it to me, without also adding it to the pile of its own training data? Banning that would mean banning all user-generated content websites, all web viewing or editing tools, web preview tools, optimizing proxies, malware scanners, corporate proxies, hell, maybe even desktop viewers and editing tools.

There are always multiple programs between your website and your user's eyeballs. Most of them do some transformations. Most of them are third-party, usually commercial software. That's how everything works. Software made by "AI company" isn't special here. Trying to make it otherwise is some really weird form of prejudice-driven discrimination.

Re: Perplexity AI is lying about their user agent

#489

Earlier quoted context omitted.

The "scumbag AI company" in question is making money by offering me a way to access information while skipping any and all attention economy bullshit you may have on your site, on top of being just plain more convenient. Note that the author is confusing crawling (which is done with documented User Agent and presumably obeys robots.txt) with browsing (which is done by working as one-off user agent for the user). As f…

If you want summaries from my website, go to my website. I want a way to deny any licence to any third-party user agent that will apply machine learning on my content, whether you initiated the request or not. LLMs — and more importantly the companies that train and operate them — should not be trusted at all, especially for so-called "summarization": https://ea.rna.nl/2024/05/27/when-chatgpt-summarises-it-actu... Wh…

> If you want summaries from my website, go to my website.

I will. Through Perplexity. My lifespan is limited, and I have better ways to spend it than digging out information while you make a buck from making me miserable (otherwise there isn't much reason to complain, other than some anti-AI ideology stance).

> I want a way to deny any licence to any third-party user agent that will apply machine learning on my content, whether you initiated the request or not.

That's not how the Internet works. Allowing for that would mean killing user-generated content sites, optimizing proxies, corporate proxies, online viewers and editors, caches, possibly desktop software too.

Also, my browser probably already does some ML on the side anyway. You'd catch a lot of regular browsing this way.

Ultimately, the rules of the road are what they always have been: whatever your publicly accessible web server spouts out on a request is fair game for the requester to consume however they like, in part or entirely. If you want to limit access for particular tools or people, put up a goddamn paywall. All the noise about scrapping and stuff is attention economy players trying to have their cake and eat it too. As the user in - i.e. the victim of - attention economy, I don't feel much sympathy for that plight.

Also:

> LLMs — and more importantly the companies that train and operate them — should not be trusted at all, especially for so-called "summarization"

That's not your problem. That's my problem. If I use a shitty tool from questionable vendor to parse your content, that's on me. You should not care. In fact, being too interested in what I use for my Internet consumption can be seen as surveillance, which is not nice.

Re: Perplexity AI is lying about their user agent

#490
post #10

What incentive does anybody have to be honest about their user agent?

It's good etiquette, for one, and encouraging good etiquette (both on the parts of website operators and website requestors) is a good thing.

As a website operator, I've actually increased ratelimits for a service I ran , from a particular crawler, that's normally much more stringent just because it was the easiest way to identify the people crawling and I liked what they were doing.

I know some web services effectively require you not to lie about your user agent (this applies more to APIs, but they'll block or severely ratelimit user agents that are browser-like or are generic "requests" or what have you).

Post reply on HN