Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

521–530 of 555 posts

Re: Perplexity AI is lying about their user agent

#521

Earlier quoted context omitted.

That's you, because you are a researcher or coder or someone who uses his brain much more than average, hence not an average joe. I ran a news site for 15 years and the stats showed that from 10000 views on an article, only a miniscule amount of clicks were made on the source links. Average people do not care where the info is coming from. Also perplexity shows the videos on their site, you cannot go to youtube, you…

You said "NOBODY" (pretty sure the all caps means it's extra true).

Well, that's the reality. The tech savvy people here are the exception, and represent only a very minor percentage of users.

Re: Perplexity AI is lying about their user agent

#522

Earlier quoted context omitted.

I totally get that it killed your traffic. If a thousand people a day typing in "what time is best to visit the Sagrada Familiar" stopped clicking on the link to your page because Google just told them "4 PM on Thursdays" at the top of the page, you lost a bunch of traffic. But why did you want the traffic? Was your revenue from ad impressions, or were you perhaps being paid by the city of Barcelona to provide useful…

Moreover, if it's the former, then good riddance . An ad-backed site is harming users a little on the margin for the marginal piece of information. Getting the same from a search engine is saving users from that harm. Parent has the right question here: why did you want the traffic? Did you intend for anything good to happen to those people? . I'm going to guess not; there's hardly a scenario where people who complai…

Google Search is ad-backed site. Especially for highly commercial queries.

They just prefer internet users to consume their ads, rather than the ads of the content creators.

Re: Perplexity AI is lying about their user agent

#523

Earlier quoted context omitted.

One more point on this, lest some people think, "hey Kanye, or Taylor Swift, don't need any more money!" I 100% agree. But the problem with streaming is that is disproportionately rewards the biggest artists at the expense of the smaller ones. It's the small artist, barely making a living from their craft, who were most hurt by the switch from albums to streaming, not those making millions.

As a musician, Spotify is the best thing to happen to musicians. Imagine trying to distribute your shit via burned CDs you made yourself. The entitlement of thinking "I have a garage band and Spotify isn't paying me enough" is fucking ridiculous. 99.99% of bands have never made it. The ability to easily distribute your music worldwide is crazy. If people don't like it, you're either bad at marketing, or, more likely,…

Read up on how Spotify remunerates artists.

Re: Perplexity AI is lying about their user agent

#524

Earlier quoted context omitted.

It's scraping content to then serve up that content to users who can now get that content from you (via a paid subscription service, or maybe ad-sponsored) instead of visiting the content creator and paying them (i.e., via ads on their website) It's the same reason I can't just take NYT archives or the Britannica and sell an app that gives people access to their content through my app. It totally undercuts content cr…

Serve it in a better way or wall it. The Internet is supposed to be free. If you don't want unauthorized eyes to see it, you have the ability to hide it behind logins.

Free to access != free to copy and redistribute for profit

Re: Perplexity AI is lying about their user agent

#525

Earlier quoted context omitted.

It's scraping content to then serve up that content to users who can now get that content from you (via a paid subscription service, or maybe ad-sponsored) instead of visiting the content creator and paying them (i.e., via ads on their website) It's the same reason I can't just take NYT archives or the Britannica and sell an app that gives people access to their content through my app. It totally undercuts content cr…

Serve it in a better way or wall it. The Internet is supposed to be free. If you don't want unauthorized eyes to see it, you have the ability to hide it behind logins.

This will further push websites to paywalls making the internet less feee.

Re: Perplexity AI is lying about their user agent

#526

Earlier quoted context omitted.

As a musician, Spotify is the best thing to happen to musicians. Imagine trying to distribute your shit via burned CDs you made yourself. The entitlement of thinking "I have a garage band and Spotify isn't paying me enough" is fucking ridiculous. 99.99% of bands have never made it. The ability to easily distribute your music worldwide is crazy. If people don't like it, you're either bad at marketing, or, more likely,…

Read up on how Spotify remunerates artists.

I have multiple Spotify artists. I get it and think it's a fantastic service. Anyone complaining about it probably gets a couple dozen monthly plays because they don't know how to market, gig, and tour, or more likely their music sucks.

Re: Perplexity AI is lying about their user agent

#527

Earlier quoted context omitted.

1) If your blog posts are private, why are they on publicly accessible websites? Why not put it behind a paywall of some sort? 2) How many novels have bibliographies? How many musicians cite their influences? Citing sources is all well and good in academic papers, but there’s a point at which it just becomes infeasible. The more transformative the work, the harder it is to cite inspiration. 3) What about libraries? S…

> 1) If your blog posts are private, why are they on publicly accessible websites? Why not put it behind a paywall of some sort? If I grow apple trees in front of my house and you come and take all apples and then turn up at my doorstep trying to sell me apple juice made from the apples you nicked that doesn't mean you had the right to do it, because I chose not to build a tall fence around my apple trees. Public con…

All of these responses were so quality, there's really no need to add. I Especially like the apple argument about a product in your front yard. You still have no basis to take them from my front yard.

If there was the equivalent of what a lot of other sites have (gems, gold, ribbons) I'd give you one. Got a lot of gems, I'll send you an admittedly teeny heliodore, tourmaline, or peridot at cost if you want one. Gemstone market's junk lately with the economy.

Re: Perplexity AI is lying about their user agent

#528

Earlier quoted context omitted.

If you want summaries from my website, go to my website. I want a way to deny any licence to any third-party user agent that will apply machine learning on my content, whether you initiated the request or not. LLMs — and more importantly the companies that train and operate them — should not be trusted at all, especially for so-called "summarization": https://ea.rna.nl/2024/05/27/when-chatgpt-summarises-it-actu... Wh…

> If you want summaries from my website, go to my website. I will. Through Perplexity. My lifespan is limited, and I have better ways to spend it than digging out information while you make a buck from making me miserable (otherwise there isn't much reason to complain, other than some anti-AI ideology stance). > I want a way to deny any licence to any third-party user agent that will apply machine learning on my cont…

I addressed this in a different response: I do not care if your browser does local ML or if there is an extension which takes content that you have already downloaded and applies ML on it (as long as the results of the ML on my licensed content are not stored in third party services without respecting my licence). I do care that an agent controlled by a third party (even if it is on your behalf) browses instead of you browsing.

My goal is to licence my content for first party use, not third party derived use.

Your statement "Ultimately, the rules of the road are what they always have been: whatever your publicly accessible web server spouts out on a request is fair game for the requester to consume however they like" is both logically and legally incorrect in pretty much every single jurisdiction in the world, even if it cannot be controlled as such without expensive legal proceedings.

> > LLMs — and more importantly the companies that train and operate them — should not be trusted at all, especially for so-called "summarization" > That's not your problem. That's my problem. If I use a shitty tool from questionable vendor to parse your content, that's on me. You should not care. In fact, being too interested in what I use for my Internet consumption can be seen as surveillance, which is not nice.

Actually, it is my problem, because it's my words that have been badly summarized.

If the LLM provides a so-called summary that is the exact opposite of what I wrote (as the link I shared previously shows happens), and that summary is then used to write something about what I supposedly wrote, then I have been misrepresented at best.

I have a moral right to work that I have created (under Canadian law and most European laws) to ensure that my work is not misrepresented. The best way that I can do that is to forbid its consumption by machine learning companies, including Perplexity.

> The moral rights include the right of attribution, the right to have a work published anonymously or pseudonymously, and the right to the integrity of the work. The preserving of the integrity of the work allows the author to object to alteration, distortion, or mutilation of the work that is "prejudicial to the author's honor or reputation". Anything else that may detract from the artist's relationship with the work even after it leaves the artist's possession or ownership may bring these moral rights into play. Moral rights are distinct from any economic rights tied to copyrights. Even if an artist has assigned his or her copyright rights to a work to a third party, he or she still maintains the moral rights to the work.

https://en.wikipedia.org/wiki/Moral_rights

Of course, Perplexity operates under the Wild West of copyright law where they and their users truly do not give one whit about the damage they cause. Eventually, this will be their downfall, because they are going to find themselves on the wrong side of legal judgements for their unwillingness to play by rules that have been in place for a fairly long time.

Re: Perplexity AI is lying about their user agent

#529

Earlier quoted context omitted.

You are definitionally incorrect. From Wikipedia: > robots.txt is the filename used for implementing the Robots Exclusion Protocol, a standard used by websites to indicate to visiting web crawlers and other web robots which portions of the website they are allowed to visit. From robotstxt.org/orig.html (the original proposed specification), there is a bit about "recursive" behaviour, but the last paragraph indicates…

Every paragraph that you've included up there just reinforces my point. The recursive behavior isn't incidental, it's literally part of the definition of a crawler. You can't just skip past that and pretend that the people who specifically included the word recursive (or the phrase "many pages") didn't really mean it. The first paragraph of the two about access controls is the context for what "should not be accessed…

Perplexity's ad hoc requests are still made by a crawler — whether you believe it or not. A web browser presents the content directly to the user. There may be extensions or features (reader mode) which modify the retrieved content in browser, but Perplexity's summarization feature does not present the content directly to the user in any way.

It honestly just feels like you have no critical thinking when it comes to LLM tech and want to pretend that an autonomous crawler that only retrieves a single page to process it isn't a crawler.

I have used, with permission of the site owner, a crawler to retrieve data from a single URL on a scheduled basis. It is fully automated data retrieval not intended for direct user consumption. THAT is what makes it a crawler. If the page from which I was retrieving the data was included in `/robots.txt`, the site owner would expect that an automated program would not pull the data. Recursiveness is not required to make a web robot. Unattended and/or disconnected requests do.

Re: Perplexity AI is lying about their user agent

#530

Earlier quoted context omitted.

I see no competition. I use Perplexity regularly to give me summaries of articles or to do preliminary research. If I like what I'm seeing, then I go to the source. If a source chooses to block their content because they don't want it to be accessed by AI bots then they reduce even further the chance of me - and increasingly more persons - touching their site at all.

You can say that, it doesn't matter. The statistics show that these tools reduce views. And really, "I'm going to replace my entire news intake with the AI slop even if it's entirely hallucinated lies or propaganda" is perhaps not something you ought to say out loud.

Reality is view stats, etc don't matter to most users. All we want is to read/research in peace.

You're missing something crucial. Yes, there may be hallucinations, but that's why there's preference for AI that provides citations; they can be easily checked. And also, in my experience, the summaries are usually decent; the sources which tend to yield pretty broken summaries also tend to intersperse unrelated material (ads, previews to other things, etc) in the content, or they're doing wild transforms with JS for example to create unnecessary eye candy. And coincidentally, I'd rather avoid those sources where possible, so it becomes an even greater win for me as user to have AI prefilter said content. Thus the sources getting those precious views in the end become the ones that respect their users' time and aversion to distractions.

Post reply on HN