I am not sure I will ever stop being weirded out, annoyed at, confused by, something... people asking these sorts of questions of an LLM. What, you want an apology out of the LLM?
I don't get it either. How is the LLM meant to know the details of how the perplexity headless browser works?
Perplexity AI is lying about their user agent
191–200 of 555 posts
Re: Perplexity AI is lying about their user agent
#192What incentive does anybody have to be honest about their user agent?
It's useful in the few cases where UAs support different features in ways that the standard feature-detection APIs can't detect. I think that's supposed to be fairly rare these days.
Instead, today there are different sets of features supported by engines with the same user agent.
Re: Perplexity AI is lying about their user agent
#193Earlier quoted context omitted.
I think that's called a school
If you think going to school to get an education is the same thing as training an LLM then you are just so misguided. Normal people read books to gain an understanding of a concept, but do not retain the text verbatim in memory in perpetuity. This is not what training an LLM does.
LLMs wouldn't hallucinate so much if they did that, either.
Re: Perplexity AI is lying about their user agent
#194There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…
N O T H I N G.
Could you change the way printed news magazines showed their content? No. Then, why is that a problem?
Btw nobody clicks on sources. NOBODY.
Re: Perplexity AI is lying about their user agent
#195There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…
It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…
Perplexity has source references. I find myself visiting the source references. Especially to validate the LLM output. And to learn more about the subject. Perplexity uses a Google search API to generate the reference links. I think a better strategy is to treat this as a new channel to receive visitors.
The browsing experience should be improved. Mozilla had a pilot called Context Graph. Perhaps Context Graph should be revisited?
> In a world where users ask Perplexity these Help questions about my SaaS, Perplexity may answer them and I would lose all the insights because I never get any traffic.
This seems like a missing feature for analytics products & the LLMs/RAGs. I don't think searching via an LLM/RAG is going away. It's too effective for the end user. We have to learn to work with it the best we can.
Re: Perplexity AI is lying about their user agent
#196Earlier quoted context omitted.
I think that’s the ideal as the server may provide different data depending on UA. Does anyone actually do this, though?
I fake my UA the way I like.
But my question should have been phrased, “are there any frameworks commonly in use these days that provide different js payloads to different clients?
I’ve been out of that part of the biz for a very long time so this could be a naive question.
Re: Perplexity AI is lying about their user agent
#197Earlier quoted context omitted.
> The second concern, though, is can perplexity do a live web query to my website and present data from my website in a format that the user asks for? Arguing that we should ban this moves into very dangerous territory. This feels like the fundamental core component of what copyright allows you to forbid. > Everything from ad blockers to reader mode to screen readers do exactly the same thing that Perplexity is doing…
I'm curious to know where you draw the line for what constitutes legitimate manipulation by a person and when it becomes distribution. I'm assuming that if I write code by hand for every part of the TCP/IP and HTTP stack I'm safe. What if I use libraries written by other people for the TCP/IP and HTTP part? What if I use a whole FOSS web browser? What about a paid local web browser? What if I run a script that I wrot…
> What if I decide to offer that script as a service for free to friends and family, who can use my cloud server?
You're allowed to make copies and adaptations in order to utilize the program (website), which probably covers a cloud server you yourself are controlling. You aren't allowed to do other things with those copies though, like distribute them to other people.
Payment only matters if we're getting into "free use" arguments, and I don't think any really apply here.
I think you're probably already in trouble with just offering it to family and friends, but if you take the next step offering it to the public that adds more issues because the copyright act includes definitions like "To perform or display a work “publicly” means (1) to perform or display it at a place open to the public or at any place where a substantial number of persons outside of a normal circle of a family and its social acquaintances is gathered; or (2) to transmit or otherwise communicate a performance or display of the work to a place specified by clause (1) or to the public, by means of any device or process, whether the members of the public capable of receiving the performance or display receive it in the same place or in separate places and at the same time or at different times."
Re: Perplexity AI is lying about their user agent
#198Earlier quoted context omitted.
If using copyrighted material to train an LLM is theft, so is reading a book.
Computers are not people. Laws differ and consequences can be different based on the actor (like how minors are treated differently in courts). Just because a person can do it does not automatically mean those same rights transfer to arbitrary machines.
Re: Perplexity AI is lying about their user agent
#199Earlier quoted context omitted.
How a user views a page isn't the same as a startup scraping the internet wholesale for financial gain.
But it's not scraping, it's retrieving the page on request from the user.
Re: Perplexity AI is lying about their user agent
#200There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…
It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…
If I visit your site from Google with my browser configured to go straight to Reader Mode whenever possible, is my visit more useful to you than a summary and a link to your site provided by Perplexity? Why does it matter so much that visitors be directly on your content?