Some people are blind, others have physical disabilities, some of us have astigmatisms or ADHD and can’t use badly designed ad-laden websites.
Perplexity AI is lying about their user agent
421–430 of 555 posts
Re: Perplexity AI is lying about their user agent
#422There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…
Re: Perplexity AI is lying about their user agent
#423Earlier quoted context omitted.
> What exactly is Perplexity doing here that isn't okay that people don't already do with their local user agents? It's in the title of TFA: they're being dishonest about who they are. PerplexityBot seems to understand that robots.txt is addressed to it . It's understood that site operators have a right to use the User-Agent to discriminate among visitors; that's why robots.txt is a standard. Crawlers that disrespect…
> It's in the title of TFA: they're being dishonest about who they are. PerplexityBot seems to understand that robots.txt is addressed to it. First, I'm ignoring the output of Perplexity. I have no reason to believe that they gave the LLM any knowledge about its internal operations, it's just riffing off of what OP is saying. Second, PerplexityBot is the user agent that they use when crawling and indexing. They never…
Is that a distinction without a difference?
I think the robots.txt RFC was addressed specifically to crawlers; so technically "ad hoc" requests generated automatically (i.e. by robots) aren't included. But the distinction operators would like to make is between humans and automata. Whether some automaton is a crawler or not isn't relevant.
Re: Perplexity AI is lying about their user agent
#424Re: Perplexity AI is lying about their user agent
#425Earlier quoted context omitted.
I'm having difficulty grasping the concept. Only a fool would trust any HTTP headers such as User-Agent sent by a random unauthenticated client. Your expenses are your problem.
… and I have absolutely no obligation to provide any particular response to any particular client. Parsing, rendering, and trusting that the payload is consistent from request to request is your problem . You can connect to my server, or not. I really don’t care. What you cannot do is dictate how my server responds to your request.
The client is under no obligation to be truthful in its communications with a server. Spoofing a User-Agent doesn't "dictate" anything. Your server dictates how it responds all on its own when it discriminates against some User-Agents.
Re: Perplexity AI is lying about their user agent
#426Earlier quoted context omitted.
> Btw nobody clicks on sources. NOBODY. I always click on sources to verify what an LLM in this case says. I also hear the claim that a lot about people not reading sources (before LLM it was video content with references) but I always visited the sources. Is there a statistics or studies that actually support this claim? Or is it just a personal experience, of people (including me) enforcing it as generic behavior o…
That's you, because you are a researcher or coder or someone who uses his brain much more than average, hence not an average joe. I ran a news site for 15 years and the stats showed that from 10000 views on an article, only a miniscule amount of clicks were made on the source links. Average people do not care where the info is coming from. Also perplexity shows the videos on their site, you cannot go to youtube, you…
Re: Perplexity AI is lying about their user agent
#427Earlier quoted context omitted.
But it's not scraping, it's retrieving the page on request from the user.
> it's not scraping, it's retrieving the page on request from the user Search engines already tried it. It’s not retrieving on request because the user didn’t request the page, they requested a bot find specific content on any page.
Re: Perplexity AI is lying about their user agent
#428Earlier quoted context omitted.
Here is a simple example. If you made your website only work in say, Microsoft Edge, and blocked everyone else telling them to download Edge. I'd think you're an asshole. Whether or not being an ass is unethical I'll leave to the philosophers. Clearly there are many other scenarios, and many that are more muddy, but overall when we get in to the business of trying to force people to consume content in particular ways…
The entire premise of the parent posters comment was that this is specifically unethical , so you lost me at the part where you deliberately decided to not address that in your reply.
?
Re: Perplexity AI is lying about their user agent
#429Earlier quoted context omitted.
Traffic numbers, regardless if it using reader mode or not, are used as a basic valuation of a website or page. This is why Alexa rankings have historically been so important. If Perplexity visit the site once and cache some info to give to multiple users, that is stealing traffic numbers for ad value, but also taking away the ability from the site owner to get realistic ideas of how many people are using the informa…
> Traffic numbers, regardless if it using reader mode or not, is used as a basic valuation of a website. I have another comment that says something similar, but: is valuing a website based on basic traffic still a thing? Feels very 2002. It's not my wheelhouse, but if I happened to be involved in a transaction, raw traffic numbers wouldn't hold much sway.
Re: Perplexity AI is lying about their user agent
#430Earlier quoted context omitted.
I'm curious about the tourism sector problem. In tourism, I would think the goal would be to promote a location. You want people to be able to easily discover the location, get information about it, and presumably arrange to travel to those locations. If Google gets the information to the users, but doesn't send the tourist to the website, is that harmful? Is it a problem of ads on the tourism website? Or is more of…
We would employ local guides all around the world to craft itinerary plans to visit places, give tips, tricks, recommend experiences and places (we made money by selling some of those through our website) and it was a success. Customers liked the in depth value of that content and it converted to buys (we sold experiences and other stuff, sort of like getyourguide). One day all of our content ended up on Google "what…
There was never a guarantee -- for anyone in any industry at all -- that what worked in the past will always continue to work. That is a regressive attitude.
However I do have concerns about Google and other monopolies replacing large swaths of people who make their livings doing things that can now be automated. I am not against automation but I don't think the disruption of our entire societal structure and economy should be in the hands of the sociopaths that run these companies. I expect regulation to come into play once the shit hits the fan for more people.