Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

251–260 of 555 posts

Re: Perplexity AI is lying about their user agent

#251
post #229

Earlier quoted context omitted.

> A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content. If I visit your site from Google with my browser configured to go straight to Reader Mode whenever possible, is my visit more useful to you than a summary and a link to your site provid…

Traffic numbers, regardless if it using reader mode or not, are used as a basic valuation of a website or page. This is why Alexa rankings have historically been so important. If Perplexity visit the site once and cache some info to give to multiple users, that is stealing traffic numbers for ad value, but also taking away the ability from the site owner to get realistic ideas of how many people are using the informa…

>Traffic numbers, regardless if it using reader mode or not, is used as a basic valuation of a website.

I have another comment that says something similar, but: is valuing a website based on basic traffic still a thing? Feels very 2002. It's not my wheelhouse, but if I happened to be involved in a transaction, raw traffic numbers wouldn't hold much sway.

Re: Perplexity AI is lying about their user agent

#252

Earlier quoted context omitted.

> We paid gazzilions to write quality content for tourists about the most different places just so Google could put it on their homepage. It's just depressing It's a legitimate complaint, and it sucks for your business. But I think this demonstrates that the sort of quality content you were producing doesn't actually have much value.

I'd argue it only demonstrates that it doesn't produce much value for the creator.

The Google summaries (before whatever LLM stuff they're doing now) are 2-3 sentences tops. The content on most of these websites is much, much longer than that for SEO reasons.

It sucks that Google created the problem on both ends, but the content OP is referring to costs way more to produce than it adds value to the world because it has to be padded out to show up in search. Then Google comes along and extracts the actual answer that the page is built around and the user skips both the padding and the site as a whole.

Google is terrible, the attention economy that Google created is terrible. This was all true before LLMs and tools like Perplexity are a reaction to the terrible content world that Google created.

Re: Perplexity AI is lying about their user agent

#253
post #211
post #183

Earlier quoted context omitted.

What will happen if: Website owners decide to stop publishing because it’s not rewarded by a real human visit anymore? Then perplexity and the like won’t have new information to train their models on and no sites to answer the questions. I think there is a real content dilemma here at work. The incentives of Google and website owners were more or less aligned. This is not the case with perplexity.

How would an LLM training on your writing reduce your reward? I guess if you're doing it for a living sure, but most content I consume online is created without incentive (social media, blogs, stack overflow). I write a fair amount and have been for a few years. I like to play with ideas. If an llm learned from my writing and it helped me propagate my ideas, I'd be happy. I lose on social status imaginary internet po…

I think a concern for people who contribute on Stack Overflow is that an LLM will pollute the water with so many subtly wrong answers that the collective work of answering questions accurately will be overwhelmed by a tsunami of inaccurate LLM-generated answers, more than an army of humans can keep up with checking and debugging (or debunking).

Re: Perplexity AI is lying about their user agent

#254
post #50

I am not sure I will ever stop being weirded out, annoyed at, confused by, something... people asking these sorts of questions of an LLM. What, you want an apology out of the LLM?

That's an interesting point you're making. I wonder what the policy is regarding the questions people ask an LLM and the developers behind the service reading through the questions (with unsuccessful responses from the LLM?)

Re: Perplexity AI is lying about their user agent

#256
post #3

I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.

If an AI agent is performing a search on behalf of a user, should its user agent be the same as that user’s?

It should, erm sorry, must pass all the info it got from the user to you, so you would have an idea who wanted info from your site.

Re: Perplexity AI is lying about their user agent

#257

Earlier quoted context omitted.

> We paid gazzilions to write quality content for tourists about the most different places just so Google could put it on their homepage. It's just depressing It's a legitimate complaint, and it sucks for your business. But I think this demonstrates that the sort of quality content you were producing doesn't actually have much value.

That line of thinking makes no sense. If the "content" had no value, why would google go through the effort of scraping it and presenting it to the user?

>If the "content" had no value, why would google go through the effort of scraping it and presenting it to the user?

They don't present it all, they summarize it.

And let's be serious here, I was being polite because I don't know the OPs business. But 99% of this sort of content is SEO trash and contributes to the wasteland that the internet is becoming. Feel free to point me to the good stuff.

Re: Perplexity AI is lying about their user agent

#258

Respecting robots.txt is something their training crawler should do, and I see no reason why their user agent (i.e. user asks it to retrieve a web page, it does) should, as it isn't a crawler (doesn't walk the graph). As to "lying" about their user agents - this is 2024, the "User-Agent" header is considered a combination bug and privacy issue, all major browsers lie about being a browser that was popular many years…

They might be "lying" because of all sorts of reasons, but a specific version of Chrome on a specific OS still sends a unique user agent string.

Re: Perplexity AI is lying about their user agent

#259
post #229

Earlier quoted context omitted.

> A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content. If I visit your site from Google with my browser configured to go straight to Reader Mode whenever possible, is my visit more useful to you than a summary and a link to your site provid…

Traffic numbers, regardless if it using reader mode or not, are used as a basic valuation of a website or page. This is why Alexa rankings have historically been so important. If Perplexity visit the site once and cache some info to give to multiple users, that is stealing traffic numbers for ad value, but also taking away the ability from the site owner to get realistic ideas of how many people are using the informa…

> The only way to confirm that, or to get the correct information in the first place, is to read the original site yourself.

As someone who uses Perplexity, I often do do this. And I don't think I'm particularly in the minority with this. I think their UI encourages it.

Re: Perplexity AI is lying about their user agent

#260
post #250

Earlier quoted context omitted.

Should OP be allowed to demand a license for redistribution from Orion Browser [0]? They make money selling a browser with a built-in ad blocker. Is that substantially different than what Perplexity is doing here? [0] https://kagi.com/orion/

Orion browser presuming it does what does what it's name says it does doesn't redistribute anything... so presumably not.

I asked you this in the other subthread, but what exactly is the moral distinction (I'm not especially interested in the legal one here because our copyright law is horribly broken) between these two scenarios?

* User asks proprietary web browser to fetch content and render it a specific way, which it does

* User asks proprietary web service to fetch content and render it a specific way, which it does

The technical distinction is that there's a network involved in the second scenario. What is the moral distinction?

Post reply on HN