Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

111–120 of 555 posts

Re: Perplexity AI is lying about their user agent

#111
post #84

The author has misunderstood when the perplexity user agent applies. Web site owners shouldn’t dictate what browser users can access their site with - whether that’s chrome, firefox, or something totally different like perplexity. When retrieving a web page _for the user_ it’s appropriate to use a UA string that looks like a browser client. If perplexity is collecting training data in bulk without using their UA that…

It’s not retrieving a web page though is it? It’s retrieving the content then manipulating it. Perplexity isn’t a web browser.

> It’s retrieving the content then manipulating it. Perplexity isn’t a web browser.

So a browser with an ad-blocker that's removing / manipulating elements on the page isn't a browser? What about reader mode?

Re: Perplexity AI is lying about their user agent

#112
For what it's worth, Brave Search lies about their User Agent too. I found it fishy as well, but they claim that many websites only allow Googlebot to crawl and ban other UAs. I remember searching for alternative search engines and finding an article that said most new engines face this exact problem: they can't crawl because any unusual bots are blocked.

I have tried programming scrappers in the past and one thing I noticed is that there doesn't seem to be a guide in how to make a "good" bot, since there are so few bots with legitimate use cases. Most people use Chrome, too. So I guess now UA is pointless as the only valid UA is going to be Chrome or Googlebot.

Re: Perplexity AI is lying about their user agent

#113
post #3

I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.

I agree with that, but I also think that they should at least identify themselves instead of using a generic user agent.

Re: Perplexity AI is lying about their user agent

#114
post #84

The author has misunderstood when the perplexity user agent applies. Web site owners shouldn’t dictate what browser users can access their site with - whether that’s chrome, firefox, or something totally different like perplexity. When retrieving a web page _for the user_ it’s appropriate to use a UA string that looks like a browser client. If perplexity is collecting training data in bulk without using their UA that…

It’s not retrieving a web page though is it? It’s retrieving the content then manipulating it. Perplexity isn’t a web browser.

[deleted]

Re: Perplexity AI is lying about their user agent

#115

Earlier quoted context omitted.

Reading a book is not theft. Building a business on processing other people's copyrighted material to produce content is.

I think that's called a school

Main issues:

1) Schools use primarily public domain knowledge for education. It's rarely your private blog post being used to mostly learn writing blog posts.

2) There's no attribution, no credit. Public academia is heavily based (at least theoretically) on acknowledging every single paper you built your thesis on.

3) There's no payment. In school (whatever level) somebody's usually paying somebody for having worked to create a set of educational materials.

Note: Like above. All very theoretical. Huge amounts of corruption in academia and education. Of Vice/Virtue who wants to watch the Virtue Squad solve crimes? What's sold in America? Working hard and doing your honest 9 to 5? Nah.

Re: Perplexity AI is lying about their user agent

#116

Earlier quoted context omitted.

Reading a book is not theft. Building a business on processing other people's copyrighted material to produce content is.

You should be able to judge whether something is a copyright violation based on the resulting work. If a work was produced with or without computer assistance, why would that change whether it infringes?

It helps. If it's at stake whether there is infringement or not, and it comes that you were looking at a photograph of the protected work while working on yours (or any other type of "computer assistance") do you think this would not make for a more clear cut case?

That's why clean room reverse engineering and all of that even exists.

Re: Perplexity AI is lying about their user agent

#117

Earlier quoted context omitted.

I think that's called a school

Schools use books that were paid for and library lending falls under PLR (in the UK), so authors of books used in schools do get compensated. Not a lot, but they are. AI companies are run by people who will loot your place when you're not looking and charge you for access to your own stuff. Fuck that lot.

> AI companies are run by people who will loot your place when you're not looking and charge you for access to your own stuff.

Funnily enough they do understand that having your own product used to build a competing product is uncool, they just don't care unless it's happening to them.

https://openai.com/policies/terms-of-use/

> What you cannot do. You may not use our Services for any illegal, harmful, or abusive activity. For example [...] using Output to develop models that compete with OpenAI.

Re: Perplexity AI is lying about their user agent

#118
post #41

Earlier quoted context omitted.

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

If using copyrighted material to train an LLM is theft, so is reading a book.

Computers are not people. Laws differ and consequences can be different based on the actor (like how minors are treated differently in courts). Just because a person can do it does not automatically mean those same rights transfer to arbitrary machines.

Re: Perplexity AI is lying about their user agent

#120
If we can feed all the knowledge we have into a system that will be able to create novel ideas, help us in a myriad of use cases, isn’t this justification enough to do it?

Isn’t the situation akin to scihub? Or library genesis? Btw: There are endless many people around the globe who cannot pay 30 USD for one book, let alone several books.

Post reply on HN