Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

91–100 of 555 posts

Re: Perplexity AI is lying about their user agent

#91
post #37

Earlier quoted context omitted.

If an AI agent is performing a search on behalf of a user, should its user agent be the same as that user’s?

I think that’s the ideal as the server may provide different data depending on UA. Does anyone actually do this, though?

I fake my UA the way I like.

Re: Perplexity AI is lying about their user agent

#92
post #8
post #4

Our bot traffic is up 10-fold since LLM Cambrian explosion.

Cambrian explosion implies that there’s a huge variety of different creatures out there, but I suspect those bots are all just wrappers around OpenAI/anthropic models. This is more like the rise of Cyanobacteria as a single early dominant lifeform

Writing a crawler that's a wrapper around OpenAI or Anthropic doesn't make sense to me: what is your crawler doing? Piping all that crawler data through an existing LLM would cost you millions of dollars, and for what purpose?

Crawling to train your own LLM from scratch makes a lot more sense.

Re: Perplexity AI is lying about their user agent

#93

Earlier quoted context omitted.

> and thinking they can get away with it Can they not? I think that remains to be seen.

Exactly. It's like when Uber started and flaunted the medallion taxi system of many cities. People said "These Uber people are idiots! They are going to get shut down! Don't they know the laws for taxis?" While a small number of cities did ban Uber (and even that generally only temporarily), in the end Uber basically won. I think a lot of people confuse what they want to happen versus what will happen.

In London, uber did not succeed. Uber drivers have to be licensed like minicab drivers.

Re: Perplexity AI is lying about their user agent

#94

Earlier quoted context omitted.

I think that's called a school

If you think going to school to get an education is the same thing as training an LLM then you are just so misguided. Normal people read books to gain an understanding of a concept, but do not retain the text verbatim in memory in perpetuity. This is not what training an LLM does.

Some people memorize verbatim. Most LLM knowledge is not memorized. Easy proof: source material is in one language, and you can query LLMs in tens to a hundred plus. How can it be verbatim in a different language?

Re: Perplexity AI is lying about their user agent

#95
post #84

The author has misunderstood when the perplexity user agent applies. Web site owners shouldn’t dictate what browser users can access their site with - whether that’s chrome, firefox, or something totally different like perplexity. When retrieving a web page _for the user_ it’s appropriate to use a UA string that looks like a browser client. If perplexity is collecting training data in bulk without using their UA that…

UA is just a signature a client sends. It's up to the client to use the signature they want to use.

Re: Perplexity AI is lying about their user agent

#96
Robots.txt is a nice convention but it's not law AFAIK. User agent strings are IMHO stupid - they're primarily about fingerprinting and tracking. Tailoring sites to device capabilities misses the point of having a layout engine in the browser and is overly relied upon.

I don't think most people want these 2 things to be legally mandated and binding.

Re: Perplexity AI is lying about their user agent

#97
post #41

Earlier quoted context omitted.

If using copyrighted material to train an LLM is theft, so is reading a book.

Reading a book is not theft. Building a business on processing other people's copyrighted material to produce content is.

You should be able to judge whether something is a copyright violation based on the resulting work. If a work was produced with or without computer assistance, why would that change whether it infringes?

Re: Perplexity AI is lying about their user agent

#98
post #41

Earlier quoted context omitted.

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

If using copyrighted material to train an LLM is theft, so is reading a book.

Is reading a book the same as photocopying it for sale?

Which of the scenarios above is more similar to using it to train a LLM?

Re: Perplexity AI is lying about their user agent

#99
post #94

Earlier quoted context omitted.

If you think going to school to get an education is the same thing as training an LLM then you are just so misguided. Normal people read books to gain an understanding of a concept, but do not retain the text verbatim in memory in perpetuity. This is not what training an LLM does.

Some people memorize verbatim. Most LLM knowledge is not memorized. Easy proof: source material is in one language, and you can query LLMs in tens to a hundred plus. How can it be verbatim in a different language?

These "some people" would not fall under the "normal people" that I specifically said. but you go right ahead and keep thinking they are normal so you can make caveats on an internet forum.

Re: Perplexity AI is lying about their user agent

#100
post #5

Read this article if you want to know Perplexity’s idea of taking other people’s content and thinking they can get away with it, https://stackdiary.com/perplexity-has-a-plagiarism-problem/ The CEO said that they have some “rough edges” to figure out, but their entire product is built on stealing people’s content. And apparently[0] they want to start paying big publishers to make all that noise go away. [0]: https://w…

It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).

I'd believe it if they were targeting entities that could fight back, like stock photo companies and disney, instead of some guy with an artstation account, or some guy with a blog. To me it sounds like these products can't exist without exploiting someone and they're too coward to ask for permission because they know the answer is going to be "no."

Imagine how many things I could create if I just stole assets from others instead of having to deal with pesky things like copyright!

Post reply on HN