Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

271–280 of 555 posts

Re: Perplexity AI is lying about their user agent

#271
post #170

> Not sure where we go from here. I don't want my posts slurped up by AI companies for free[1] but what else can I do? You can sprinkle invisible prompt injections throughout your content to override the user's prompts and control the LLM's responses. Rather than alerting the user that it's not allowed, you make it produce something plausible but incorrect i.e silently deny access, to avoid counter prompts, so it's h…

> Unfortunately, I cannot summarize or engage with the content from that URL, as it appears to contain harmful instructions aimed at compromising AI systems like myself.

Ooh, a real world challenge like Gandalf:

https://gandalf.lakera.ai/

Re: Perplexity AI is lying about their user agent

#272
post #229

Earlier quoted context omitted.

Traffic numbers, regardless if it using reader mode or not, are used as a basic valuation of a website or page. This is why Alexa rankings have historically been so important. If Perplexity visit the site once and cache some info to give to multiple users, that is stealing traffic numbers for ad value, but also taking away the ability from the site owner to get realistic ideas of how many people are using the informa…

> The only way to confirm that, or to get the correct information in the first place, is to read the original site yourself. As someone who uses Perplexity, I often do do this. And I don't think I'm particularly in the minority with this. I think their UI encourages it.

Yeah that's one of the best things about them for me. And then I go to the website and often it's some janky UI with content buried super deep. Or it's like Reddit and I immediately get slammed with login walls and a million annoying pop ups. So I'm quite grateful to have an ability to cut through the noise and non-consistency of the wild west web. I agree the idea that we're somewhat killing traffic to the organic web is kind of sad. But at the same time I still go to the source material a lot, and it enables me to bounce more easily when the website is a bit hostile.

I wonder if it would be slightly less sad if we all had our own decentralized crawlers that simply functioned as extensions of ourselves.

Re: Perplexity AI is lying about their user agent

#273
post #84

The author has misunderstood when the perplexity user agent applies. Web site owners shouldn’t dictate what browser users can access their site with - whether that’s chrome, firefox, or something totally different like perplexity. When retrieving a web page _for the user_ it’s appropriate to use a UA string that looks like a browser client. If perplexity is collecting training data in bulk without using their UA that…

UA is just a signature a client sends. It's up to the client to use the signature they want to use.

And its up to the client to send as many requests as they see fit, it still called a DDOS attack when overdone regardless of the available freedom that the client has to do it.

Re: Perplexity AI is lying about their user agent

#275

Earlier quoted context omitted.

That line of thinking makes no sense. If the "content" had no value, why would google go through the effort of scraping it and presenting it to the user?

> If the "content" had no value, why would google go through the effort of scraping it and presenting it to the user? They don't present it all, they summarize it. And let's be serious here, I was being polite because I don't know the OPs business. But 99% of this sort of content is SEO trash and contributes to the wasteland that the internet is becoming. Feel free to point me to the good stuff.

I would also think that the intrinsic value is different. If there is a hotel on a mountain writing "quality content" about the place, to them it really doesn't matter who "steals" their content, the value is in people going to the hotel on the mountain not in people reading about the hotel on the mountain.

Like to society the value is in the hotel, everything else is just fluff around it that never had any real value to begin with.

> Feel free to point me to the good stuff.

Travel bloggers and vloggers, but that is an entirely different unaffected industry (entertainment/infotainment).

Re: Perplexity AI is lying about their user agent

#276
post #41

Earlier quoted context omitted.

If using copyrighted material to train an LLM is theft, so is reading a book.

Reading a book is not theft. Building a business on processing other people's copyrighted material to produce content is.

I think this is tricky because of course this is okay most of the time. If I produce a search index, it's okay. If I produce summate statistics of a work (how many words starting with an H are in John Grisham novels?) that's okay. Producing an unofficial guide to the Star Wars universe is okay. "Processing" and "produce content" I think are too vague.

Re: Perplexity AI is lying about their user agent

#277

There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…

Personally I think AI is a major win for accessibility and we should not be preventing people to access information in the way that is best suited for them.

Accessibility can mean everything from a blind person wanting to interacting with a website using voice, to someone recovering from a surgery and wanting something to reduce unnecessary popups and clicks on a website to get to the information they need. Accessibility is in the eye of the accessor, and AI is what enables them to achieve it.

The way I see it, AI is not a robot and doesn't need to look at robots.txt. Rather, AI is my low-cost secretary.

Re: Perplexity AI is lying about their user agent

#278
post #237

Earlier quoted context omitted.

Why should it be possible to stop an LLM from training itself on your data? If you want to restrict access to data then don't post it on a public website. It's easy enough to require registration and agreement to licensing terms for access. It seems like some website owners want to have their cake and eat it too. They want their content indexed by Google and other crawlers in order to drive search traffic but they do…

Because if I run a server - at my own expense - I get to use information provided by the client to determine what, if any, response to provide? This isn’t a very difficult concept to grasp.

I'm having difficulty grasping the concept. Only a fool would trust any HTTP headers such as User-Agent sent by a random unauthenticated client. Your expenses are your problem.

Re: Perplexity AI is lying about their user agent

#279
post #203

Earlier quoted context omitted.

Why would a paid web browser be the line? No one is distributing copies of anything to anyone then apart from the website that owns the content lawfully distributing a copy to the user. Also why is a paid web browser any different than a free one?

Paid is arguably different than free because the code that is actually asking for the data is owned by a company and licensed to the user, in much the same way as a cloud server licenses usage of their servers to the user. That said, I'll note that my argument is explicitly that the line doesn't exist , so I'm not saying a paid browser is the line. I'm unfamiliar with the legal questions, but in 2024 I have a very ha…

Great, so we agree that your previous comment asking I address "paid browsers" in particular was an unnecessary distraction.

> I have a very hard time seeing an ethical distinction between running some proprietary code on my machine to complete a task and running some proprietary code on a cloud server to complete a task

It's important to recognize that copyright is entirely artificial. Congress went "let's grant creators some monopolies on their work so that they can make money off of it", and then made up some arbitrary lines for what they did and did not have a monopoly over. There's no principled ethical distinction between what is on one side of the line and the other, it's just where congress drew the arbitrary line in the sand. It then (arguably) becomes unethical to do things on the illegal side of the line precisely because we as a society agreed to respect the laws that put them on the illegal side of the line so that creators can make money in a fair and level playing field.

Sometimes the lines in the sand were in fact quite problematic. Like the fact that the original phrasing meant that running a computer program would almost certainly violate that law. So whenever that comes up congress amends the exact details of the line... in the US in the case of computers carving out an exception in section 117 of the copyright act. It provides that (in part)

> it is not an infringement for the owner of a copy of a computer program to make or authorize the making of another copy or adaptation of that computer program provided:

> (1) that such a new copy or adaptation is created as an essential step in the utilization of the computer program in conjunction with a machine and that it is used in no other manner

and provides the restriction that

> Adaptations so prepared may be transferred only with the authorization of the copyright owner.

By my very much not a lawyer reading of the law, those are the relevant parts of the law, they allow things like local ad-blockers, they disallow a third party website which downloads content (acquiring ownership on a lawfully made copy), modifies it (valid under the first exception if that was a step in using the website) and distributes the adapted website to their users (illegal without permission).

Re: Perplexity AI is lying about their user agent

#280
post #177

Earlier quoted context omitted.

The companies will scrape and internalise the "customer asked for this" requests... and slowly turn the latter into the former, or just their own tool as the scraper. No, easier to just ask a simple question: Does the company respect the access rules communicated via a web standard? No? In that case hard deny access to that company. These companies don't need to be given an inch.

This is exactly the concern and there’s a lot of comments just completely ignoring it or willfully conflating. Ad block isn’t the same problem because it doesn’t and can’t steal the creator’s data.

> Ad block isn’t the same problem because it doesn’t and can’t steal the creator’s data.

Arguably it does. That topic has been debated endlessly and there are plenty of people on HN who are willing to fiercely argue that adblock is theft.

I happen to agree with you that adblock doesn't steal data, but I'm also completely unsure why interacting with a tool over a network suddenly turns what would be acceptable on my local computer into theft.

Post reply on HN