Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

411–420 of 555 posts

Re: Perplexity AI is lying about their user agent

#411

Earlier quoted context omitted.

I, on the other hand, hope NYT refuses a settlement and OpenAI loses in court.

Be careful what you wish for, because, depending on how broad the reasoning in such a decision would be, it is not impossible that the precedent would be used to then target ad blockers and similar software.

Fair point, but it's a risk I'd be willing to take.

Re: Perplexity AI is lying about their user agent

#412

Earlier quoted context omitted.

I really like this idea. Someone needs to implement this. I'm not sure what the ideal poison would be. Randomly constructed sentences that follow the basic rules of grammar?

fun! but a few ill-intentioned agitators can use up the ability and resources of those trying to fight back. This phenomenon is well-known in legal circles I believe..

> This phenomenon is well-known in legal circles I believe..

I think you’re referring to spoliation, but in this context it could be considered a special-case of a document dump.

https://en.wikipedia.org/wiki/Tampering_with_evidence#Spolia...

https://en.wikipedia.org/wiki/Document_dump

Re: Perplexity AI is lying about their user agent

#413

Earlier quoted context omitted.

I cannot imagine how viewing/scraping a public website could ever be illegal, wrong, immoral etc. I just don't see the argument for it.

It's scraping content to then serve up that content to users who can now get that content from you (via a paid subscription service, or maybe ad-sponsored) instead of visiting the content creator and paying them (i.e., via ads on their website) It's the same reason I can't just take NYT archives or the Britannica and sell an app that gives people access to their content through my app. It totally undercuts content cr…

One more point on this, lest some people think, "hey Kanye, or Taylor Swift, don't need any more money!" I 100% agree. But the problem with streaming is that is disproportionately rewards the biggest artists at the expense of the smaller ones. It's the small artist, barely making a living from their craft, who were most hurt by the switch from albums to streaming, not those making millions.

Re: Perplexity AI is lying about their user agent

#414

Earlier quoted context omitted.

> they are decreasing the probability that this user would come to by content (via Google, for example). Google has been providing summaries of stuff and hijacking traffic for ages. I kid you not, in the tourism sector this has been a HUGE issue, we have seen 50%+ decrease in views when they started doing it. We paid gazzilions to write quality content for tourists about the most different places just so Google could…

I'm curious about the tourism sector problem. In tourism, I would think the goal would be to promote a location. You want people to be able to easily discover the location, get information about it, and presumably arrange to travel to those locations. If Google gets the information to the users, but doesn't send the tourist to the website, is that harmful? Is it a problem of ads on the tourism website? Or is more of…

Google snippets are hilariously wrong, absurdly often; I was recently searching for things while traveling and I can easily imagine relying on snippets getting people into actual trouble.

Re: Perplexity AI is lying about their user agent

#415

Earlier quoted context omitted.

I don't, because if it will, then someone like the author of the article will do the obnoxious thing and ban it. We've been there before, 30 years ago. That's why all browsers' user agent strings start with "Mozilla".

Why is the author here obnoxious, and not Perplexity? I don't want these scumbag AI companies making money off me, end of story.

The "scumbag AI company" in question is making money by offering me a way to access information while skipping any and all attention economy bullshit you may have on your site, on top of being just plain more convenient. Note that the author is confusing crawling (which is done with documented User Agent and presumably obeys robots.txt) with browsing (which is done by working as one-off user agent for the user).

As for why this behavior is obnoxious, I refer you to 30 years worth of arguing on this, as it's been discussed ever since User-Agent header was first added, and then used by someone to discriminate visitors based on their browsers.

Re: Perplexity AI is lying about their user agent

#416
post #409

Earlier quoted context omitted.

Perplexity isn't playing in the attention economy unless they upsell you, advertise to you, or put any other kind of bullshit between you and your goal. Attention economy is (as the name suggests) about monetizing attention; it does so through friction.

I didn’t write they would. I said “like”. The next perplexity will show ads. The attention economy will not die. Because it’s hasn’t for the last 100 years. The profits just shift to where the attention is now.

Fair enough, I agree with that. Hell, we may not need a next Perplexity, this one may very well enshittify couple years down the line - as it happens to almost any service offered commercially on the Internet. I was just saying it isn't happening now - for the moment, Perplexity has arguably much better moral standing than most of the websites they scrape or allow users to one-off browse.

Re: Perplexity AI is lying about their user agent

#417
post #177

Earlier quoted context omitted.

The companies will scrape and internalise the "customer asked for this" requests... and slowly turn the latter into the former, or just their own tool as the scraper. No, easier to just ask a simple question: Does the company respect the access rules communicated via a web standard? No? In that case hard deny access to that company. These companies don't need to be given an inch.

> Does the company respect the access rules communicated via a web standard? No? In that case hard deny access to that company. So should Firefox not allow changing the user agent in order to bypass websites that erroneously claim to not work on Firefox?

Similarly, for sites which configure robots.txt to disallow all bots except Googlebot, I don't lose sleep about new search engines taking that with a grain of salt.

Re: Perplexity AI is lying about their user agent

#418

Earlier quoted context omitted.

> spoofing an agent seems in dubious territory. Just to clarify, Perplexity is not spoofing a user agent, they're legitimately using a headless Chrome to fetch the page. The author just misunderstood their docs [0]: when they say that "you can identify our web crawler by its user agent", they're talking about the crawler, not the browser they use for ad hoc requests. As you note, crawling is different. [0] https://do…

This is completely false, the user agent being used by Perplexity its _not_ the headless-chrome user agent, wich is close similar to this (emphasis on HeadlessChrome): Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/119.0.0.0 Safari/537.36 They are spoofing it to pretend to be a desktop Chrome one: Mozilla/5.0 (Windows NT 10.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/111.0.…

There's a difference here between "headless chrome" as a concept and "headless-chrome" the software. It's still pretty common to run browser automation with a full "headful" browser, in which case you would just get the normal user agent. headless-chrome is sort of an optimized option that comes with some downsides.

Re: Perplexity AI is lying about their user agent

#420

Earlier quoted context omitted.

it's comparable exactly in the way 0.001% can be compared to 10^100 humans learning is the old-school digital copying. computers simply do it much faster, but it's the same basic phenomenon consider one teacher and one student. first there is one idea in one head but then the idea is in two heads. now add book technology1 the teacher writes the book once, a thousand students read it. the idea has gone from being in o…

> humans learning is the old-school digital copying. computers simply do it much faster, but it's the same basic phenomenon Train an LLM on the state of human knowledge 100,000 years ago - language had yet to be invented and bleeding edge technology was 'poke them with the pointy side.' It's not going to be able to do or output much of anything, and it's going to be stuck in that state for perpetuity until somebody g…

> bleeding edge technology was 'poke them with the pointy side.'

Relevant: https://www.smbc-comics.com/comic/rise-of-the-machines

Post reply on HN