Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

401–410 of 555 posts

Re: Perplexity AI is lying about their user agent

#401
post #3

I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.

Why should I have to differentiate Perplexity's services?

Re: Perplexity AI is lying about their user agent

#402
post #225

Earlier quoted context omitted.

For me, the irony is the opposite side of the same coin, 30 years of "information wants to be free" and "copyright infringement isn't piracy" and "if you don't want to be indexed, use robots.txt"… …and then suddenly OpenAI are evil villains, and at least some of the people denounced them for copyright infringement are, in the same post, adamant that the solution is to force the model weights to become public domain.

I broadly agree with you, but I don't see what's contradictory about the solution of model weights becoming public domain. When it comes to piracy, the people who have viewed it as ethical on the grounds that "information wants to be free" generally also drew the line at profiting from it: copying an MP3 and giving it to your friend or even a complete stranger is ethical, charging a fee for that (above and beyond wha…

To me, it's like trying to "solve The Pirate Bay" by making all the stuff they share public domain.

But thank you for sharing your perspective, I appreciate that.

Re: Perplexity AI is lying about their user agent

#403
post #312

Earlier quoted context omitted.

This is why this conversation is making me insane. How are people saying straight-faced that the user is requesting a specific page? They aren't, they're doing a search of the web. That's not at all the same as a browser visiting a page.

Because that's literally what the author does in TFA and then complains about when Perplexity complies. > What is this post about https://rknight.me/blog/blocking-bots-with-nginx/

Am I the only one that sees a difference between “show me page X” and “what is page X about”?

The first is how browsers work. The second is what perplexity is doing.

Those two are clearly different imo.

Re: Perplexity AI is lying about their user agent

#404
I was going to reply in thread, but this comment and my reply are directed at the whole thread generally, so I’ve chosen to reply-all in hopes of promoting wider discussion.

https://news.ycombinator.com/item?id=40692432

> And if the answer is "scale", that gets uncomfortably close to saying that it's okay for the rich but not for the plebs.

This is the correct framing of the issues at hand.

In my view, the issue is one of class as viewed through the lens of effort vs reward. Upper middle class AI developers vs middle class content creators. Now that lower class content creators can compete with middle and upper class content creators, monocles are dropping and pearls are clutched.

I honestly think that anyone who is able to make any money at all from producing content or cultural artifacts should count themselves lucky, and not take such payments for granted, nor consider them inherently deserved or obligatory. On an average individual basis, those incomes are likely peaking and only going down outside of the top end market outliers.

Capitalism is the crisis. Copyright is a stalking horse for capital and is equally deserving of scrutiny, scorn, and disruption.

AI agents are democratizing access to information across the world just like search engines and libraries do.

Those protesting AI acting on behalf of users seems entitled to me, like suing someone for singing Happy Birthday. Copyright was a mistake. If you don’t want others to use what you made anyway they want, don’t sell it on the open market. If you don’t want other to sing the song you wrote, why did you give it away for a song?

Recently YouTube started to embed ads in the content stream itself. Others in the comments have mentioned Cloudflare and other methods of blocking. These methods work for megacorps who already benefit from the new and coming AI status quo, but they likely will do little to nothing to stem the tide for individuals. It’s just cutting your nose off to spite your face.

If you have any kind of audience now or hope to attract one in the future, demonstrate value, build engagement, and grow community, paid or otherwise. A healthy and happy community has value not just to the creator, but also to the consumer audience. A good community is non-rivalrous; a great community is anti-rivalrous.

https://en.wikipedia.org/wiki/Rivalry_(economics)

https://en.wikipedia.org/wiki/Anti-rival_good

Re: Perplexity AI is lying about their user agent

#405
post #352

Earlier quoted context omitted.

The problem that Perplexity has that ad blockers don't is that they're an independent site that is publishing content based on work they didn't produce. That runs afoul of both copyright laws and section 230 which let's sites like Google and Facebook operate. That's pretty different from an ad blocker running on your local machine. The ad blocker isn't publishing the page it edited for you.

> they're an independent site that is publishing content based on work they didn't produce. What distinguishes these two situations? * User asks proprietary web browser to fetch content and render it a specific way, which it does * User asks proprietary web service to fetch content and render it a specific way, which it does The technical distinction is that there's a network involved in the second scenario. What is…

I don't have a horse in this race, but:

> * User asks proprietary web service to fetch content and render it a specific way, which it does

That sounds like Google Translate to me, when pasting a URL.

Bonus points if instead of pasting a URL directly, it is submitted to one of the Internet Archive-like sites; and then submit that archive URL to Google Translate. That would be download and adaptation (by Google Translate) of the download and adaptation[1] (by Internet Archive) of the original content.

[1]: These archive sites usually present the content in a slightly different way. Granted, it's usually just adding stuff around the page, e.g. to let you move around different snapshots, but that's still showing stuff that was originally not there.

Re: Perplexity AI is lying about their user agent

#406

A lot of comments here are confusing the two use cases for crawling: training and summarization. Perplexity's utility as an answer engine is RAG (retrieval augmented generation). In response to your question, they search the web, crawl relevant URLs and summarize them. They do include citations in their response to the user, but in practice no one clicks through on the tiny (1), (2) links to go to the source. So if y…

The real question here is whether websites are entitled to that traffic, or even more specifically, to human eyes - and to what extent that should allow them to override users' preferences (which are made fairly clear by the very act of using Perplexity in the first place; the reason why you'd do it instead of doing a Google Search and then manually sifting through the links yourself is because most of what you see i…

I like your comment a lot, so much so that I replied to it on the top-level in hopes of promoting wider discussion of the points you have raised:

https://news.ycombinator.com/item?id=40693140

Re: Perplexity AI is lying about their user agent

#407
post #344

Earlier quoted context omitted.

What, users won't share anything? I said I wanted Perplexity to identify themselves in the user agent instead of using the generic "Mozilla/5.0 (Windows NT 10.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/111.0.0.0 Safari/537.3" they're using right now for the "non-scraper bot". How does that impact users at all?

I don't, because if it will, then someone like the author of the article will do the obnoxious thing and ban it. We've been there before, 30 years ago. That's why all browsers' user agent strings start with "Mozilla".

Why is the author here obnoxious, and not Perplexity? I don't want these scumbag AI companies making money off me, end of story.

Re: Perplexity AI is lying about their user agent

#409
post #379

Earlier quoted context omitted.

The attention economy will never die. Attention will only shift. From websites to aggregators like perplexity.

Perplexity isn't playing in the attention economy unless they upsell you, advertise to you, or put any other kind of bullshit between you and your goal. Attention economy is (as the name suggests) about monetizing attention; it does so through friction.

I didn’t write they would. I said “like”. The next perplexity will show ads.

The attention economy will not die. Because it’s hasn’t for the last 100 years. The profits just shift to where the attention is now.

Re: Perplexity AI is lying about their user agent

#410
post #295

With all the ad blockers out there, which functionally demonetize content sites, why isn’t there an ad equivalent to robots.txt that says “don’t display this site if ads are blocked”? So many good comments from several points of view in this thread and the thing I can’t square is the same person championing ad blockers and condemning agents like Perplexity.

Because these are all voluntary standards. If you want your content to be discoverable and accessible, you don’t get to dictate how someone renders it. If you want to force monetization, adopt a different business model.
Post reply on HN