Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

501–510 of 555 posts

Re: Perplexity AI is lying about their user agent

#501

Earlier quoted context omitted.

> I think there is a real content dilemma here at work It's not really a dilemma. This is exactly what copyright serves to protect authors from. Perplexity copied the content, and in doing so directly competes with the original work, destroying it's market value and driving the original author out of business. Literally what copyright was invented to prevent. It's the exact same situation as journalists going after G…

I see no competition. I use Perplexity regularly to give me summaries of articles or to do preliminary research. If I like what I'm seeing, then I go to the source. If a source chooses to block their content because they don't want it to be accessed by AI bots then they reduce even further the chance of me - and increasingly more persons - touching their site at all.

You can say that, it doesn't matter. The statistics show that these tools reduce views.

And really, "I'm going to replace my entire news intake with the AI slop even if it's entirely hallucinated lies or propaganda" is perhaps not something you ought to say out loud.

Re: Perplexity AI is lying about their user agent

#502

Earlier quoted context omitted.

Just to go a little bit more into detail on this, because the article and most of the conversation here is based on a big misunderstanding: robots.txt governs crawlers . Fetching a single user-specified URL is not crawling. Crawling is when you automatically follow links to continue fetching subsequent pages. Perplexity’s documentation that the article links to describes how their crawler works. That is not the piece…

Its fairly logical to assume that robots.txt governs robots (empahsis in "bots") not just crawlers, if they are only intended to block crawlers why aren't they called crawlers.txt instead and remove all ambiguity?

It’s not logical to assume anything about a standard merely from a 30-year-old filename when you can just read the documentation.

> WWW Robots (also called wanderers or spiders) are programs that traverse many pages in the World Wide Web by recursively retrieving linked pages.

http://www.robotstxt.org/orig.html

Re: Perplexity AI is lying about their user agent

#503

Earlier quoted context omitted.

I'm curious to know where you draw the line for what constitutes legitimate manipulation by a person and when it becomes distribution. I'm assuming that if I write code by hand for every part of the TCP/IP and HTTP stack I'm safe. What if I use libraries written by other people for the TCP/IP and HTTP part? What if I use a whole FOSS web browser? What about a paid local web browser? What if I run a script that I wrot…

Where exactly you crossed the line is a question for the courts. I am not a lawyer and will there for not help with the specifics. However, please see the Aereo case [0] for a possibly analogous case. I am allowed to have a DVR. There is no law preventing me from accessing my DVR over a network. Or possibly even colocating it in a local data center. But Aereo definitely crossed a line. Also see Vidangel [1]. The fact…

>Where exactly you crossed the line is a question for the courts. I am not a lawyer and will there for not help with the specifics.

I expect you're right. Although Perplexity thinks they're well within the law[0]. Are they correct? I guess we'll see....

[0] https://www.perplexity.ai/search/Why-are-you-2wJteqZ4SUCqPjk...

Re: Perplexity AI is lying about their user agent

#504
UA aside (and presumably the spirit of the UA and robots.txt is about measuring intent), Perplexity could announce an IP range to allow people to reliably block the requests. Problem solved.

Read a few comments implying that a browser UA implies capabilities, tbf they should simply change their UA and not use a generic browser UA.

Re: Perplexity AI is lying about their user agent

#505
post #211

Earlier quoted context omitted.

How would an LLM training on your writing reduce your reward? I guess if you're doing it for a living sure, but most content I consume online is created without incentive (social media, blogs, stack overflow). I write a fair amount and have been for a few years. I like to play with ideas. If an llm learned from my writing and it helped me propagate my ideas, I'd be happy. I lose on social status imaginary internet po…

Speaking as an SO contributor, I'm perfectly fine with having an LLM read my answers and produce output based on them. What I'm not okay with is said LLM being closed-weight so that its creator can profit off it. When I posted my answers on SO, I did so under CC-BY-SA, and I don't think it's unreasonable for me to expect any derivatives to abide by both the letter and the spirit of this arrangement.

This hits the nail completely on the head.

If the issue here was "just" training LLMs, like some AI bros want to deflect it to be, the conversation around this topic would be very different, and I would be enthusiastically defending the model trainers.

But that's not this conversation. These are companies that are trying to fold our permissively-license content into weights, close source it, and make themselves the only access point, all while pre-emptively perform regulatory capture with all the right DEI buzzwords so that the open source variants are sufficiently demonized as "alt-right" and "dangerous".

The thing that truly frightens me is that (even here on Hacker News) there is an increasing number of people that have fallen for the DEI FUD and are honestly cheering on the Sam Altmans of the world to control the flow of information.

Re: Perplexity AI is lying about their user agent

#506
post #492

Earlier quoted context omitted.

Moreover, if it's the former, then good riddance . An ad-backed site is harming users a little on the margin for the marginal piece of information. Getting the same from a search engine is saving users from that harm. Parent has the right question here: why did you want the traffic? Did you intend for anything good to happen to those people? . I'm going to guess not; there's hardly a scenario where people who complai…

Now think of the 2nd order effects: they paid money to collect that useful information. If it’s no longer feasible to create such high quality content, it won’t magic itself into existence on its own. It’ll all be just crap and slop in a few years.

In my experience, the highest-quality content on the internet was created without a profit motive.

Re: Perplexity AI is lying about their user agent

#507

Earlier quoted context omitted.

If explicitly telling it to access a URL is an access by automaton, then isn't every web browser load an access by automaton?

The flaw with that example is your web browser isn't between other users and the website, turning 500 views into one. And if we took the analogy to the other end, one could argue that all crawlers have to be kicked off manually at some point... The problem is here in reality the differentiation is somewhat more understood. The honor system web is going away, that's for sure.

> your web browser isn't between other users and the website, turning 500 views into one.

There are a lot of people making this assumption about the way Perplexity is working, but there is no evidence in TFA that Perplexity is caching its ad hoc requests.

And even if they were, what's left unsaid is why it even would matter if 500 views turned into one. It matters either because of lost ad revenue or lost ability to track the users' behavior. Personally, I'm okay with moving past that phase of the internet's life and look forward to new business models that aren't built around getting large numbers of "views".

Re: Perplexity AI is lying about their user agent

#508
post #485

Earlier quoted context omitted.

I'm not. I'm asking why this flow is "distribution": * User types an address into Perplexity * Perplexity fetches the page, transforms it, and renders some part of it for the user But this flow is not: * User types an address into Orion Browser * Orion Browser fetches the page, transforms it, and renders some part of it for the user Regardless of the legal question (which I'm also skeptical of), I'm especially unconv…

The moral case is pretty obviously that Perplexity is preventing traffic from reaching the people who made the content.

How so? TFA pretty clearly shows that traffic does reach the server, how else would it show up in the logs?

Also, the author of TFA has already gotten themselves deindexed, the behavior they're complaining about now is that if someone copies and pastes a link into Perplexity it will go fetch the page for the user and summarize it.

This scenario presupposes that the user has a link to a specific page. I suspect that in nearly all cases that link will be copied from the address bar of an open tab. This means that most of the time the site will actually get double the traffic: one hit when the user opens it in the browser and a second when Perplexity asks for the page to summarize it.

Re: Perplexity AI is lying about their user agent

#509

Earlier quoted context omitted.

Isn't that just the American spelling? I always assume Americans remove 'u' from everything.

Yes, actually the very first thing we did was remove u. Sorry, I couldn’t resist. Checking the wiki page on British/American spelling differences, it looks like there are also a handful of words which have diverged completely. For example aluminum/aluminium and airplane/aeroplane.

Almost one here actually spells it aeroplane though. We do write and say aluminium though.

Re: Perplexity AI is lying about their user agent

#510

Earlier quoted context omitted.

We would employ local guides all around the world to craft itinerary plans to visit places, give tips, tricks, recommend experiences and places (we made money by selling some of those through our website) and it was a success. Customers liked the in depth value of that content and it converted to buys (we sold experiences and other stuff, sort of like getyourguide). One day all of our content ended up on Google "what…

I totally get that it killed your traffic. If a thousand people a day typing in "what time is best to visit the Sagrada Familiar" stopped clicking on the link to your page because Google just told them "4 PM on Thursdays" at the top of the page, you lost a bunch of traffic. But why did you want the traffic? Was your revenue from ad impressions, or were you perhaps being paid by the city of Barcelona to provide useful…

I think I have answered this already in the post, haven't I?

We sold experiences, thus we created a lot of free content from local experts and hoped that they would buy some of the tickets through our website.

Post reply on HN