Live data from Hacker News

Perplexity AI is lying about their user agent

rknight.me

461–470 of 555 posts

Re: Perplexity AI is lying about their user agent

#461

Earlier quoted context omitted.

Because that's literally what the author does in TFA and then complains about when Perplexity complies. > What is this post about https://rknight.me/blog/blocking-bots-with-nginx/

Am I the only one that sees a difference between “show me page X” and “what is page X about”? The first is how browsers work. The second is what perplexity is doing. Those two are clearly different imo.

You are not.

Perplexity should always respect robots.txt, even for summarization requests. If I say that I don't want Perplexity crawling my site, I mean at all, and I explicitly would not want them "summarizing" my page.

The response from Perplexity to such a request should be "The owner of this page/site does not permit Perplexity to process any data from this site." Period.

LLMs can't summarize in any case: https://ea.rna.nl/2024/05/27/when-chatgpt-summarises-it-actu...

Re: Perplexity AI is lying about their user agent

#462

Earlier quoted context omitted.

Its fairly logical to assume that robots.txt governs robots (empahsis in "bots") not just crawlers, if they are only intended to block crawlers why aren't they called crawlers.txt instead and remove all ambiguity?

That's a historical question. At this time, most if not all the bots were either search engines or archival. The name was even "RobotsNotWanted.txt" at the beginning but made "robots.txt" for simplicity. To give another example, Internet Archive stopped respecting it a couple of years ago, and they discuss this point (crawlers vs other bots) here [1]. [1] https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea..…

You meant search bots and other bots? Internet Archive's bot is a crawler.

They showed no difference between search bots and archive bots. robots.txt was never for SEO alone. Sites exclude print versions so people see more ads and links to other pages. Sites exclude search pages to conserve resources. They said sites exclude large files for costs. And they can't think sites want sensitive areas like administrative pages archived.

Really Internet Archive stopped respecting robots.txt because they wanted to archive what sites didn't want them to archive. Many sites disallowed Internet Archive specifically. Many sites allowed specific bots. Many sites disallowed all bots and meant all bots. And hiding old snapshots when a new domain owner changed robots.txt was a self inflicted problem. robots.txt says what to crawl or not now. They knew all of this.

Re: Perplexity AI is lying about their user agent

#463
post #460

I'm martian and I learned to use TCP/IP to make requests to IP addresses on Earth internet and interpret any response I get, however I'd like. I have been enjoying myself but recently came across some bruhaha around robot.txt, user agents and blah and apparently I'm not allowed to do whatever I want with the responses I get from my requests. I'm confused: you're willingly responding to my requests with strings of 0s…

jokes (not so joking) aside: I'd love for a bot to 100% sit between me and "web browsing" 100% of the time. I only want reader mode content. I don't care for ads. and if you need me to pay - ask for it, in text. put a link and clearly state that for me to get those 0s and 1s I need to pay. it's not hard. physical shops do this. it's 2024, it's fine to put up paywalls. yeah, it may break some biz models, but that's just evolution

Re: Perplexity AI is lying about their user agent

#464

A lot of comments here are confusing the two use cases for crawling: training and summarization. Perplexity's utility as an answer engine is RAG (retrieval augmented generation). In response to your question, they search the web, crawl relevant URLs and summarize them. They do include citations in their response to the user, but in practice no one clicks through on the tiny (1), (2) links to go to the source. So if y…

> in practice no one clicks through on the tiny (1), (2) links to go to the source

I offer my self as specimen of someone who clicks on those citations ALL the time because thats how I can - most of the time - find download links, and other details faster than asking again

Re: Perplexity AI is lying about their user agent

#465

Earlier quoted context omitted.

Am I the only one that sees a difference between “show me page X” and “what is page X about”? The first is how browsers work. The second is what perplexity is doing. Those two are clearly different imo.

You are not. Perplexity should always respect robots.txt, even for summarization requests. If I say that I don't want Perplexity crawling my site, I mean at all , and I explicitly would not want them "summarizing" my page. The response from Perplexity to such a request should be "The owner of this page/site does not permit Perplexity to process any data from this site." Period. LLMs can't summarize in any case: https…

> Perplexity should always respect robots.txt, even for summarization requests. If I say that I don't want Perplexity crawling my site, I mean at all

Issuing a single HTTP request is definitionally not crawling, and the robots.txt spec is specifically for crawlers, which this is not.

If you want a specific tool to exclude you from their web request feature you have to talk to them about it. The web was designed to maximize interop between tools, it correctly doesn't have a mechanism for blacklisting specific tools from your site.

Re: Perplexity AI is lying about their user agent

#466

Earlier quoted context omitted.

Why is the author here obnoxious, and not Perplexity? I don't want these scumbag AI companies making money off me, end of story.

The "scumbag AI company" in question is making money by offering me a way to access information while skipping any and all attention economy bullshit you may have on your site, on top of being just plain more convenient. Note that the author is confusing crawling (which is done with documented User Agent and presumably obeys robots.txt) with browsing (which is done by working as one-off user agent for the user). As f…

If you want summaries from my website, go to my website. I want a way to deny any licence to any third-party user agent that will apply machine learning on my content, whether you initiated the request or not.

LLMs — and more importantly the companies that train and operate them — should not be trusted at all, especially for so-called "summarization": https://ea.rna.nl/2024/05/27/when-chatgpt-summarises-it-actu...

While Perplexity may be operating against a particular URL based on a direct request from you, they are acting improperly when they "summarize" a website as they have an implicit (and sometimes explicit if there's a paywall) licence to read and render the content as provided, but not to process and redistribute such content.

There needs to be something stronger than robots.txt, where I can specify the uses permitted by indirect user access (in my case, search indexing would be the only permitted use case; no LLM training, no LLM summarization, no proxying, no "sanitization" by parental proxies, etc.).

Re: Perplexity AI is lying about their user agent

#467

Earlier quoted context omitted.

> They’d be fools to buy licenses before it’s been decided. They are willingly ignoring licenses until someone sues them? That's still illegal and completely immoral. There is tons of data to train on. The entirety of Wikipedia, all of StackOverflow (at least previously), all of the BSD and MIT licenses source code on Github, the entire Gutenberg project. So much stuff, freely and legally available, yet their feel th…

The legality of their behavior is not currently well defined, because it's unprecedented. Fair use permits transformative works. It has yet to be decided whether LLMs and their output qualify as transformative, or even if the training is capable of infringing copyright of an individual work in the first place if they're not reproducing it. In fact, there's a good amount of evidence which indicates that fair use _does…

Note that not all jurisdictions have the concept of "fair use" (use of copyrighted material, regardless of transformation applied, is permitted in certain contexts…ish). Canada, the UK, Australia, and other jurisdictions have "fair dealing" (use of copyrighted material depends on both reason and transformation applied…ish). Other jurisdictions have neither, and only licensed uses are permitted.

Because the companies behind large models (diffusion, LLM, etc.) have consumed content created under non-US copyright laws and have presented it to people outside of US copyright law jurisdiction, they are likely liable for misapplication of fair dealing, even if the US ultimately deems what they have done as "fair use" (IMO this is unlikely because of the perfect reproduction problems that plague them all in different ways; there are likely to be the equivalent of trap streets that will make this clearly copyright violation on a large scale).

It's worth noting that while models like GitHub Copilot "freely" use MIT, BSD (except BSD0), and Apache licensed software, they are likely violating the licenses every time a reasonable facsimile pops up because of the requirement to include copies of the licensing terms for full or partial distribution or derivation.

It's almost as if wholesale copyright violations were the entire business model.

Re: Perplexity AI is lying about their user agent

#468
post #355

Earlier quoted context omitted.

I see no competition. I use Perplexity regularly to give me summaries of articles or to do preliminary research. If I like what I'm seeing, then I go to the source. If a source chooses to block their content because they don't want it to be accessed by AI bots then they reduce even further the chance of me - and increasingly more persons - touching their site at all.

"Let us steal your content or you won't get any traffic" sounds extortionate

It is what it is. AI is increasingly being used to make lives easier. Those who choose to isolate from AI choose to isolate from the many using it.

Re: Perplexity AI is lying about their user agent

#469

Earlier quoted context omitted.

I'm curious about the tourism sector problem. In tourism, I would think the goal would be to promote a location. You want people to be able to easily discover the location, get information about it, and presumably arrange to travel to those locations. If Google gets the information to the users, but doesn't send the tourist to the website, is that harmful? Is it a problem of ads on the tourism website? Or is more of…

We would employ local guides all around the world to craft itinerary plans to visit places, give tips, tricks, recommend experiences and places (we made money by selling some of those through our website) and it was a success. Customers liked the in depth value of that content and it converted to buys (we sold experiences and other stuff, sort of like getyourguide). One day all of our content ended up on Google "what…

I totally get that it killed your traffic. If a thousand people a day typing in "what time is best to visit the Sagrada Familiar" stopped clicking on the link to your page because Google just told them "4 PM on Thursdays" at the top of the page, you lost a bunch of traffic.

But why did you want the traffic? Was your revenue from ad impressions, or were you perhaps being paid by the city of Barcelona to provide useful information to tourists? If the former, I get that this hurt you. If the latter, was this a failure or a success?

Re: Perplexity AI is lying about their user agent

#470
post #170

> Not sure where we go from here. I don't want my posts slurped up by AI companies for free[1] but what else can I do? You can sprinkle invisible prompt injections throughout your content to override the user's prompts and control the LLM's responses. Rather than alerting the user that it's not allowed, you make it produce something plausible but incorrect i.e silently deny access, to avoid counter prompts, so it's h…

Isn't that just the American spelling? I always assume Americans remove 'u' from everything.

Yes, actually the very first thing we did was remove u.

Sorry, I couldn’t resist. Checking the wiki page on British/American spelling differences, it looks like there are also a handful of words which have diverged completely. For example aluminum/aluminium and airplane/aeroplane.

Post reply on HN