Quibble with the headline-- I don't see a lie by Perplexity, they just aren't complying with a voluntary web standard.[1] [1] https://en.m.wikipedia.org/wiki/Robots.txt
Perplexity AI is lying about their user agent
31–40 of 555 posts
Re: Perplexity AI is lying about their user agent
#32OpenAI scraped aggressively for years. Why should others put themselves behind an artificial moat? If you want to block access to a site, stop relying on arbitrary opt-in voluntary things like user agent or robots.txt. Make your site authenticated only, that’s literally the only answer here.
Re: Perplexity AI is lying about their user agent
#33I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.
Is it actually retrieving the page on the fly though? How do you know this? Even if it were - it’s not supposed to be able to.
Re: Perplexity AI is lying about their user agent
#34I think we need to define the difference between a software (my browser) returning some web content and another software (an agent) doing the same thing.
Re: Perplexity AI is lying about their user agent
#35GDPR doesn't preclude anyone from scraping you. In fact, scraping is not illegal in any context (LinkedIn keeps losing lawsuits). Using copyrighted data in training LLMs is a huge grey area, but probably not outright illegal and will take like a decade (if not more) before we'll have legislative clarity.
Re: Perplexity AI is lying about their user agent
#36Earlier quoted context omitted.
Or detect the LLM and serve up an LLM rewritten version of the page. That way you feed it poisonous garbage.
I really like this idea. Someone needs to implement this. I'm not sure what the ideal poison would be. Randomly constructed sentences that follow the basic rules of grammar?
Mix up the verbs, add/delete "not", "but", "and".
Change names.
Re: Perplexity AI is lying about their user agent
#37I don’t think we should lump together “AI company scraping a website to train their base model” and “AI tool retrieving a web page because I asked it to”. At least, those should be two different user agents so you have the option to block one and not the other.
If an AI agent is performing a search on behalf of a user, should its user agent be the same as that user’s?
Does anyone actually do this, though?
Re: Perplexity AI is lying about their user agent
#38Earlier quoted context omitted.
It's been debated at length, but to make it short: piracy is not theft, and everyone in the LLM space has been taking other people’s content and so far getting away with it (pending lawsuits notwithstanding).
Can’t wait for OpenAI to settle with The New York Times. For a billion dollars no less.
Re: Perplexity AI is lying about their user agent
#39OpenAI scraped aggressively for years. Why should others put themselves behind an artificial moat? If you want to block access to a site, stop relying on arbitrary opt-in voluntary things like user agent or robots.txt. Make your site authenticated only, that’s literally the only answer here.
Agree - the first movers who scraped before changes to websites terms and robots files shouldn’t get an unfair advantage. That’s overall bad for society in terms of choice and competition
Re: Perplexity AI is lying about their user agent
#40Earlier quoted context omitted.
Or detect the LLM and serve up an LLM rewritten version of the page. That way you feed it poisonous garbage.
I really like this idea. Someone needs to implement this. I'm not sure what the ideal poison would be. Randomly constructed sentences that follow the basic rules of grammar?
ChatGPT, write a short story that warns about the dangers of artificial intelligence stealing people's intellectual property, from the perspective of a hamster in a cage beside a computer monitor.