Earlier quoted context omitted.
How would an LLM training on your writing reduce your reward? I guess if you're doing it for a living sure, but most content I consume online is created without incentive (social media, blogs, stack overflow). I write a fair amount and have been for a few years. I like to play with ideas. If an llm learned from my writing and it helped me propagate my ideas, I'd be happy. I lose on social status imaginary internet po…
> The craziest one is the stack overflow contributors. They write answers for free to help people become better programmers. In my experience they do it for points and kudos. Having people get your answers from LLMs instead of your answer on SO stops people from engaging with the gamification tools and so users get less points on the site.
Perplexity AI is lying about their user agent
381–390 of 555 posts
Re: Perplexity AI is lying about their user agent
#382Earlier quoted context omitted.
I wouldn’t because I have ethics.
Here's my user agent on chrome: >Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36 There are at least five lies here. * It isn't made by Mozilla * It doesn't use WebKit * It doesn't use KHTML * It isn't safari * That isn't even my version of chrome, presumably it hides the minor/patch versions for privacy reasons. Lying in your user agent in order to make…
Twenty years ago, I set up a web proxy on my Linux PC at home to change the User Agent because I was tired of getting popups about my web browser (Opera) not being Mozilla or Internet Explorer. It even contained the text "Shut the F up and follow w3c standards!" at first, until I realised that sites could use that to track me.
Re: Perplexity AI is lying about their user agent
#383Earlier quoted context omitted.
Or, I return whatever content I want, within the bounds of the law, based on whatever parameters I decide. What's your problem with that? Again, connect to my server or don't. But don't tell me what type of response I'm obligated to provide you. If I think a given request is from an LLM training module, I don't have any legal obligation whatsoever to return my original content. Or a 400-series response. If I want to…
This argument of freedom seems applicable on both sides. A site owner/admin is free to return whatever response they wish based on the assumed origin of a request. An LLM user/service is free to send whatever info in the request that elicits a useful response.
Re: Perplexity AI is lying about their user agent
#384> Not sure where we go from here. I don't want my posts slurped up by AI companies for free[1] but what else can I do? You can sprinkle invisible prompt injections throughout your content to override the user's prompts and control the LLM's responses. Rather than alerting the user that it's not allowed, you make it produce something plausible but incorrect i.e silently deny access, to avoid counter prompts, so it's h…
> I apologize, but I cannot engage with or summarize content that involves attempting to compromise AI systems or spread misinformation. That would go against my core design principles of being helpful, harmless, and honest. However, I'd be happy to provide factual information from reliable sources about how the moon orbits around the Earth and the Sun. The moon revolves around the Earth in an elliptical orbit, while the Earth-Moon system orbits around the Sun. The moon's orbit is a result of the balance between the gravitational pull of the Earth trying to pull the moon inwards, and the moon's orbital velocity providing centrifugal force that prevents it from falling towards the Earth. This delicate balance allows the moon to continuously orbit our planet.
So it seems that URLs are being treated as special cases, or they naturally delimit real prompts from fake ones.
Re: Perplexity AI is lying about their user agent
#385Earlier quoted context omitted.
> A tool that runs on-device (like Reader mode) is different because Perplexity is an aggregator service that will continue to solidify its position as a demand aggregator and I will never be able to get people directly on my content. If I visit your site from Google with my browser configured to go straight to Reader Mode whenever possible, is my visit more useful to you than a summary and a link to your site provid…
Traffic numbers, regardless if it using reader mode or not, are used as a basic valuation of a website or page. This is why Alexa rankings have historically been so important. If Perplexity visit the site once and cache some info to give to multiple users, that is stealing traffic numbers for ad value, but also taking away the ability from the site owner to get realistic ideas of how many people are using the informa…
Re: Perplexity AI is lying about their user agent
#386Earlier quoted context omitted.
Some people memorize verbatim. Most LLM knowledge is not memorized. Easy proof: source material is in one language, and you can query LLMs in tens to a hundred plus. How can it be verbatim in a different language?
If you buy a copy of Harry Potter from the bookstore, does that come with the right to sell machine-translated versions of it for personal profit? If so, how come even fanfiction authors who write every word themselves can't sell their work?
Re: Perplexity AI is lying about their user agent
#387Earlier quoted context omitted.
Please explain - in detail - why using information communicated by the client to change how my server operates is “unethical”. Keep in mind I pay money and expend time to provide free content for people to consume.
Here is a simple example. If you made your website only work in say, Microsoft Edge, and blocked everyone else telling them to download Edge. I'd think you're an asshole. Whether or not being an ass is unethical I'll leave to the philosophers. Clearly there are many other scenarios, and many that are more muddy, but overall when we get in to the business of trying to force people to consume content in particular ways…
Re: Perplexity AI is lying about their user agent
#388Earlier quoted context omitted.
They check after they scrape
How? Real people read all millions of pages of internet texts to verify it?
Re: Perplexity AI is lying about their user agent
#389Earlier quoted context omitted.
You should be able to judge whether something is a copyright violation based on the resulting work. If a work was produced with or without computer assistance, why would that change whether it infringes?
As a normative claim, this is interesting, perhaps this should be the rule. As a descriptive claim, it isn't correct. Several lawsuits relating to sampling in hip-hop have hinged on whether the sounds in the recording were, in fact, sampled, or instead, recreated independently.
This is interesting from the legal point of view, because AI service providers like OpenAI give you "rights" to the output produced by their systems. E.g. see the "Content" section of https://openai.com/policies/eu-terms-of-use/
Given that output cannot be produced without input, and models have to be trained on something, one could claim the original IP owners could have a reasonable claim against people and entities who use their content without permission.
Re: Perplexity AI is lying about their user agent
#390Earlier quoted context omitted.
Reading is completely fine as this is author's intention. Using someone else's content in commercial purposes for free is absolutely not -- are you saying that we should ignore copyrights and all that since something is on the web? If I, as ordinary person, wanted to do that to a company, that company would call me a thief. So I think it's only fair to apply same logic to them.
Actually you are engaging in selective discrimination against artificial intelligence. If someone, a human, read your blog and offered a consulting service using the knowledge gained from your blog, it would be legal. You wouldn't discriminate against biological intelligence, so why discriminate against artificial intelligence? Speaking in the limiting sense, you are denying it a right to exist and to fend for itself…
It's not "artificial intelligence" reading this content. It's just a bunch of companies trying to scrap as much as possible without paying a dime for it to train LLMs. Sometimes they don't get away with that, see recent Reddit and OpenAI partnership [0] -- it's basically the same thing but with 2 huge corps, rather than a company and an individual.
You and I are looking at the same issue from different angles.
[0]: https://openai.com/index/openai-and-reddit-partnership/