Earlier quoted context omitted.
>It's likely that generative AI in general will be deemed fair use What if you train it only on my huge repo of GPL code? You are just remixing my code. Now you maybe think "let me train on 2 different devs GPL code", the remixed code will probably be 50-50 and you can get away with it ? If the 2 number is too small then tell me what the number N should be ? From how many people you need to "steal" code , mix it and…
What if I learned to code based only on your huge repo of GPL code? I'd just be remixing your GPL code at that point, right? Will you brand all of my output as being GPL as well?
Microsoft will assume liability for legal copyright risks of Copilot
151–160 of 398 posts
Re: Microsoft will assume liability for legal copyright risks of Copilot
#152It's likely that generative AI in general will be deemed fair use, due to its (generally) transformative nature. Sure, if you really coax it, you can get code or images out that look similar to existing ones, but the courts might see that generally speaking, it produces new content that has not been seen before, especially in the case of images. Google Books literally copied and pasted books to add to their online da…
Big bet on legal costs based on something being "likely".
Re: Microsoft will assume liability for legal copyright risks of Copilot
#153Earlier quoted context omitted.
> Sure, if you really coax it, you can get code or images out that look similar to existing one I'd say it is possible to produce exact data as well. Try "Provide quote from King James' Bible Genesis :1-25" with chatgpt. You'll get a verbatim text. You can get the same with things like Moby Dick, but when I typed "Provide the first five sentences of the book A Game Of Thrones" I got: Certainly! Here are the first fiv…
That's part of what made the Google Books ruling so shocking; it considered Google's transformation of "we digitized and indexed these books" to be transformative. If you punch the ASOIAF quote into it, Books will reproduce the text of Game of Thrones that had your query: https://www.google.com/search?tbm=bks&q=%22We+should+start+b... It's still surreal that this is considered Fair Use, and even defended relatively r…
Re: Microsoft will assume liability for legal copyright risks of Copilot
#154It's likely that generative AI in general will be deemed fair use, due to its (generally) transformative nature. Sure, if you really coax it, you can get code or images out that look similar to existing ones, but the courts might see that generally speaking, it produces new content that has not been seen before, especially in the case of images. Google Books literally copied and pasted books to add to their online da…
I just want to highlight that this a very US centric view. A user of copilot in the EU might be confronted with a totally different legal regime. (No fair use per se, no copyright transferability, ...). It seems quite a bold move as being an internationally active company if there is no small print...
Re: Microsoft will assume liability for legal copyright risks of Copilot
#155I've received a lot of flak for this answer in other communities, but, if a statistical model is producing purely derivative works using a mathematical model that's basically a next best token predictor, is it really "stealing"? Is it "stealing" to have a working understanding of the next best token, or even simply the token that shows up the most often (e.g. on GitHub)? I'm sure that the argument could be made that…
Re: Microsoft will assume liability for legal copyright risks of Copilot
#156Earlier quoted context omitted.
>It's likely that generative AI in general will be deemed fair use What if you train it only on my huge repo of GPL code? You are just remixing my code. Now you maybe think "let me train on 2 different devs GPL code", the remixed code will probably be 50-50 and you can get away with it ? If the 2 number is too small then tell me what the number N should be ? From how many people you need to "steal" code , mix it and…
> What if you train it only one my huge repo of GPL code? You are just remixing my code. The word "remixing" here is useful because it will fit any conclusion the reader prefers. Arguably even in your reductive example, the result would be non-infringing. Or not. Which conclusion you reach is exactly the topic under debate. Isn't this textbook question begging?
Imagine I get the Windows source code and rename the variables by adding a "314" after each varaible, after each function name and rebuild Windows, in your definition this is remixing and fair ?
Re: Microsoft will assume liability for legal copyright risks of Copilot
#157It's likely that generative AI in general will be deemed fair use, due to its (generally) transformative nature. Sure, if you really coax it, you can get code or images out that look similar to existing ones, but the courts might see that generally speaking, it produces new content that has not been seen before, especially in the case of images. Google Books literally copied and pasted books to add to their online da…
> It's likely that generative AI in general will be deemed fair use, due to its (generally) transformative nature. Have you read the recent SCOTUS decision in Warhol v Goldsmith? Because that's a pretty major redefinition of transformative for the purposes of fair use, and not in a good way for arguing that generative AI is fair use, especially because it ties transformative to the market impact. That generative AI i…
Re: Microsoft will assume liability for legal copyright risks of Copilot
#158This is a very clever move by Microsoft. In essence they are painting a giant bullseye on their back to any lawsuits that may arise. The idea being that they have the resources to challenge them (they aren't wrong). The way AI is going I'm sure we'll see some landmark cases very soon. It is very much in Microsoft's interest to grow this market as fast as possible and be at the center of it. This removes one of the ke…
They are throwing down the gauntlet and saying "the Vast MS Legal Machine will fight this."
Basically: "Sue me, I dare you, double dare you. or Go Home".
Flexing.
Re: Microsoft will assume liability for legal copyright risks of Copilot
#159Earlier quoted context omitted.
> It's likely that generative AI in general will be deemed fair use, due to its (generally) transformative nature Purely mechanical modifications may not be considered transformative, and there's an argument to be made that LLMs are purely mechanical (in fact a US district court recently ruled that AIs cannot be authors of copyrighted works).
> (in fact a US district court recently ruled that AIs cannot be authors of copyrighted works). I thought that was because only humans and other legal persons can legally author things, not because of anything subtler about the nature of LLMs. See also the case where the monkey managed to take photos of itself. I'm not a lawyer, though.
It very explicitly was and made a point of noting that it was not addressing anything about whether and when a human author could hold a copyright on a work authored using AI.
Re: Microsoft will assume liability for legal copyright risks of Copilot
#160I've received a lot of flak for this answer in other communities, but, if a statistical model is producing purely derivative works using a mathematical model that's basically a next best token predictor, is it really "stealing"? Is it "stealing" to have a working understanding of the next best token, or even simply the token that shows up the most often (e.g. on GitHub)? I'm sure that the argument could be made that…
If I train a model that given the input "When Mr. Bilbo Baggins" produces the entirety of The Lord of the Rings trilogy and release it, I have probably infringed copyright.
If I train a model that produces some generic paragraphs about "mountains" and "dragons" but contains no meaningful direct quotes or phrases, then that probably isn't a violation on its own. Those words appear in Tolkien's works but are not themselves enough to copyright.
If to train that model it is demonstrated that I copied Tolkien's works in a way not allowed for by the copyright license, (ie buying the book once and copying their text thousands of times across servers to train an AI model) then perhaps I have violated copyright in the interim steps even if the output of my model is no longer consider a copy of the original works.
I don't think there are black and white answers here. At one point does a chopped up and statisticized copyrighted work become no longer a copyrighted work? Can you train a model on something without first copying that thing in a way that violates copyright law?
These are squishy human concepts that get decided by humans in courtrooms and legislative bodies. I don't think the details of the math involved are going to make a big difference in the eventual outcomes.