Live data from Hacker News

Microsoft will assume liability for legal copyright risks of Copilot

blogs.microsoft.com

321–330 of 398 posts

Re: Microsoft will assume liability for legal copyright risks of Copilot

#321

Earlier quoted context omitted.

Google absolutely has their own internal models that do exactly this. It wouldn't surprise me if Microsoft indeed does have an internal Copilot that is trained on their data, but even on the smallest risk that they leak their code, they wouldn't share that particular model.

What does "absolutely has" mean here? Have you actually heard anything about such internal models?

Why wouldn’t they? Meta does, and they write openly about it

Re: Microsoft will assume liability for legal copyright risks of Copilot

#322
post #296
post #276

Earlier quoted context omitted.

Napster and the Pirate Bay struggled because the vast majority of content was pirated. You would be hard pressed to say a significant minority of generative ai has any copyright issues, much less copyright issues as blatant as straight piracy.

I’m not convinced any of the output of these generative AI is free from copyright issues. Consider, a ROT13 copy of a book may at first glance look nothing like the original, but distributing digital copies would be clear copyright infringement. Feature extraction is literally a form of lossy compression. You can prod DALEE to make obvious copies of some of the works it was trained on, but even seemingly novel images…

I think this is basically the same as the sentiment that there is no such thing as a truly novel idea.

The standard isn’t “I think that looks like an Andy Warhol picture”, it is “That is substantially similar to a specific Andy Warhol piece”. Copyright doesn’t protect style.

> Feature extraction is literally a form of lossy compression.

This is one way think of neural nets, another is that they find the topological space of pictures.

But these are just models of computation, which aren’t especially relevant in the same way that it isn’t relevant what produces an infringing image, just that it is produced.

Which brings me back to my original point: there are a two different barriers for generative ai: is the model itself transformative, and is the primary purpose of the model to generate copyright infringing material.

With respect to the first… I have no idea how someone could argue that the model itself isn’t transformative enough. It isn’t “substantially similar” to any of the works that it is trained on. It might be able to generate things that are “substantially similar”, but the model itself isn’t… it’s just a bunch of numbers and code.

Regarding the second: I have less experience with image models, but I use chatgpt regularly without even trying to violate copyright, and I don’t think I’m alone, so I doubt you could make an argument that llms have a primary purpose of committing copyright infringement.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#323

Earlier quoted context omitted.

Take a look at the prompts people use in these examples. They are always so contrived. Sure if you ask it to "Take this function exactly as it is from this file at this repo and output it without changes" it can do that.

Is the contrivedness relevant to the legal question? It shows the model contains the copyrighted content and can reproduce it on demand.

Yes. Courts will generally assign blame to whoever did the thing that caused a breach of the law, which in this case is the user.

In other words, law isn’t a programming language.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#324
post #276

Earlier quoted context omitted.

Napster and the Pirate Bay struggled because the vast majority of content was pirated. You would be hard pressed to say a significant minority of generative ai has any copyright issues, much less copyright issues as blatant as straight piracy.

> You would be hard pressed to say a significant minority of generative ai has any copyright issues, much less copyright issues as blatant as straight piracy. Where generative AI ingests copyrighted works in order to work and bases its output on it, then it is copyright infringement, equivalent to 'straight piracy' of all that it ingested, unless it's deemed fair use. What Google does with its search engine, for exam…

>Where generative AI ingests copyrighted works in order to work and bases its output on it, then it is copyright infringement

This is an absurd standard. Is it copyright infringement when a human "ingests" copyrighted work and bases their output on it? Because that's commonly called inspiration and is how every artist creates their work - through experiencing other works and using that cumulative inspiration to form their own product.

Copyright infringement is already ridiculously restrictive as it is, this proposal not only fundamentally misunderstands how generative AI works but penalises AI for doing what humans do everyday.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#325
post #296

Earlier quoted context omitted.

I’m not convinced any of the output of these generative AI is free from copyright issues. Consider, a ROT13 copy of a book may at first glance look nothing like the original, but distributing digital copies would be clear copyright infringement. Feature extraction is literally a form of lossy compression. You can prod DALEE to make obvious copies of some of the works it was trained on, but even seemingly novel images…

The foundational models most coding models are built on May have comments and code in them. They’re almost certainly built on a number of legal violations, including copyright infringement. I have details in “Proving Wrongdoing” section here: https://www.heswithjesus.com/tech/exploringai/index.html I’ve also seen GPT spit out proprietary content word for word that’s not licensed for commercial use that I’m aware of.…

The problem with this line of thinking is that a person can also cut and paste code that they don’t have a license to use… but until they do, they haven’t done anything wrong by reading the code.

So either we carve out an explicit exception that machines aren’t allowed to do things that are remarkably similar to what humans do… which would be a massive setback for AI in the US.

Or we agree that generative models are subject to the same rules that humans are — they can’t commit copyright infringement, but are able to appropriately consume copyrighted material that a human would be able to consume.

The second option seems to me to be much simpler, nicer, and more appropriate than the first.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#326

Earlier quoted context omitted.

> It's likely that generative AI in general will be deemed fair use, due to its (generally) transformative nature. This isn't how "fair use" works, in the sense that there can never be a blanket assurance like that. Also, whether the result is "transformative" is just one of many factors (see audio sampling/remixing).

It's not how fair use works now , but many things about copyright law will have to change radically over the next few years. There's too much at stake.

This is what people on the "this is copyright infringement" side don't understand. Even if it somehow is copyright infringement by the standards of today's law, those laws will inevitably change in the near future. Generative AI is far, far too lucrative and convenient to society for it to be crippled by obsolete copyright infringement laws formed in a time where generative AI was a thing of science fiction novels.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#327

Earlier quoted context omitted.

Billions of dollars in investments that will partially benefit lawmakers?

Well, content creators and publishers have trillions.

Content creators and publishers are also the primary ones who are using generative AI. They aren't a united monolith against generative AI.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#328
post #280

Earlier quoted context omitted.

You're getting a lot of pushback, but the EU seems to agree with you: https://creativecommons.org/wp-content/uploads/2021/12/CC-St... https://www.notion.so/DSM-Directive-Implementation-Tracker-3... https://eur-lex.europa.eu/eli/dir/2019/790/oj The TDM4 copyright exception allows datasets to be created consisting of copyrighted works, as long as there is a mechanism for rightsholders to opt out. This seems like the be…

> as long as there is a mechanism for rightsholders to opt out. I really don't like this--opt-out never works because the scale advantages are backwards. It places the burden in the wrong place. The aggregators should have to get opt-in. Look at YouTube. Because of "opt-out", lots of people monetize content that they have no right to and it's up to the original author to have to fight the scale of a zillion uploaders…

>I really don't like this--opt-out never works because the scale advantages are backwards. It places the burden in the wrong place. The aggregators should have to get opt-in.

>Look at YouTube. Because of "opt-out", lots of people monetize content that they have no right to and it's up to the original author to have to fight the scale of a zillion uploaders. Only the biggest entities can do that.

>YouTube (and everybody else) should have to assert "You, the uploader, own this content" when they ingest it. Nothing else works.

Thankfully, in a rare turn of fate, capital will be on the side of the laissez-faire instead of the stringent anti-copyright-infringers for once. You do not own the rights to material created by a generative AI.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#330

Earlier quoted context omitted.

Take a look at the prompts people use in these examples. They are always so contrived. Sure if you ask it to "Take this function exactly as it is from this file at this repo and output it without changes" it can do that.

Is the contrivedness relevant to the legal question? It shows the model contains the copyrighted content and can reproduce it on demand.

My brain contains loads of copyrighted info. And if I exactly reproduce it from memory, it's copyright infringement. But if I come up with my own work, even if using that copyrighted info to learn from, it isn't infringement.
Post reply on HN