Live data from Hacker News

Microsoft will assume liability for legal copyright risks of Copilot

blogs.microsoft.com

331–340 of 398 posts

Re: Microsoft will assume liability for legal copyright risks of Copilot

#331

Earlier quoted context omitted.

Is the contrivedness relevant to the legal question? It shows the model contains the copyrighted content and can reproduce it on demand.

My brain contains loads of copyrighted info. And if I exactly reproduce it from memory, it's copyright infringement. But if I come up with my own work, even if using that copyrighted info to learn from, it isn't infringement.

I don't understand why people keep comparing humans and computers. The law does not treat machinery equal to a human.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#332
post #320
post #291

Earlier quoted context omitted.

> This use is found to be illegal This being the real hurdle. With Microsoft money behind the defense, only megacorps can win.

Microsoft has lost legal battles against non-megacorps in the past. I remember some guy representing himself and winning some dispute over shrink wrap licenses and student discounts.

That's badass. Where can I read about that court case?

Re: Microsoft will assume liability for legal copyright risks of Copilot

#333
post #276

Earlier quoted context omitted.

Napster and the Pirate Bay struggled because the vast majority of content was pirated. You would be hard pressed to say a significant minority of generative ai has any copyright issues, much less copyright issues as blatant as straight piracy.

> You would be hard pressed to say a significant minority of generative ai has any copyright issues, much less copyright issues as blatant as straight piracy. Where generative AI ingests copyrighted works in order to work and bases its output on it, then it is copyright infringement, equivalent to 'straight piracy' of all that it ingested, unless it's deemed fair use. What Google does with its search engine, for exam…

I suspect that your extreme position must come with a sincere belief that nothing even close to human intelligence will ever be achieved. Imagine a robot with an AI brain. It would have to be blindfolded because just by learning to recognize, or even just viewing, the label on a can of Coke it would have “copied” it and become “illegal,” especially if it was capable of sketching it on demand. Any kind of intelligence cannot even simply view or listen to the world without encountering something IP-encumbered.

Learning, by human or machine, means extracting a copy of the essence of something and yes, storing that essence in a lossy way. It seems like learning from copyright-encumbered material ought to either be illegal for both, or legal. I know which world I would rather live in.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#334

Earlier quoted context omitted.

Without a myriad of dumbasses like me being able to commit to Microsoft vs Github, I'd assume Microsoft's average is better than Github's.

That is a... bold assumption to make. Not just for Microsoft but for any large corporation.

> That is a... bold assumption to make. Not just for Microsoft but for any large corporation.

I dunno; the average project on github isn't code-reviewed, while all the projects at Microsoft are.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#335
post #194

Let Microsoft first publish a Copilot model that's trained on the internal codebases of Azure, Windows and Office. That's the only way Microsoft can convince me that they truly believe Copilot is non-infringing technology.

I suspect Microsoft would earn more money by doing this. Their own engineers would get productivity boosts - with copilot already being familiar with data structures, code style, etc. would be a big boost to accuracy. But also, third party code would end up being more similar. Code style of the whole world would be pushed towards 'Microsoft style', which probably makes hiring easier, less training time for engineers,…

> Code style of the whole world would be pushed towards 'Microsoft style'

Yes, that's exactly what the world needs, more software like Teams.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#336

Earlier quoted context omitted.

My brain contains loads of copyrighted info. And if I exactly reproduce it from memory, it's copyright infringement. But if I come up with my own work, even if using that copyrighted info to learn from, it isn't infringement.

I don't understand why people keep comparing humans and computers. The law does not treat machinery equal to a human.

You can't really say that. All this needs to be tested in court and see what definitions of what end up winning and setting precendent.

It can go either way.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#337
post #194

Let Microsoft first publish a Copilot model that's trained on the internal codebases of Azure, Windows and Office. That's the only way Microsoft can convince me that they truly believe Copilot is non-infringing technology.

I suspect Microsoft would earn more money by doing this. Their own engineers would get productivity boosts - with copilot already being familiar with data structures, code style, etc. would be a big boost to accuracy. But also, third party code would end up being more similar. Code style of the whole world would be pushed towards 'Microsoft style', which probably makes hiring easier, less training time for engineers,…

> is probably irrelevant when outsiders can already decompile binaries and learn far more.

most, if not all microsoft products can have their sources be available for viewing, if you are one of those vip development partners. microsoft doesn't really have any secret source (pardon the pun) of which the leaking would undo their value proposition.

In fact, if microsft opened up their system a bit more, they might even gain some PR or mindshare, and have no effect on, if not increase, their bottom line.

Re: Microsoft will assume liability for legal copyright risks of Copilot

#340
post #296
post #276

Earlier quoted context omitted.

Napster and the Pirate Bay struggled because the vast majority of content was pirated. You would be hard pressed to say a significant minority of generative ai has any copyright issues, much less copyright issues as blatant as straight piracy.

I’m not convinced any of the output of these generative AI is free from copyright issues. Consider, a ROT13 copy of a book may at first glance look nothing like the original, but distributing digital copies would be clear copyright infringement. Feature extraction is literally a form of lossy compression. You can prod DALEE to make obvious copies of some of the works it was trained on, but even seemingly novel images…

Copyright is not cooties. For something to be infringing it has to be beyond the de minimis threshold. It’s not enough to show that a copyrighted work influenced another work, there needs to be some substantial level of copying.

This music industry has been going through exactly this for the last few years and the courts have recognized that the creative process necessarily involves copying and that a small amount of copying is not infringement.

Post reply on HN