Live data from Hacker News

An open source lawyer’s view on the copilot class action lawsuit

katedowninglaw.com

61–70 of 182 posts

Re: An open source lawyer’s view on the copilot class action lawsuit

#61

A hypothetical question: imagine a filmmaker, who had studied a lot of obviously copyrighted movies by famous renowned directors. This means he has trained his neural network using their copyrighted licensed content. Does he breach copyright when he composes and films a scene? Are visual quotes copyright theft? Homages? Did George Lucas infringe copyright when he was borrowing compositions from "Triumph of the will"?

Your magic box is not a film maker and the inputs you are encoding with it are verbatim file content. Said content belonging to someone else.

Please study the series of events that unfolded in the music industry after folk begun incorporating recordings made by other artists in their own work and proceeded to sell the result.

Spoiler: The deeply nuanced question of feeding a mechanical recording through a series of complex physical and mathematical apparatus and whether that constituted a transformational creative act did not come up during the proceedings or final judgements!

Re: An open source lawyer’s view on the copilot class action lawsuit

#62
post #45

Earlier quoted context omitted.

There has never been more support for tightening and enforcing copyright than there is today. This is very unlikely to change due to megacorps like Microsoft, Disney, Apple et.al. having a massive vested interest to use it to extract maximum profits.

Copyright protection for the rich and powerful, while those who cannot afford armies of lawyers get their stuff stolen by machine learning models. Sounds credible to me.

I’m not rich and powerful but I’m glad Disney is going to destroy MystickInk on principle.

Check out my essay I’ve submitted recently.

Re: An open source lawyer’s view on the copilot class action lawsuit

#63

A hypothetical question: imagine a filmmaker, who had studied a lot of obviously copyrighted movies by famous renowned directors. This means he has trained his neural network using their copyrighted licensed content. Does he breach copyright when he composes and films a scene? Are visual quotes copyright theft? Homages? Did George Lucas infringe copyright when he was borrowing compositions from "Triumph of the will"?

Just because machine learning uses the word “learning” doesn’t mean it “learns” in the same way a human mind does — that analogy is doing a lot of load bearing in your argument, and needs proving why the program’s nature of creative remixing (for lack of a better word) is the same as a human’s. Right now it seems like you’re just reusing the same word for two phenomena we don’t understand, and therefore claiming they’re equivalent.

See Marvin Minsky’s comment regarding “suitcase words”.

Re: An open source lawyer’s view on the copilot class action lawsuit

#64

A hypothetical question: imagine a filmmaker, who had studied a lot of obviously copyrighted movies by famous renowned directors. This means he has trained his neural network using their copyrighted licensed content. Does he breach copyright when he composes and films a scene? Are visual quotes copyright theft? Homages? Did George Lucas infringe copyright when he was borrowing compositions from "Triumph of the will"?

Your magic box is not a film maker and the inputs you are encoding with it are verbatim file content. Said content belonging to someone else. Please study the series of events that unfolded in the music industry after folk begun incorporating recordings made by other artists in their own work and proceeded to sell the result. Spoiler: The deeply nuanced question of feeding a mechanical recording through a series of c…

> Said content belonging to someone else.

Is CoPilot just trained on OSS, or on private repos too?

Re: An open source lawyer’s view on the copilot class action lawsuit

#65
post #15

Earlier quoted context omitted.

Copyright was originally intended to protect the creators of a work. Over many years it has now mostly become a tool for large companies to accumulate rights (on works they didn't create themselves) and monetize them. Maybe a reform is needed, to find a way back to the original purpose.

> Copyright was originally intended to protect the creators of a work. No, it wasn’t. Copyright was originally intended to protect the publishers of a work. It was later transformed to nominally focus on the creators, but even this was lobbied for by publishers in their own self-interest after the old law directly protecting them was allowed to lapse, and because it still had the same net effect since realizing value…

Wrong!

At the point of creation something is granted copyright.

Publishers in literature and music are right assholes who’ve created this system. Little middle men rent seeking.

It does need reform but it is for the creators that’s why it’s tied to the creator and not date of publication. Fix your perspective buckaroo

Re: An open source lawyer’s view on the copilot class action lawsuit

#66

A hypothetical question: imagine a filmmaker, who had studied a lot of obviously copyrighted movies by famous renowned directors. This means he has trained his neural network using their copyrighted licensed content. Does he breach copyright when he composes and films a scene? Are visual quotes copyright theft? Homages? Did George Lucas infringe copyright when he was borrowing compositions from "Triumph of the will"?

Humans are not neural networks, that's just a thesis. Even novelists do not sit all day long in a closed room reading other people's work and then do a collage of what they've read. Otherwise no books would have been written in the first place. Cut the AI off humans' work, let it interact with the real world and see what it produces. It will be nothing. Once (if ever?) an AI is capable of producing an actual original…

> Cut the AI off humans' work, let it interact with the real world and see what it produces. It will be nothing.

That "experiment" could just as well be done on humans, though, cut them off of any work that any human has done before and you may get simple cave paintings, if you're lucky.

Re: An open source lawyer’s view on the copilot class action lawsuit

#68
post #46

Earlier quoted context omitted.

There's a minimum level of complexity and creativity which constitutes a copyright violation. It's up to a legal professional to draw the line, but I believe it can be a single line of code (`i = 0x5f3759df - ( i >> 1 );`) If I saw 100 LOC which was very similar to something which I wrote, AND contained a log statement copied verbatim, it's very easy to imply that the entire piece of code is a derivative work. Let's…

Replicating copyrighted code from the training set only happens 1% of the time, it's the exception not the rule. And when it happens it's usually because the same text appears multiple times in the training set. So it will memorize boilerplate and popular code snippets, not unique stuff. Even a replicated piece of code 100 lines long is no big deal in my opinion, unless it contains some kind of unique thing never see…

I have ~400KLOC changed on GitHub. 1% of the time happens multiple times a day given scale.

Pragmatically, people are already knowingly committing commercially viable copyright violations of my work. I'd rather it wasn't encouraged further by a US-based 'big tech', especially if the people using my code aren't aware that they're doing anything questionable.

Some months, I earn over 100x less from OSS than I would in industry. I don't want people taking advantage any more than I'm comfortable with, especially for commercial purposes.

Re: An open source lawyer’s view on the copilot class action lawsuit

#69
post #45

Earlier quoted context omitted.

There has never been more support for tightening and enforcing copyright than there is today. This is very unlikely to change due to megacorps like Microsoft, Disney, Apple et.al. having a massive vested interest to use it to extract maximum profits.

Copyright protection for the rich and powerful, while those who cannot afford armies of lawyers get their stuff stolen by machine learning models. Sounds credible to me.

I find Copilot most useful for filling out debug statements such as this:

  println(“foo at {:x} is {:?}”, &foo as *const _ as usize, foo);
It almost always writes what I would have. How DARE I steal from open source contributors like that?!

Re: An open source lawyer’s view on the copilot class action lawsuit

#70
post #64

Earlier quoted context omitted.

Your magic box is not a film maker and the inputs you are encoding with it are verbatim file content. Said content belonging to someone else. Please study the series of events that unfolded in the music industry after folk begun incorporating recordings made by other artists in their own work and proceeded to sell the result. Spoiler: The deeply nuanced question of feeding a mechanical recording through a series of c…

> Said content belonging to someone else. Is CoPilot just trained on OSS, or on private repos too?

Just on public repositories, as far as I know, however regardless of license.

There are GPL repositories which force you to open your code, which is one aspect, and there are "source available" repositories, which allows you to see the code, but forbids everything else.

There are a lot of blurry areas about this, and in my opinion, an AI learns like a human is not a solid basis for fair use.

On the other hand, if private repositories are crawled too, this would be very, very bad.

Post reply on HN