Live data from Hacker News

GitHub is sued, and we may learn something about Creative Commons licensing

scholarlykitchen.sspnet.org

201–210 of 475 posts

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#201

I don’t understand the author’s position in this article. It spends a long time talking about details of the licenses, but I can’t see any way the suit will actually be about licenses, because if it’s about licenses then it seems patently obvious to me that GitHub will lose very quickly, because they have undoubtedly violated the terms of the licenses. As I see it, the only leg GitHub can possibly stand on is the “fa…

IMO fair use is still not a strong argument for Microsoft. They commercialized the product and made money out of it. Fair use is only allowed if the work you're doing is purely for the greater good. I might be wrong though, IANAL.

You are wrong, though public interest is certainly the basis of the purpose of fair use doctrine. But the simplest way of demonstrating that fair use still allows commercialisation is probably this example: search engines absolutely depend on fair use if they include any content from the linked pages. (And some countries have even tried to call the act of linking copyright infringement, though they’ve tended to back off at least a little, to requiring at least the title or other content for it to be infringement, and not just the URL.)

https://en.wikipedia.org/wiki/Fair_use, lots of good reading there.

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#202
post #51

Earlier quoted context omitted.

As well, Google is, because of these things, somewhat acting as a library. An libraries are very special entities.

Not really. Google won because Google Books was not actually a new concept; someone else had already built a book search engine the same way Google did, also got sued by the Authors Guild, and also prevailed. The only thing different about Google Books was that it'd give you two pages worth of excerpt out of the book. So it was very easy for a court to extend the fair use logic that they had already weaved into the l…

> I still think "training is fair use" still has a leg to stand on, though

If that's the case, we need to serious re-consider how we reward Open Source as a society (I think that would be fantastic anyway!) -- we have people producing knowledge and others profiting directly from this material, producing new content and new code that's incompatible with the original license.

You make GPL code, a make an AI that learns from GPL code, shouldn't its output be GPL licensed as well?

I think meanwhile the most reasonable solution is that an AI should always produce content compatible with the training material licenses. So if you want to use GPL training sets, you can only use that to create GPL-compatible code. If you use public domain (or e.g. 0BSD?) training sets, you can produce any code I guess.

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#203
post #67

Earlier quoted context omitted.

I think we're facing a copyright extinction event. The whole concept is out of touch with the new reality - when you can generate 100 variations for your text, code or image with the click of a button, what does it even mean to hold copyright over the original? "In the style of" killed copyright in 2022.

How many people want to read AI generated text in the style of Lord of the Rings vs how many people want to read Lord of the Rings?

If an AI can generate an infinite number of stories that take place in the LotR world and do it well and faithfully in the style of JRR Tolkien then I would happily read it. You can only read the trilogy and The Hobbit so many times.

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#204
post #124

Earlier quoted context omitted.

[flagged]

I'm sorry but Copyright is very much the law of the land no matter how often you post your links.

Physical slavery was once the law of the land too. I like to think I would have been on the right side of history at that time, as well.

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#205
post #186

I'm still baffled as to why people treat Github like a public library despite being owned by what was at one time the greatest enemy of free and open source software in existence. Not saying they haven't changed their tune somewhat, but a library owned by Barnes and Noble is going to have very different incentives than an actual library. Made all the more silly by the fact that it's Git. You could just host it yourse…

That's because there's no feedback loop of what you're saying, in the active lives of the people interfacing with GitHub. Consider an example of ingesting poison. If the poison tastes bad, I'll be sure to spit it out immediately, either by involuntary disgust, or because I associate with that negative feeling of being poisoned, something I don't want, so I react. But what if the poison tastes good? And what if it not only tastes good, it actually rewards me for ingesting it, in some way? People tell me it's poison, it might not say so on the label, and many are also ingesting it. Is it even believable that it's poison, given that I don't experience the negatives at all?

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#206

I don’t understand the author’s position in this article. It spends a long time talking about details of the licenses, but I can’t see any way the suit will actually be about licenses, because if it’s about licenses then it seems patently obvious to me that GitHub will lose very quickly, because they have undoubtedly violated the terms of the licenses. As I see it, the only leg GitHub can possibly stand on is the “fa…

IMO fair use is still not a strong argument for Microsoft. They commercialized the product and made money out of it. Fair use is only allowed if the work you're doing is purely for the greater good. I might be wrong though, IANAL.

It might be a problem to treat it as copyright. Copyright applies to reproduction, distribution, public performance... if I go to a library or bookstore and I read books and look at their covers, copyright does not apply. Would an android that walks around learning things be subject to copyright? To what extent does it need a body and mobility to be more like a person and less as a scraper?

It might seem stupid, but I worry that if copyright begins applying to "mining" then the next thing is that it applies to humans watching things.

Of course, if an AI re-creates copyrighted content, copyright should apply. Just like it applies when I redraw and sell the Mona Lisa, but not when I store it in my memory. I would pass on the responsibility to users. I don't fear my use of Github Copilot because it's far from infringing any reasonable copyright... then again, I'm assuming the most likely way to infringe copyright with GPT is to use a prompt that almost explicitly requests it.

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#207
post #186

I'm still baffled as to why people treat Github like a public library despite being owned by what was at one time the greatest enemy of free and open source software in existence. Not saying they haven't changed their tune somewhat, but a library owned by Barnes and Noble is going to have very different incentives than an actual library. Made all the more silly by the fact that it's Git. You could just host it yourse…

GitHub built goodwill over the years. There were many controversies, but there were also many die-hard fans. That didn't evaporate overnight. Microsoft bought GitHub (and minted 3 billionaires in the process) specifically to acquire that goodwill and monetize it.

> acquire that goodwill and monetize it

Embrace

Extend Extinguish

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#208

Earlier quoted context omitted.

Not really. Google won because Google Books was not actually a new concept; someone else had already built a book search engine the same way Google did, also got sued by the Authors Guild, and also prevailed. The only thing different about Google Books was that it'd give you two pages worth of excerpt out of the book. So it was very easy for a court to extend the fair use logic that they had already weaved into the l…

> I still think "training is fair use" still has a leg to stand on, though If that's the case, we need to serious re-consider how we reward Open Source as a society (I think that would be fantastic anyway!) -- we have people producing knowledge and others profiting directly from this material, producing new content and new code that's incompatible with the original license. You make GPL code, a make an AI that learns…

If open source code authors (and other content creators) don't want their IP to be used in AI training data sets then they can simply change the license terms to prohibit that use. And if they really want to control how their IP is used then they shouldn't host it on GitHub in the first place. Of course Microsoft is going to look for ways to monetize that data.

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#209
post #51

Earlier quoted context omitted.

As well, Google is, because of these things, somewhat acting as a library. An libraries are very special entities.

Not really. Google won because Google Books was not actually a new concept; someone else had already built a book search engine the same way Google did, also got sued by the Authors Guild, and also prevailed. The only thing different about Google Books was that it'd give you two pages worth of excerpt out of the book. So it was very easy for a court to extend the fair use logic that they had already weaved into the l…

> But it doesn't save GitHub Copilot because they're not merely training a model; they're selling access to its outputs and telling people they have "full commercial rights" to its outputs (i.e. sublicensing).

But if you read the source code of 100 different projects to learn how they worked and then someone hired you to write a program that uses this knowledge, that should be legit. I'm not sure if the law currently makes a distinction between learning vs. remixing, and if Copilot would qualify as learning.

Re: GitHub is sued, and we may learn something about Creative Commons licensing

#210

Earlier quoted context omitted.

On the other side, they could argue that it's like a human learning how to code over a decade of looking at the internet, and that human doesn't need to DM every code author to ask if they can learn from their content (and the risk for the author is similar given the human might one day recall some author's code verbatim and not give attribution).

But the thing is that we explicitly allow humans to learn and develop their own skills learning from other humans, but we have our own taboos around directly copying peoples work without permission and passing it off as your own. The debate is that copilot isn’t a human, it’s a machine that outputs copied work on a statistical basis. Humans are allowed to be unoriginal, uncreative, boring, mediocre, and all sorts of…

Whenever I've used Copilot it never seems to copy whole sections of code. Can you provide examples of this?

From what I've seen it is producing fairly generic boilerplate that has been modified based on the rest of the code in my repo so that it works with the other functions and even incorporates other pieces of my code in the same style that I'm using. The boilerplate aspect makes sense because this would be the most common sequence of tokens that it observed during training. It's somewhat miraculous that it can incorporate code on the fly from my repo. I've never seen anything that looks like a direct copy paste from elsewhere though. If you have a different observation I'd love to see it.

Post reply on HN