Live data from Hacker News

I do not agree with Github's use of copyrighted code as training for Copilot

thelig.ht

421–430 of 545 posts

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#421
post #127

I've given thousands of hours to open source projects, I really think open source is a pillar of modern society. So you would think I am all for something like copilot, but no. At first I thought this was a great feature, because easier access to code, but after some reflection, I am also very skeptical. I am able to make my code open source, because I can make a living out of it, and I have a lot of open source code…

> I am surprised you are able to DMCA a twitch stream because someone whistle india jones theme but in this case it is considered fair use Why is it surprising? Indiana Jones is private IP. The code was published with an OSS license explicitly authorizing its use.

From the Copilot FAQ, notice there is no specific mention of the term "Open Source":

  > It has been trained on a selection of English language and source code from publicly available sources, including code in public repositories on GitHub.
Meaning GitHub might have also been sourcing from public source / source-available projects that were not OSS licensed at all.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#422
post #290

Earlier quoted context omitted.

I think in this new era of endless security breaches at cloud firms and M1-style processing innovation we'll see a slow but steady migration away from the cloud.

I used to think like this but the total cost of ownership of on-prem is substantially higher, and it has its own security implications too.

I laugh at the thought that every company trying to host 20 vendor apps is more secure than having the company that wrote the apps host them. Self hosted apps get updated on a much much slower schedule and don't have access to the brains who built the software.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#423
post #34

I thought this was a pretty good thread (by an ex-Wikipedia lawyer) on Twitter about the IP meaning of Copilot. https://twitter.com/luis_in_brief/status/1410242882523459585... And this is a longer article about how IP and AI interact: https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright... I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allow…

> I am capable of summarizing the thoughts of lawyers

Then, I'm sorry, but you seem to have done a pretty bad job of summarizing that paper ?

My own take on summarizing it's conclusion would be :

"If the program can pass the Turing Test, then it should be legally liable, just like a human would."

(Yes, emphasis on the should here, but the way that you're presenting that quote might make readers think that that paper's conclusion is the opposite one !)

----

> Nobody is really being hurt when a new tool makes it easier to copy little bits of code from the internet.

Some of the examples from that paper are A&M Records, Inc. v. Napster (no comment) and White Smith Music Publ’g Co. v. Apollo Co. (piano rolls), and in that latter case we pretty much (?) got the whole Copyright Act of 1909 the very next year, where these “mechanical reproductions” were subjected to a statutory compulsory license.

So at the very least there should be concern about how Copilot might be eventually considered by the courts to facilitate copyright infringement and at the very least have to provide the source of its "insights" ?

(Attribution being the bare minimum that most of the software licenses require.)

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#424

Earlier quoted context omitted.

"reading by robots doesn’t count." It should be obvious that if the robot is simply scraping web sites and reproducing their text verbatim (without permission and without giving credit) that would be an infringement. There are a lot of shades of gray between that and the other extreme, which is where it is scraping millions of sites, learning from them, and producing something that isn't all that similar to any of th…

Copyright requires a certain amount of creativity involved in its creation. I strongly suspect most code snippets of a few lines just don't qualify.

Yeah well I haven't looked at what this thing actually produces. Regardless the statement that "robots don't count" is ridiculous.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#425
post #129

I think it's clear that Copilot pushes boundaries...technological and legal. It makes people uncomfortable and challenges a lot of assumptions that we have about the current world. But this is exactly what I expect from the next revolutionary change in computing.

it doesn't at all challenge any assumption we have about the current world, unless you've been living with the assumption that you can't automate the process of producing code snippets. Copilot is no different than stackoverflow with the exception that your selection mechanism is done by an artifical neural net rather than a bunch of people with up and downvotes in front of their screen. The reason people are uncomfo…

It sounds like you have it all figured out, so I just have one question: if it's not innovative (it's "no different than stackoverflow with X"), then what is your explanation for why the Copilot announcement here received over 2.8k votes? It's so obviously just stackoverflow with a better selection mechanism, so why aren't people treating it as such?

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#426
post #34

I thought this was a pretty good thread (by an ex-Wikipedia lawyer) on Twitter about the IP meaning of Copilot. https://twitter.com/luis_in_brief/status/1410242882523459585... And this is a longer article about how IP and AI interact: https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright... I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allow…

> Copyright has concluded that reading by robots doesn’t count. Infringement is for humans only; when computers do it, it’s fair use.

This is a ridiculous conclusion. The ultimate destination of the robot actions' product is its user, i.e. a human. It is a clear corollary of the transitive law. Therefore, all human-focused legal concepts, including infringement, are applicable in such cases.

The absurdness of the conclusion cited above can be easily illustrated, as follows. Suppose that a person owns or rents an advanced robot (say, like Boston Dynamics' Spot, but better). He/she then programs it to break into someone's house and steal something valuable. All goes by the plan, the robot delivers the stolen goods to the rendezvous point and, if rented, gets returned. Now, according to the conclusion's logic, since a robot has done the actual "work", "it's fair use". Nonsense, right?

Just to clarify: I like the concept of GitHub Copilot (even though I have not yet tried this particular product). It offers various benefits, from pedagogical to adopting software engineering best practices to improving engineering productivity. However, I think that IP and legal aspects of this approach and specific product should be carefully studied and resolved in a consistent manner (e.g., prevent the model or system to output exact source code snippets).

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#427
post #295

I'm glad that Copilot is bringing the grey areas of copyright into discussion. If I write a book and it is copyright, what's the smallest unit which is covered by that copyright? Each word is obviously not. Some sentences will be fairly generic and I will not be the first person to write them. But some sentences will be characteristic of the work or my own style. Clearly how we apply copyright to subdivisions of an o…

This. You realize it doesn’t make any sense. All ideas are shared creations, by definition. If you’ve created something that has meaning for other people, the meaning comes from the ideas you are incorporating into your own tree. There is no defending copyright. It is indefensible from first principles. It makes no logical sense. Though it sure has proven to be a profitable con.

If there is no copyright, what incentive is there to ever create anything digital? Adobe would never invest in Photoshop if any random person was legally able to sell copies for $1 each. A production company will never publish a book or create a TV show if anyone can just undercut them by taking what they have produced and resell it. Doesn't seem like a con to me, it sounds pretty critical for any kind of functional digital market.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#428
post #89
post #34

I thought this was a pretty good thread (by an ex-Wikipedia lawyer) on Twitter about the IP meaning of Copilot. https://twitter.com/luis_in_brief/status/1410242882523459585... And this is a longer article about how IP and AI interact: https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright... I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allow…

there's a nice example here of it reproducing carmack's famous inverse square root function from Quake 3 (sans GPL, of course) https://twitter.com/mitsuhiko/status/1410886329924194309 this is clearly copyright infringement, and if it isn't: it should be

That function appears in hundreds, if not thousands of GitHub repos. It's plausible that it's the most famous block of code ever. Are all those repos guilty of copyright infringement?

https://github.com/search?q=float+Q_rsqrt%28+float+number+%2...

The only ways this argument could be less alarming to me is if people were bothered that it was writing the same Hello World as somebody else, or that it was naming variables "foo" and "bar".

Let's wait until we have a bulletproof, egregious, and inexcusable case of it lifting code until we panic.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#429

For anyone who thinks this is even remotely okay, answer me this. Imagine that Copilot is actually just a gigantic PR campaign, and when you send a completion request, they do this: They send the request off to a sweatshop in Bangladesh where a bunch of mechanical turk workers scroll through licensed codebases, find an appropriate snippet of code, agree on the best one, strip all of the licenses and attributions out,…

In your hypothetical 100% of completions are copied verbatim from existing code. But if GitHub is to believed, only 0.1% are.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#430
post #425

Earlier quoted context omitted.

it doesn't at all challenge any assumption we have about the current world, unless you've been living with the assumption that you can't automate the process of producing code snippets. Copilot is no different than stackoverflow with the exception that your selection mechanism is done by an artifical neural net rather than a bunch of people with up and downvotes in front of their screen. The reason people are uncomfo…

It sounds like you have it all figured out, so I just have one question: if it's not innovative (it's "no different than stackoverflow with X"), then what is your explanation for why the Copilot announcement here received over 2.8k votes? It's so obviously just stackoverflow with a better selection mechanism, so why aren't people treating it as such?

as far as I have seen it that is exactly how people are using it, which is why the last thread about it was (rightfully) full of people pointing out what a bad idea it is of copying half-cooked code snippets into your own codebase.

Why does it get a lot of upvotes? Because it's an AI product that makes things appear on your screen and that's the threshold for hype in our current age. Of course the number of upvotes on HN regardless doesn't speak to the innovation of anything. As best as I can tell all pre-covid posts on HN concerning mrna vaccines have a total of one upvote. Useless tech right?

Post reply on HN