Live data from Hacker News

I do not agree with Github's use of copyrighted code as training for Copilot

thelig.ht

281–290 of 545 posts

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#281
post #155

Feels like everyone is missing the point: Copilot will ultimately serve to weaken the arguments in support of software patents and copyright. That can only be a good thing for society (though perhaps not for rent seekers).

It is certainly fascinating to see people start running away from " information wants to be free " and other Free Software principles full tilt when, all of a sudden, it's their livelihoods that are on the line. Unless my recollection is off, the GPL was never the goal of the original Free Software movement; it was merely a tool to get to the end state where all code becomes available for use by anyone for any reason…

Your recollection is off, majorly. I'd recommend looking up the origins of the FSF/GPL/Copyleft. The entire movement essentially got started because Stallman gave Symbolics his (public domain) Lisp interpreter, then Symbolics improved it but refused to share the improvements.

"No restrictions" has never been the goal and to claim that they're egoistic hypocrites who are just scared for their own livelihood because of this is just an absurd strawman.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#282
post #34

I thought this was a pretty good thread (by an ex-Wikipedia lawyer) on Twitter about the IP meaning of Copilot. https://twitter.com/luis_in_brief/status/1410242882523459585... And this is a longer article about how IP and AI interact: https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright... I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allow…

Just when I thought tweetstorms couldn't get any worse, here's one where every tweet is a quote-tweet of the author. I don't even understand how I'm supposed to read this. > Copyright has concluded that reading by robots doesn’t count. Infringement is for humans only; when computers do it, it’s fair use. Surely there's a limit to this. If I use a machine to produce something that just happens to exactly match a copyr…

That quote is basically entirely nonsensical. 'copyright' hasn't decided anything (nor has any legislative body nor the courts). All that's happened is that OpenAI has put forward an argument that using large quantities of media scraped from the internet as training data is fair use. This argument for the most part does not rely on the human vs machine distinction (in fact it leans on the idea that the process is not so different from a human learning). The main place this comes up is the final test of damage to the original in terms of lost market share where it's argued that because it's a machine consuming the content there's no loss of audience to the creator (which is probably better phrased as the people training the neural net weren't going to pay for it anyway). A lot does ride on the idea that the neural net, if 'well designed', does not generally regurgitate its training data verbatim, which is in fairly hot dispute at the moment. OpenAI somewhat punts on this situation and basically says the output may infringe copyright in this case, but the copyright holder should sue whoever's generating and using the output from the net, not the person who trained and distributed the net.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#283
post #106

Earlier quoted context omitted.

That's not what GP is saying. In general, you're absolutely allowed to learn programming techniques from anywhere . You can contribute software almost anywhere even if you've read Windows source code. Re-using everything you've learned, in your own creative creation, is part of fair use. Your example is the very specific scenario where you're attempting to replicate an entire program, API, etc., to identical specific…

This is true, but there's also a murkier middle option. I used to work for a company that made a lot of money from its software patents but I was in a division that worked heavily in open-source code. We were forbidden to contribute to the high-value patented code because it was impossible to know whether we were "tainted" by knowledge of GPL code.

Same here. I worked at a NAS storage (NFS) vendor and this was a common practice. Could not look at server implementation in Linux kernel and open source NFS client team could not look at proprietary server code.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#284
Maybe now is the time to release a GPLv4 extending-restricting-relating the four freedoms to non-humans.

I expect the best lawyers from Microsoft have had a look into this and maybe there a weaknesses in GPLv3 ready to exploit for corporate AIs. What is the response from the FSF?

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#285
post #51

Earlier quoted context omitted.

> Who is this person? One of the beautiful things about HN is that you don't need to be anything, you just have to have something interesting to say.

Right, but you either need a solid argument or some authority, and this guy has neither. He's effectively a nobody and he has just jumped to the conclusion that CoPilot is illegal. If he had a good argument for that, fine. But without that he really needs to be someone whose opinion I care about.

This is somehow inverse logic. Does rape victim needs authority to voice raping in order to validate it? What is there that is not solid, CoPilot is using community code that is under GPL licence therefore Microsoft should not be able to charge for CoPilot but give it for free, or not create another revenue stream.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#286
post #46

Some unknown person is trying to get some hype on “cancel github” cry. I don't give a shit about the Copilot, but I care even less about Rian Hunter and his statements.

>Some unknown person is trying to get some hype on “cancel github” cry. >I don't give a shit about the Copilot, but I care even less about Rian Hunter and his statements. This is untrue because you had a choice of not saying anything at all and carrying on (clearly not giving a shit) or take the time to leave such a comment (giving enough of a shit to inform everyone you don't give a shit.) So far this and Lloyd's is…

Is it not obvious that you can care about a post on HN without caring about the page it links to?

Back away from this specific situation for a second: If you would ignore something entirely if it wasn't being shoved in your face, complaining about it being shoved in your face and saying it's stupid wouldn't mean you suddenly "care" about the underlying item.

(And no, I'm not saying that an HN post is shoved in your face. It's a more extreme example to make the point more clear.)

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#287
Seems to me like they need to back out of this fast and at very least limit it such that it is only trained and then used on "license compatible" projects. eg: train it in isolation on MIT licensed projects and then have the user explicitly confirm what license the code they are working on is to enable it. Possibly they even need to auto-enable a mechanism to detect when code has been reused verbatim and enable some kind of attribution (or respect for other constraints) where that is required by the license.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#288
post #34

I thought this was a pretty good thread (by an ex-Wikipedia lawyer) on Twitter about the IP meaning of Copilot. https://twitter.com/luis_in_brief/status/1410242882523459585... And this is a longer article about how IP and AI interact: https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright... I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allow…

Autonomous programming will be explored. Potentially, Copilot is a proof of concept, an early step in that direction. If it is, the corrections made by Copilot users will be applied to the development of the future of unattended programming. Whether it is or not, it's close enough that any legal outcomes experienced by Copilot users will contribute to the definition of liability boundaries relevant to the future of autonomous programming. Copilot users are numerous enough that the incidence of risk is low of ending up under the foot of a copyright owner with the means and will to crush a user, but no one should take such a risk to use a novelty like Copilot in production code.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#290
post #155

Feels like everyone is missing the point: Copilot will ultimately serve to weaken the arguments in support of software patents and copyright. That can only be a good thing for society (though perhaps not for rent seekers).

Not really. Back when free software was strong, it would have been a good thing for society since Microsoft was selling software in boxes on actual store shelves. Now 'the edge' is already mostly open source. All the lock-in and value has moved into either infrastructure or in software you don't even get to touch since it runs in the Cloud and you just provide IO to it.

I think in this new era of endless security breaches at cloud firms and M1-style processing innovation we'll see a slow but steady migration away from the cloud.
Post reply on HN