Live data from Hacker News

I do not agree with Github's use of copyrighted code as training for Copilot

thelig.ht

261–270 of 545 posts

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#261
I'm looking for a new place, because of GitHub's new policy of not supporting password authentication.

I sometimes code from devices which are not my own and on which key management is a major impediment and accessibility issue for me.

Does anyone know how those listings were generated? I like their simplicity, and would like to do something similar.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#262

This is typical Microsoft behavior: embrace, extend, and extinguish. They embraced open source with the intention of controlling (GitHub) and exploiting it. And the interesting thing is that many people fell for this already ancient strategy.

Well, I did try to warn the Copilot fanatics [0]. They just downvoted me days ago and here we are. We have a GitHub Copilot backlash against the hype squad.

The GitHub CEO is no where to be found to answer the important questions on software licenses, copyright and the legal implications on scraping the source code of tons of projects with those licenses for Copilot.

The fact you can only use it in VSCode and with Microsoft having an exclusive deal with OpenAI screams an obvious 'embrace and extend',

As for 'Extinguish', they will need to be very creative on that.

[0] https://news.ycombinator.com/item?id=27685104

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#263

Of all the hills to die on, this seems like an odd choice. Why not work with others to iron out the legal and technical issues with this new technology?

It's a reflex with some people and that's OK.

Not everyone has the patience and ability to discuss their objections in a public forum while their rights are being violated (in their view).

Some people have a passion and a very strong belief in their ideals and I applaud them for following through with it, even if I don't necessarily share their opinion on the matter.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#264

Earlier quoted context omitted.

> Copyright has concluded that reading by robots doesn’t count. Infringement is for humans only; when computers do it, it’s fair use. This would be interesting to test with AI and pop music.

This is a stupid argument that the Twitter author made. Saving music digitally is reading by robot, so recording music that wasn't digital into a digital format is fair use.

> recording music that wasn't digital into a digital format is fair use

If you’re doing it from an analog format you bought for your own use (format shifting), it is fair use.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#265
post #141

Earlier quoted context omitted.

> Nobody is really being hurt when a new tool makes it easier to copy little bits of code from the internet. Quite the opposite. We all get a tiny bit better with good information like this. This is what the internet should be for, evolving, learning from past mistakes, information availability. If the discussion was “I clicked this button and got someone’s entire chat platform” that would be different. Words and sen…

> Words and sentences aren’t copy written, books are, so when exactly are a collection of words a book? If that were true, then 20 people could each steal a single chapter from a book, and one of the people could combine those 20 chapters into a new copyright-free book. That's clearly false.

Did I say anything about paragraphs or chapters? Didn’t I specially write there is nuance?

And for your strawman, the assembly of uncopywritable components into a copywriten work, would still be a violation.

So we agree that copywriter is somewhere between the paragraphs and chapters and the book. So why are tiny code excerpts “a problem”?

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#266
post #185
post #139

Earlier quoted context omitted.

So, to verify, your claim is that GPT-3, when trained on a corpus of human text, isn't merely managing to string together a bunch of high-probability sequences of symbol constructs--which is how every article I have ever read on how it functions describes the technology--but is instead managing to build a model of the human world and the mechanism of narration required to describe it, with which it uses to write new…

What? I don't think I made any claim of the sort. I'm claiming that it does more than mere regurgitation and has done some amount of abstraction, not that it has human-level understanding. As an example, GPT-3 learned some arithmetic and can solve basic math problems not in its training set. This is beyond pattern matching and replication, IMO. I'm not really sure why we should consider Copilot legally different from…

The argument I was responding to--made by the user crazygringo--was that GPT-3 trained on a model of the Windows source code is fine to use nigh unto indiscriminately, as supposedly Copilot is abstracting knowledge like a human engineer. I argued that it doesn't do that: that GPT-3 is a pattern recognize that not only theoretically just likes to memorize and regurgitate things, it has been shown to in practice. You then responded to my argument claiming that GPT-3 in fact... what? Are you actually defending crazygringo's argument or not? Note carefully that crazygringo explicitly even stated that copying little bits and pieces of a project is supposedly fair use, continuing the--as far as I understand, incorrect--assertion by lacker (the person who started this thread) that if you copied someone's binary tree implementation that would be fair use, as the two of them seem to believe that you have to copy essentially an entire combined work (whatever that means to them) for something to be infringing. Honestly, it now just seems like you decided to skip into the middle of a complex argument in an attempt to made some pedantic point: either you agree that GPT-3 is a human that is allowed to, as crazygringo insists, read and learn from anything and the use that knowledge in any way they see fit, or you agree with me that GPT-3 is a fancy pattern recognizer and it can and will just generate copyright infringements if used to solve certain problems. Given your new statements about Copilot being a "fancy pen" that can in fact be used incorrectly--something crazygringo seems to claim isn't possible--you frankly sound like you agree with my arguments!!

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#267
post #216

Earlier quoted context omitted.

It's not "literally" stealing, because it doesn't deprive anyone of the use the source code. Those two points were somehow extremely obvious to everyone here as long as it was music and movies we were talking about. And Github themselves have stated that only 0.1% of the Copilot output contains chunks taken verbatim from the learning set. Of those, the vast majority are likely to be boilerplate so generic it's silly…

> It's not "literally" stealing, because it doesn't deprive anyone of the use the source code. That's simply not true. You might be confusing idealism about software freedom with how both law and society define theft. Edit: In this comment I refer to the US.

It is actually true, in the UK at least the legal definition of theft includes the deprivation of the owner of the property in question.

The copyright lobby hedge the term as "copyright theft" (i.e. not actual theft) in order to shift the societal understanding. Whish appears to have worked.

This is not a value judgement on copyright infringement. Just that technically it doesn't meet the legal definition of theft.

cf. The rather amusing satire of the "you wouldn't steal a handbag" campaign in the UK, which ran "you wouldn't download a bear!"

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#269
post #34

I thought this was a pretty good thread (by an ex-Wikipedia lawyer) on Twitter about the IP meaning of Copilot. https://twitter.com/luis_in_brief/status/1410242882523459585... And this is a longer article about how IP and AI interact: https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright... I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allow…

Well, maybe the interpretation will change if the right people are pissed off.

At this point, how hard would it be to produce a structurally similar "content-aware continuation/fill" for audio producers, film makers, etc, which suggests audio snippets or film snippets, trained from copyrighted source material?

If prompted by a black screen with some white dots, the video tool could suggest a sequence of frames beginning with text streaming into the distance "A long time ago in a galaxy far far away ..." and continue from there.

Normally we don't try to train models to regurgitate their inputs, but if we actually tried, I'm sure one could be made to reproduce the White Album or Thriller or whatever else.

Re: I do not agree with Github's use of copyrighted code as training for Copilot

#270
post #34

I thought this was a pretty good thread (by an ex-Wikipedia lawyer) on Twitter about the IP meaning of Copilot. https://twitter.com/luis_in_brief/status/1410242882523459585... And this is a longer article about how IP and AI interact: https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright... I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allow…

> As a human, I am allowed to read copyrighted code and learn from it. Of course not. Reading some copyrighted code can have you entirely excluded from some jobs - you can't become a wine contributor if it can be shown you ever read Windows source code and most likely conversely. Likewise, you can't ever write GPL VST 2 audio plug-ins if you ever had access to the official Steinberg VST2 SDK. Etc etc... Did people fo…

Wasn't that the entire premise of "Halt and Catch Fire"?
Post reply on HN