Live data from Hacker News

All public GitHub code was used in training Copilot

twitter.com

681–690 of 734 posts

Re: All public GitHub code was used in training Copilot

#681

Earlier quoted context omitted.

I have about 300,000 photos that haven't been scanned by AI (unless someone at Backblaze did it without permission). I'm sure there are lots of other photographers out there who miss Picassa, which Google killed off to push everyone's data to their service. (It did really well in matching faces, even across age, but the last version has a bug when there are multiple faces in a picture, sometimes it swaps the labels)…

You don't encrypt your data before uploading to backblaze?

Oh heck no, I never encrypt data.

I run windows. It can't ever be secure, anyone who wanted to hack me could.

Scrambling the data really makes things worse as any accident requiring recovery of my data is also probably going to lose the encryption key.

The only time I ever lost any significant chunk of data (a persons lifetime set of photos!) was because Windows encrypted data at rest, and thus it couldn't be recovered after a disk crash.

Unless there is some corporate or legal requirement to do so, I'll never encrypt a whole disk, or backup.

Re: All public GitHub code was used in training Copilot

#682

Earlier quoted context omitted.

If you're thinking about the liability waiver found in many licenses and contracts and EULA and other, they are often void, depends on the jurisdiction. The official answer from Github that they take all input on purpose doesn't play in their favor.

I'm speaking specifically about the indemnity agreement that you get as part of a purchased license. It is the opposite of the liability waiver -- it is saying that the software publisher will take on responsibility in certain cases and with certain limits. For example, if I purchase certain corporate Linux licenses, i'm protected against being sued if something in the distribution ends up having misappropriated code…

Having an indemnity clause doesn't mean that a company will automatically defend you or cover your expenses. You may have to sue them to enforce the contract and agree on the costs.

The answer to what will happen when companies are bitten, is that there will be series of lawsuits involving various parties (including GitHub), dragging on for a while and costing a fortune. The court will decide everything in the end (who's responsible for what, who cover the fees, who own the IP, etc...).

The SCO case was rather frivolous, I don't think there is much to take from it, except that if a US company is determined to sue and they have a billion dollar to go on forever, there's nothing stopping them to and it's a lot of troubles.

Which is relevant I suppose. It's only a matter of time until there's a major case putting GitHub Copilot in the spotlight and an aggressive company with deep pocket (think Oracle and the likes). We will certainly be reading about it everywhere the day it starts.

Re: All public GitHub code was used in training Copilot

#683
post #551

To me, the particular use case and whether it is fair use or not, is of minor interest. A far more pressing matter is at hand: AI centralization and monopolization. Take Google as an example, running Google Photos for free for several years. And now that this has sucked in a trillion photos, the AI job is done, and they likely have the best image recognition AI in existence. Which is of course still peanuts compared…

Interesting, given that HN thinks that it is yandex who has SOA image search, not google https://news.ycombinator.com/item?id=23976172 which kinda counters your logic.

It is yandex who now collects massive amounts of data to improve their image search now, while google apparently doesn't.

Yandex is a giant, for sure, but google is, like, 10 times bigger and still doesn't provide the best service.

Re: All public GitHub code was used in training Copilot

#685
post #185

If the training set contains verbatim (A)GPL code does this mean that Copilot also should be distributed by Microsoft under GPL? Because without it Copilot (as it is distributed by Microsoft) couldn't be built, wouldn't it make it a derivative work of GPL'd code (and obviously every other license)? I see a lot of people comparing human learning to machine learning in the comments, but there is a huge difference - we…

No, see Authors Guild v. Google. Even without a license or permission, fair use permits the mass scanning of books, the storage of the content of those books, and rendering verbatim snippets of those books. The Google Books site is not a derivative work of the millions of authors they copied from, and if they did copy any coincidentally GPL, AGPL, or creative commons copyleft work, the fair use exception applies befo…

> Fair use is an exception to copyright itself.

And copyright itself is an exception to the normal state of things : the public domain, copyright being only a temporary monopoly.

Re: All public GitHub code was used in training Copilot

#686

Earlier quoted context omitted.

> By comparison, Copilot is even more obviously fair use. Not sure I see it that way. If I take your hard work that you clearly marked with a GPL license and then make money from it, not quite directly, but very closely, how is that fair use? Or legal? Copying and storing a book isn't recreating another book from it. Copilot is creating new stuff from the contents of the "books" in this case. Edit: I misunderstood fa…

> If I take your hard work that you clearly marked with a GPL license and then make money from it, not quite directly, but very closely, how is that fair use? Or legal? If I'm Google, and I scan your code and return a link to it when people ask to find code like that (but show an ad next to that link for someone else's code that might solve their problem too), that's fair use and legal. My search engine has probably…

It's fine because a search engine is a generic tool the main purpose of which is not to replicate the code verbatim to be used as code.

Re: All public GitHub code was used in training Copilot

#687

Earlier quoted context omitted.

> Copilot should be inspiring people to figure out how to do better than it, not making hackers get up in arms trying to slap it down. One of the (many) problems is that GitHub/Microsoft already benefit from runaway network effects so it’s difficult to “do better”. Where will you get all of that training code if not off GitHub? The real answer to this is to yank your projects from GitHub now while you search for alte…

Even if you do that, what's to stop them from using open source software from all over the web and not just what's on GitHub? The only way to stop them then is to go closed source.

They make you give up some of your monopoly rights when you put stuff on Github (some parts of those ToS might or might not be legal).

You would have a much stronger case if they had taken your code from elsewhere.

Re: All public GitHub code was used in training Copilot

#688

Earlier quoted context omitted.

Try doing any type of deal (fundraising, M&A) where you can't point to the provenance of your application's code. This isn't good for programmers, programmers WANT clean and knowable copyrights. This is good for lawyers, who'll now have another way to extract thousands of $$ from companies to launder their code.

If you do get sued, the Copilot page is written in a way that would make Github legally responsible for it, not you. "Just like with a compiler, the output of your use of GitHub Copilot belongs to you."

Yeah, right... This isn't going to fly in court any more than if the Pirate Bay page was written in a way that says that it's solely responsible for what you do with the magnet links that they share.

Re: All public GitHub code was used in training Copilot

#689
post #384

Earlier quoted context omitted.

> ...Should you be able to train with material you don't own? If relating this to how humans learn, books and other sources are used to inform understanding and human knowledge. One can purchase or borrow a book without actually owning the copyright to it. Indeed, a given passage may be later quoted verbatim, provided it is accompanied with a reference to its source. Otherwise, a verbatim use without attribution in a…

"If relating this to how humans learn" seems like a big IF though right? Are we going to treat computer neural nets as human from a legal standpoint? At some point Neural Nets like GameGAM might be good enough to duplicate (and optimize) a commercial game. Can you then release your version of the game? Do you just need to make a few tweaks? Are we going to get a double standard because commercial interests are oppose…

> Are we going to treat computer neural nets as human from a legal standpoint?

Maybe we will some day, but for now this isn't the case, where the law is concerned :

https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright...

Re: All public GitHub code was used in training Copilot

#690
post #525

Earlier quoted context omitted.

At that point, can we all just agree IP is the stupidest concept to ever be layered on top of math (which programming is) and move on with non-copyrightable code?

Only if you agree that copyleft licenses are also stupid; without copyright, there's no way to prevent companies from making closed-source forks of code you wrote and intended to stay open.

The whole point of copyleft was as a stepping stone to get to RMS's four freedoms (https://www.gnu.org/philosophy/free-sw.en.html) which effectively eliminates copyright for software.
Post reply on HN