Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

571–580 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#571
post #529

A sizable, possibly plurality cohort of fully adult tech people is young enough to not know about United States v. Microsoft Corp. This would explain a lot of comments I see on this topic. If you don't know Microsoft's history, a lot of what more informed people are worried about seems overblown. Copilot was Microsoft's first test of people's trust after the GitHub acquisition. It's going very, very, very poorly. The…

> Too many people are focused on what's legal. It's fine to think of, but law is the last stop before the breakdown of society. Microsoft skipped society and went straight to sparking an inevitable test of and possible reshaping of copyright law.

Maybe it's illuminating of a trait of human nature. On the stable diffusion webui repo many people have stated that they would continue to use the code even if it were stolen or unlicensed. These people aren't a part of a corporation; they are average netizens handed a technology essentially indistinguishable from magic with nothing in place to prevent its use.

If the tech is simply so impeccable as to be irresistible then a higher order framework needs to be in place to teach people not to bite because they will be bitten back.

Re: GitHub Copi­lot inves­ti­ga­tion

#572
post #486

Here are a few thoughts I haven't formulated before: It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human culture (g…

What stops humans from reproducing thinly disguised copies of their influences is, essentially, their ethical judgement.

Which amounts to saying, humans are trained with a model that they can use to recognize when something they are thinking of producing is 'too similar' to something they have seen before.

And, of course, some humans choose not to apply that filter and go ahead and plagiarize anyway; some humans try to apply that model but get back a false negative, thinking they're producing something original when they aren't. And we have ways of dealing with humans who do that.

In the case where an AI is coming up with the work, perhaps the mistake is in relying on humans to try and apply their own trained judgement to figuring out if the result is unoriginal. We need an AI that scores work for how likely it is to be infringing on a prior copyright.

Then you use that AI to train the creator AI, and teach it 'originality'.

Re: GitHub Copi­lot inves­ti­ga­tion

#573

Earlier quoted context omitted.

Personally I'm not worried about the end user using copyrighted code. That is their responsibility. If you have verbatim GPL code in your commercial closed source code base that is a liability and it might be dangerous to use copilot. What I have more of a problem with is Microsoft charging for copilot which was trained on copyrighted code without any permission whatsoever which they really have no right to utilize/c…

As a human, if I learn how to program by studying copyrighted code, is it unethical for me to use that knowledge to make a living ?

The difference is scale. You will get old and die before you ever even reach 10 percent of the corpus.

Re: GitHub Copi­lot inves­ti­ga­tion

#574
post #539
post #530

Earlier quoted context omitted.

I use Copilot all the time and I’ve never once used it to generate a whole prepackaged function that’s more than maybe three lines. So no, I don’t benefit from its reproducing other people’s code at all. Tell me you don’t use Copilot without telling me about it.

> Tell me you don’t use Copilot without telling me about it. You don't accept arguments against the use of Copilot from people unless they... use it? That's a nifty way to ignore any and all criticism of Copilot, or indeed any discussion about any ethical issue ever.

>> Tell me you don’t use Copilot without telling me about it.

> You don't accept arguments against the use of Copilot from people unless they... use it?

> That's a nifty way to ignore any and all criticism of Copilot, or indeed any discussion about any ethical issue ever.

"I only listen to people who agree with me, but to make that sound legitimate, I have a somewhat indirect way of saying so."

Re: GitHub Copi­lot inves­ti­ga­tion

#575
post #138

Good bye and good riddance. Even just the idea that GitHub should be allowed to train their proprietary AI on other people's work is insane. Much less distribute that AI in a paid package which lets you spit out other people's code verbatim. Anyone who supports open-source and the (ab)use of copyright law to create free works should be vehemently opposed to Copilot.

The fact that it's the first major development to be started at GitHub after their acquisition by Microsoft is such a hit too. Way to spend their social capital. I can't imagine the money they made from Copilot subscriptions was worth it since companies have certainly stayed away from this...

Re: GitHub Copi­lot inves­ti­ga­tion

#576
post #138

Good bye and good riddance. Even just the idea that GitHub should be allowed to train their proprietary AI on other people's work is insane. Much less distribute that AI in a paid package which lets you spit out other people's code verbatim. Anyone who supports open-source and the (ab)use of copyright law to create free works should be vehemently opposed to Copilot.

> Even just the idea that GitHub should be allowed to train their proprietary AI on other people's work is insane. You explicitly agree to this when you upload code to GitHub. FOSS folks shouldn’t have sold their soul to the proprietary devil but they did and now they have to deal with it.

You meant to say "implicitly"? Even then, no, the terms are much more specific.

Re: GitHub Copi­lot inves­ti­ga­tion

#577
post #529

A sizable, possibly plurality cohort of fully adult tech people is young enough to not know about United States v. Microsoft Corp. This would explain a lot of comments I see on this topic. If you don't know Microsoft's history, a lot of what more informed people are worried about seems overblown. Copilot was Microsoft's first test of people's trust after the GitHub acquisition. It's going very, very, very poorly. The…

Microsoft has owned Github for how many years... and this is the _first_ test?

Re: GitHub Copi­lot inves­ti­ga­tion

#578
post #491

It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. The only way you can get it to do that is to bait it with the function names of functions that have already been copy and pasted thousands of times onto GitHub without proper licenses. Luckily, someone will probably come out with a "renegade" vers…

Fine, but what happens if Copilot is so successful that it ends up actively harming the very projects that it requires for training material? If you value Copilot then you should be concerned about that, even if you don't care at all about any ethical considerations.

Seems very similar to a "tragedy of the commons" type of situation.

Re: GitHub Copi­lot inves­ti­ga­tion

#580
post #52

I'd be rather saddened if Copilot was shut down or neutered because of a few vocal few protesting against it. It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.

If the tech becomes open (and all indicators point to it being open in the near future) then it will become impossible to shut down. This has already happened with Stable Diffusion and the related model leaks.

People's expectations have already been set by this technology, and they are only going to want more. Also, AI researchers are still publishing their work out in the open for anyone to reproduce.

If there was a Copilot model out in the wild like with Stable Diffusion then this ceases to be a valid question, regardless of the model's legality. All it takes is a single leak or decision by another entity to release their own code generation model.

Post reply on HN