Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

741–750 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#742
post #727

Earlier quoted context omitted.

> It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law Which laws are considered in this case? I understand that fair use is a US concept. For example how does that apply to my projects, published and licensed by a European living in a European country? I would expect the majority of GitHub contributors to not be based in the US, so what laws should…

From GitHub’s Terms of Service [0]: > Except to the extent applicable law provides otherwise, this Agreement between you and GitHub and any access to or use of the Website or the Service are governed by the federal laws of the United States of America and the laws of the State of California, without regard to conflict of law provisions. You and GitHub agree to submit to the exclusive jurisdiction and venue of the cou…

Anyone can say that, but that doesn’t make it real, especially with regards to European consumer protection.

Re: GitHub Copi­lot inves­ti­ga­tion

#743

Copilot is trained on and returns AGPL code verbatim. It’s game over. If these licenses are not enforced it defeats the entire purpose.

That's a problem of the licensers, not for Microsoft or the CoPilot users. If you released AGPL code but never intended to ever sue anyone. Why did you release it like that? If you did and if someone is able to use your code without any damage to you, without reputation loss, and via a way they have access to the innocent infringer defense after you overcome fair use, after you sue them. How is that game over?

So suppose you go out and about and a Microsoft representative punches you in the face. Now, the Microsoft representative has a billion dollar corporation backing him, willing to defend him at all cost through every institution, while you're just John Doe who went on a trip.

If you ever went on a hike but never intended to sue anyone. Why did you go out in the first place?

If you did and someone is able to punch you in the face without any lasting damage, without reputation loss, and via a way they have access to the myriad legal defenses you couldn't come up with if you tried, after you sued them.

How is that game over?

Just because someone corporation is, because of its sheer size, over the law (as far as a John Doe is concerned anyways), does that make it a right? We could probably do away with laws at that point and just accept getting punched in the face by Microsoft whenever they feel like it as the new reality.

Re: GitHub Copi­lot inves­ti­ga­tion

#744

While the moral and legal discussions here are interesting and worth exploring, I find this text hyperbolic. Its premise is that the main way that people currently interact with open-source projects is by digging into their source code, copy-pasting away a snippet of code that solves a particular problem, and then of course giving the authors the required attribution. This is far from the truth. The main usage of mos…

I haven’t used Copilot but do its samples give links on where it was from? If so, that seems to be a sufficient funnel back to the OSS repo itself without the community harming aspects mentioned in the article.

They don't. It would be a profoundly difficult problem to find the right links for each suggestion.

Re: GitHub Copi­lot inves­ti­ga­tion

#745

While the moral and legal discussions here are interesting and worth exploring, I find this text hyperbolic. Its premise is that the main way that people currently interact with open-source projects is by digging into their source code, copy-pasting away a snippet of code that solves a particular problem, and then of course giving the authors the required attribution. This is far from the truth. The main usage of mos…

I haven’t used Copilot but do its samples give links on where it was from? If so, that seems to be a sufficient funnel back to the OSS repo itself without the community harming aspects mentioned in the article.

Not at all. There isn't even a way to get the "source" if you wanted.

Re: GitHub Copi­lot inves­ti­ga­tion

#746
post #685

Don't confuse what you want with what the law says "Your work is under copyright protection the moment it is created and fixed in a tangible form that it is perceptible either directly or with the aid of a machine or device" [ https://www.copyright.gov/help/faq/faq-general.html ] A copy is made whenever that text is displayed, e.g., in GitHub's UI. Even that copy is subject to copyright. Is there an excuse/exception?…

No post body was provided.

Re: GitHub Copi­lot inves­ti­ga­tion

#747
post #535

Earlier quoted context omitted.

I don't think most people are concerned that Copilot is going to be reproducing verbatim copyrighted code, it's more that it sucks that a giant corporation is going to make a billion dollars from a tool that is entirely built off of millions of peoples' work who were never asked permission and will never be compensated.

That's hardly a new thing! For instance, Google search makes billions of dollars by indexing content that other people make.

Google helps you find someone's content. Copilot helps you rip off someone's content.

Re: GitHub Copi­lot inves­ti­ga­tion

#748

Earlier quoted context omitted.

Personally I'm not worried about the end user using copyrighted code. That is their responsibility. If you have verbatim GPL code in your commercial closed source code base that is a liability and it might be dangerous to use copilot. What I have more of a problem with is Microsoft charging for copilot which was trained on copyrighted code without any permission whatsoever which they really have no right to utilize/c…

As a human, if I learn how to program by studying copyrighted code, is it unethical for me to use that knowledge to make a living ?

You eventually, at some point, write your own code which is not reproduced. From what I see Copilot authors state that copilot is reproducing code, it is not inventing new code like human being would. It seems to be big and efficient database of somebody else's code that it reproduces.

Re: GitHub Copi­lot inves­ti­ga­tion

#749

Earlier quoted context omitted.

I really appreciated the argument in the book "The New Breed," which is that we should adapt ideas around the governance of animals to governance of ML: You can train your dog to attack random passersby, but if you do, you're a monster and ultimately responsible for the dog's actions. Likewise, you can tell Copilot to crank out code specific algorithms written by specific people, but if you do so, you're still creati…

This just sounds like blaming the researchers to me. How would i ever know if my "boring code completion" was actually copyright infringement? Your argument just disallows discussing the problem while doing absolutely nothing about it. If you train your dog to NOT attack random passersby and it still does, that dog is euthanized no matter your intentions.

If you train your dog to NOT attack random passersby and it still does, that dog is euthanized no matter your intentions.

Of course, but you will not face manslaughter charges in that case.

So, following the same logic, if you train your copilot NOT to infringe on other people's copyright and it still does, it should be destroyed no matter your intentions. But at least you won't be charged with copyright violation yourself.

That said, I don't believe Microsoft's actions to be benign. I think this copyright whitewashing scheme is fully in line with their old MO, purposefully creating a legal quagmire surrounding all open source code.

Re: GitHub Copi­lot inves­ti­ga­tion

#750
post #78
post #52

I'd be rather saddened if Copilot was shut down or neutered because of a few vocal few protesting against it. It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.

Copilot wouldn't be shut down or neutered because "a few vocal people" protested against it. It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus. You act like Microsoft is trying to do a public service and people are angry about it. The reality is that they're taking billions of hours of work and using it to build a product t…

> It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus

I'm almost entirely certain you're wrong about the desires bit. 99% of the developers who wrote that code won't mind.

Post reply on HN