Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

811–820 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#811
Although I'm aware that this tool is a boon to many, particularly those with impediments like RSI, I still have to echo what a number of other comments say: There really is a very large proportion of adult software developers in the market who are simply too young to have lived through the EEE Microsoft era. Add on to that the proportion of old-enough Microsoft-brand "dotnetter" software developers who simply don't care as long as they get to sit comfortably within C#, Visual Studio and Azure.

After that, what are you left with? A small enough proportion of developers, and Microsoft evidently thinks so, who don't know, and/or don't care, and/or don't have the time to fight their Extend-Embrace phase of take-over of Github.

One could argue that the purchase of Github was Extend, and their involvement with OpenAI, the Codex, and the potentially illegal use of OSS (subject to the legal investigations) is Embrace.

It's my own personal view that Microsoft held-back the progress of software development by probably a decade or so with their shady commingling with academia, blatant crippling of C# .NET to sell Visual Studio, and endlessly so forth. So I am, along with many, upset to see a business like this EEE their way into OSS, something which is dear and special to so many.

In the end, and I must state in my own opinion (since there is an element of speculation here), I am just pleased that there are still people out there who are not letting Microsoft continue their old ways.

Re: GitHub Copi­lot inves­ti­ga­tion

#812
Reality: *GPL licenses are proprietary licenses.

I hope Copilot and similar technologies weakens the copyright establishment.

Do Business WITHOUT Intellectual Property - Stephen Kinsella http://www.stephankinsella.com/wp-content/uploads/publicatio...

Against Intellectual Property - Stephen Kinsella https://mises.org/library/against-intellectual-property-0

Re: GitHub Copi­lot inves­ti­ga­tion

#813
post #491

It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. The only way you can get it to do that is to bait it with the function names of functions that have already been copy and pasted thousands of times onto GitHub without proper licenses. Luckily, someone will probably come out with a "renegade" vers…

> It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. Sure — in the same way that hacking into a competitor's GitHub account and copying their private source code is "genuinely useful" to you. As the person benefitting from unlawfully using their source code, of course you wouldn't care that it rep…

Straw man. That isn't what Github/Copilot is doing.

If open source communities are worried about having their source code copied... then don't open the source. Keep it closed, keep it off GitHub... I mean the genie is already out of the bottle, so doesn't really matter what they do now.

You can't prompt Copilot with things like: # Function that detects spam accurately

And get anything useful/sensitive/competitive

Is there really super sensitive algorithms out there that Copilot is exposing that are otherwise unknown?

Re: GitHub Copi­lot inves­ti­ga­tion

#814

Earlier quoted context omitted.

> You have the choice to turn that filter on or off during setup. Notice that Copilot often gives code that verbatim matches opens source software, even when that filter is on. For example: https://twitter.com/DocSparse/status/1581461734665367554?s=2... Their approach of "matches or near matches (ignoring whitespace)" is clearly inadequate, and it's honestly insulting that they think this is enough. Even if Copilot j…

>Notice that Copilot often gives code that verbatim matches opens source software, even when that filter is on. I saw a few examples, but I don't see how that extrapolates to often. It's quite possible I've missed something in the article since I kinda skimmed it. :) >and it's honestly insulting that they think this is enough. They don't. - "We plan on continuing to evolve this approach and welcome feedback and comme…

>They don't. - "We plan on continuing to evolve this approach and welcome feedback and comment."

That is corporate speak for "we plan to do nothing about this".

Re: GitHub Copi­lot inves­ti­ga­tion

#815
post #491

It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. The only way you can get it to do that is to bait it with the function names of functions that have already been copy and pasted thousands of times onto GitHub without proper licenses. Luckily, someone will probably come out with a "renegade" vers…

I don't think most people are concerned that Copilot is going to be reproducing verbatim copyrighted code, it's more that it sucks that a giant corporation is going to make a billion dollars from a tool that is entirely built off of millions of peoples' work who were never asked permission and will never be compensated.

If Copilot makes a billion dollars, it is only because it is generating at least a billion dollars worth of value to the community of developers who want to use it.

The people painting Microsoft as a big, greedy trust conveniently ignore that Copilot would actually be empowering the ecosystem of tech companies to develop services that compete with Microsoft faster and more easily.

Re: GitHub Copi­lot inves­ti­ga­tion

#816
My biggest concern regarding GitHub Copilot is that it is cloud based and opens up our previously private coding activities to continuous surveillance by third-parties.

It's only a matter of time before intelligence agencies will get their hands on the data. And if use of Copilot becomes an industry wide practice then those who wish to preserve their privacy will become uncompetitive.

I really hope we have some decent offline alternatives eventually.

Re: GitHub Copi­lot inves­ti­ga­tion

#818

Earlier quoted context omitted.

The law does not concern itself with trifles.[0] Programmers tend to think of copyright as a Boolean valued function. Either something is infringement or it isn’t. Judges think of copyright infringement as a real-valued function of many arguments corresponding to the circumstances of the parties (e.g. what actual damage was done?). A human quoting a human without attribution, without any profit made or identifiable d…

> A human quoting a human without attribution ... as opposed to AI. At the heart of the matter lurks a debate whether AI is an independent phenomenon which behaves in its own right, or a just a tool that's created and wielded by humans against a backdrop of clear incentives and motivations. The argument isn't about whether or not the law deals in a absolutes - it's a basic principle that law is tested in courts throu…

I disagree with your statement about ai being an independent behavior.

Suppose we add a button to a visual studio plugin called 'Copy me a function' and when you click it, it 100% grabs some random code from github and plops it as-is into your code base.

I don't have to argue the ethics of if the button is 'thinking for itself'

Re: GitHub Copi­lot inves­ti­ga­tion

#819
post #640
post #486

Here are a few thoughts I haven't formulated before: It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human culture (g…

I see this basic logic in almost AI ethics threads, and it starts with a big assumption: "humans learn from copyrighted source material without copyright violation". This then gets tenuously extended to "ai also learns, so it must not be in violation of copyright law". The first assumption is highly flawed though. Humans routinely do violate copyright law. Plagiarism is a huge problem in many sectors; un-cited direct…

> If an AI is ingesting and perfectly reproducing someone else's copyrighted works, it is in violation of copyright law in the same way a human would be if they reproduce someone else's copyrighted works.

"In computer programs, concerns for efficiency may limit the possible ways to achieve a particular function, making a particular expression necessary to achieving the idea. In this case, the expression is not protected by copyright."

https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...

Consider the impact on innovation if Microsoft or Oracle were allowed to claim a copyright over utilitarian aspects of their works such as the Java or Windows API!

BTW, Copilot seems to be reproducing copyrightable material when the tool reproduces comments verbatim!

Re: GitHub Copi­lot inves­ti­ga­tion

#820
So happy to learn of this and I wish them best of luck in their efforts. And I'm surprised to find so many people klinging to Copilot.

We shouldn't shed any tears for a megacorporation which shows such blatant disregard for the licensed works of people's labour.

Yes, AI is here to stay but we should be able to build AI that respects copyright. Yes, it's easier to just steal data and call it fair use. Whether or not that's stealing will be interesting to try in court.

Post reply on HN