Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

191–200 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#191
post #173

If you're against Copilot as developer, you're shooting yourself in the foot. Locking up code under non-permissive licenses stymies the pace of code development and increases the costs of progress dramatically. We all stand on the shoulders of others before us. Including the organisations that stand to benefit the most from aggressive licensing.

Not my problem.

I put my time and my effort to open source a program for free and I want to make sure that my code creates an incentive to create more free software, by using a copyleft license.

Re: GitHub Copi­lot inves­ti­ga­tion

#192

Earlier quoted context omitted.

Why should Copilot engineers, and the company that invests in it, not be rewarded for their incredible product and SaaS offering they spend resources on providing?

Quoted post unavailable.

That's not theft. Was the original code deleted? No. Then it's copying. And not even that, is the model replicating the training set like Google search? No. Then it's some kind of derivative work. And especially for Github it's OK because user agreements allow MS to do it.

Is "imagining" the same with "copying"? Does copy-right cover learning-right? Can learning and practicing be restricted by the authors? Can visual styles, algorithms and facts be copyrighted? I say no to all of them.

Re: GitHub Copi­lot inves­ti­ga­tion

#193
post #163

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

> "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" https://twitter.com/docsparse/status/1581461734665367554 An english description plus three characters of a function name is enough to coax CoPilot into distributing LGPL-licensed code out of context, without a proper license. That's neither "emotional" nor "cherry-picked", it's a clear-cut license violation.

No post body was provided.

Re: GitHub Copi­lot inves­ti­ga­tion

#194
post #163

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

> "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" https://twitter.com/docsparse/status/1581461734665367554 An english description plus three characters of a function name is enough to coax CoPilot into distributing LGPL-licensed code out of context, without a proper license. That's neither "emotional" nor "cherry-picked", it's a clear-cut license violation.

I mean, he clearly knew about that code in advance and used his prior knowledge to coax Copilot into spitting it out, yeah? Three characters can get you pretty far, that's 1 combination out of 125,580 (considering all english letters, upper and lower, along with most of the numbers and symbols on my keyboard), plus the description of a fairly complex algorithm.

Also, this code is really just executing a mathematical operation in what I would assume is fairly standard, so it may not even fall under copyright. IANAL.

Even if it is a copyright violation, that is one out of, IDK, millions, maybe billions already of Copilot completions?

That sounds cherry-picked. Just because some high profile, highly-retweeted person says something doesn't make it not cherry picked.

Re: GitHub Copi­lot inves­ti­ga­tion

#195
post #78

Earlier quoted context omitted.

Copilot wouldn't be shut down or neutered because "a few vocal people" protested against it. It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus. You act like Microsoft is trying to do a public service and people are angry about it. The reality is that they're taking billions of hours of work and using it to build a product t…

Why should Copilot engineers, and the company that invests in it, not be rewarded for their incredible product and SaaS offering they spend resources on providing?

They built it using some publicly available resources. These resources are available conditionally, subject to licenses (such ad GPL). Which is fine.

The problem is that their product sometimes produces verbatim copies of licensed works, without attaching licensing information. This not only goes against the licenses under which the original authors made these works available. It can also put the product's users in danger of anything from bad publicity to a copyright lawsuit.

CoPilot is a very interesting research project. It's not yet an acceptably mature product though.

Re: GitHub Copi­lot inves­ti­ga­tion

#196
post #52

I'd be rather saddened if Copilot was shut down or neutered because of a few vocal few protesting against it. It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.

The people protesting aren't a "vocal few"; We're the people who made copilot possible. We are frustrated that our work is being used to profit a massive corporation without any compensation and in a way that we at best did not intend to allow and at worst is in direct violation of the terms we set.

Re: GitHub Copi­lot inves­ti­ga­tion

#197

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

I'm expecting to see dual-licensing used as precedent here.

Github/OpenAI should have to pay a licensing fee to use GPL and similarly-licensed code in their closed-source derivative IP (CoPilot).

Re: GitHub Copi­lot inves­ti­ga­tion

#198

I wonder what will happen when a company pays some overseas developers $50 for some code, they copy it from Copilot and it copies a bug from a US developer and that company gets hacked for $10 million. Will the lawsuit fall on the overseas developer, US developer or Github?

No one? They'd probably stop doing business with the overseas developer and that's it.

Re: GitHub Copi­lot inves­ti­ga­tion

#199

One issue I see with Copilot is that they get free access to all open-source data on GitHub, but using GitHub APIs to download the data yourself isn't possible (rate limiting). This is an unfair advantage. Copilot is not only making money off of open-source, they are making money off of open-source in a way others can't. I would love to see a lawsuit which requires GitHub to provide their full Copilot dataset.

Don't forget that GitHub is a closed-source proprietary commercial service.

Their freemium product is useful to many open-source projects and communities, but you do not have any more rights to use Microsoft's GitHub than you have to Microsoft's Windows.

Re: GitHub Copi­lot inves­ti­ga­tion

#200
post #143

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

If they continue that path, the future will be that OpenAI, Microsoft, Google etc. will pay larger and larger fines at least in the EU, until they are blocked entirely.

Which may be entirely justifiable

Earlier HN thread today on a large chunk of OS code pasted almost verbatim by the CoPilot engine into a project, but stripped of any licensing references.

Within the last few days, another thread on artists who have spent decades developing a unique and valuable style are making parallel complaints about Dall-E/SD/etc., where inputting "Xyz in the style of [Artist]" produces exactly a copy of [Artist]'s unique style, barely distinguishable from the original.

These engines are fairly literally giant collage engines, able to parse language inputs and output a collage of the input works. Maybe some are small snippets so it could be fair use, but they are also evidently capable of outputs of a far larger scope, amounting to wholesale ripoff.

Opting out or not posting on Github or whatever prevents nothing, as stuff is posted everywhere by many, and with code, it's totally legit posting a fork under OS licensing.

Is there a solution analogous to a flag? How do we verify it? Will there soon be HaveIBeenUsedAsTraining adversarial systems to probe these output engines?

Not sure of the solution, but this seems to rather rapidly overstepping boundaries of creators.

Post reply on HN