Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

211–220 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#211
post #163

Earlier quoted context omitted.

> "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" https://twitter.com/docsparse/status/1581461734665367554 An english description plus three characters of a function name is enough to coax CoPilot into distributing LGPL-licensed code out of context, without a proper license. That's neither "emotional" nor "cherry-picked", it's a clear-cut license violation.

I mean, he clearly knew about that code in advance and used his prior knowledge to coax Copilot into spitting it out, yeah? Three characters can get you pretty far, that's 1 combination out of 125,580 (considering all english letters, upper and lower, along with most of the numbers and symbols on my keyboard), plus the description of a fairly complex algorithm. Also, this code is really just executing a mathematical…

Of course it is cherry picked. The idea is that it allows you to INTENTIONALLY void any copyright you want.

So let's say I obtain an illegal copy of microsoft windows' source code. Under this precedent, what stops me from just (overfitting) training a neural network to produce the source code verbatim, sans any license notice?

But it doesn't end there. What stops me from making a neural network that exactly reproduces the bytes of Illegally_Ripped_Disney_Movie.mp4 that I obtain from the pirate bay? Copyright need not apply.

At what point can the neural network I've described (which is intentionally designed to violate copyright) distinguishable from a neural network like Copilot and others which violate copyright extrinsically?

Re: GitHub Copi­lot inves­ti­ga­tion

#212
post #169

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

> What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? A better future to me. I don't want pictures of my face training ML models, nor do I want my art, or my code. I don't want my face to be more recognizable to AI, and I don't want my work to contribute to the consolidation of power to a few big firms. And for what, wh…

IDK how generative models can really be used for surveillance?

Certainly facial recognition models etc, can be, but those would seem to be appropriately covered by the Google Books ruling dealing with discriminitive models: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... They've also been a thing for quite a while, so I think that cat escaped the bag a long time ago.

Remember clip art in Microsoft Word back in the day? Now there is an infinite supply of that. Stock images? Infinite supply. Solo filmmakers are going to have a much easier time creating their own films that can rival the best movie studios in the world. Any text anywhere will be read to you in any voice or voices you like, with tone and setting appropriate sound-effects. Smaller things will just be better too, noise cancellation on microphones? Easy and free. Image editing? Trivial to remove, relight, reposition, etc, etc, etc.

So many other things too. It's going to be magnificent. If you're not into then I guess to each their own, but I do think we are looking at something that can be a net good for everyone in the world, so long as it's available and cheap for everyone in the world.

Re: GitHub Copi­lot inves­ti­ga­tion

#213
post #42

It seems like there's no good license that places absolutely no restrictions or requirements on people using your code (such as attribution and respecting patent rights) worldwide. I want my code to be used the way people treated text in the old days. There's texts that have been re-written, added to and edited by thousands of people over the centuries and yet they don't come with thousands of pages of attribution no…

CC0 or unlicense, probably. Although I concede that it's hard to do in all jurisdictions globally.

Re: GitHub Copi­lot inves­ti­ga­tion

#214
post #163

Earlier quoted context omitted.

> "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" https://twitter.com/docsparse/status/1581461734665367554 An english description plus three characters of a function name is enough to coax CoPilot into distributing LGPL-licensed code out of context, without a proper license. That's neither "emotional" nor "cherry-picked", it's a clear-cut license violation.

I mean, he clearly knew about that code in advance and used his prior knowledge to coax Copilot into spitting it out, yeah? Three characters can get you pretty far, that's 1 combination out of 125,580 (considering all english letters, upper and lower, along with most of the numbers and symbols on my keyboard), plus the description of a fairly complex algorithm. Also, this code is really just executing a mathematical…

> Even if it is a copyright violation, that is one out of, IDK, millions, maybe billions already of Copilot completions?

If you, only once, steal lines of code that you don't have license to do so and use them to make money, that's the same exact thing. "Trusting the algo" and saying "whoops I'm sorry" doesn't make a strong legal defense.

In a company of 1000 programmers, what are the odds that copilot increases the risk of using improperly licensed code because "well because microsoft gave it to us it has to be legit!"

And sure, stackoverflow copying is a thing, but they clearly tell you the license by which you can use said code: https://creativecommons.org/licenses/by-sa/4.0/

If copilot gives you CC-by-sa code, will it tell you so you can properly credit?

Re: GitHub Copi­lot inves­ti­ga­tion

#215
post #103

One issue I see with Copilot is that they get free access to all open-source data on GitHub, but using GitHub APIs to download the data yourself isn't possible (rate limiting). This is an unfair advantage. Copilot is not only making money off of open-source, they are making money off of open-source in a way others can't. I would love to see a lawsuit which requires GitHub to provide their full Copilot dataset.

Github does plenty of stuff that you can't do. e.g. the contributions graph as just one example that comes to mind. That's not unfair, and not the basis for a lawsuit, it's just business. > Copilot is not only making money off of open-source, they are making money off of open-source in a way others can't. Of course! That's why MS paid squillions to buy Github.

Its anticompetitive and illegal (if the government ever decides to start prosecuting people under those laws again).

Re: GitHub Copi­lot inves­ti­ga­tion

#216

Maybe I'm in the minority, but I think the prospect of someone autocompleting and getting a snippet that came from me, they found it useful, and are going to incorporate it is great. It means my thoughts and logic are shaping culture in a mimetic feedback loop.

That's not the problem. If you want to license your work in a way that allows that, you are free to do so. The issue is that Microsoft did that with code that was published under licenses that either did not grant that right, or which explicitly forbade them from doing so.

Re: GitHub Copi­lot inves­ti­ga­tion

#218

Earlier quoted context omitted.

Any law where the penalty is a fine only exists for the poor. Any regulation where the penalty is in the millions only exists for small businesses.

I guess that's true. If you consider laws to be strictly transactional then you can totally do the crime if you're willing to do the time. I'm just not convinced by the idea that any penalty less than death isn't a penalty.

Penalties that scale off the offender's income/revenues work a lot better. They're common in some countries.

Re: GitHub Copi­lot inves­ti­ga­tion

#219
post #162
post #143

Earlier quoted context omitted.

If they continue that path, the future will be that OpenAI, Microsoft, Google etc. will pay larger and larger fines at least in the EU, until they are blocked entirely.

And the EU will continue to fall farther and farther behind in software development.

Yeah, regulation is definitely the only reason Europe is behind. /s

Re: GitHub Copi­lot inves­ti­ga­tion

#220
post #163

Earlier quoted context omitted.

> "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" https://twitter.com/docsparse/status/1581461734665367554 An english description plus three characters of a function name is enough to coax CoPilot into distributing LGPL-licensed code out of context, without a proper license. That's neither "emotional" nor "cherry-picked", it's a clear-cut license violation.

I mean, he clearly knew about that code in advance and used his prior knowledge to coax Copilot into spitting it out, yeah? Three characters can get you pretty far, that's 1 combination out of 125,580 (considering all english letters, upper and lower, along with most of the numbers and symbols on my keyboard), plus the description of a fairly complex algorithm. Also, this code is really just executing a mathematical…

Microsoft Copilot is easily the greatest theft of intellectual property in the history of man.

You want to use my code, without ever knowing I wrote it? You want to use my hard work, regurgitated anonymously, stripped of all credit, stripped of all attribution, stripped of all identity and ancestry and citation? FUCK YOU

There's no need to defend something so obviously harmful, so why do you do it?

The law should be amended to make this kind of theft illegal.

It's not ambiguous.

Post reply on HN