Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

931–940 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#931
post #491

It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. The only way you can get it to do that is to bait it with the function names of functions that have already been copy and pasted thousands of times onto GitHub without proper licenses. Luckily, someone will probably come out with a "renegade" vers…

> It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It enables the large scale theft of code. It completely ignores licenses. There are plenty of open source licenses that allow code use with proper attribution yet Copilot doesn't (and probably can't) figure a way to comply with all of them. Copilot, as the article suggests, is a marketing stunt. To me it's more than that…

> you can't just take 20 lines of completely stolen code, modify a few things, and call it your own.

Now I hope that Copilot sticks around for this exact reason: cause endless inane lawsuits claiming that actual original code was stolen and laundered through Copilot or reverse engineers going cowboy. Make it enough of a problem clogging the courts that they start dropping copyright cases.

Re: GitHub Copi­lot inves­ti­ga­tion

#932

Earlier quoted context omitted.

A few snippets of code is not a product. If there was an open-source money-making product and someone builds a competing product using considerable help from CoPilot then that is a stronger case for damages then if someone just used some snippets of code in their own product. But at that point, it would be just like someone cloning the Github code without following the license and in that case, it should become obvio…

Music samples are a natural parallel. You cannot sample music without permission no matter how short the sample may be. Similarly you cannot steal a snippet of someone else's code without permission or the correct licensing.

> You cannot sample music without permission no matter how short the sample may be.

Which is a blatantly mistaken court ruling and one which I will not enforce if I am on a jury in such a trial.

Re: GitHub Copi­lot inves­ti­ga­tion

#933

Earlier quoted context omitted.

> It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It enables the large scale theft of code. It completely ignores licenses. There are plenty of open source licenses that allow code use with proper attribution yet Copilot doesn't (and probably can't) figure a way to comply with all of them. Copilot, as the article suggests, is a marketing stunt. To me it's more than that…

> you can't just take 20 lines of completely stolen code, modify a few things, and call it your own. Now I hope that Copilot sticks around for this exact reason: cause endless inane lawsuits claiming that actual original code was stolen and laundered through Copilot or reverse engineers going cowboy. Make it enough of a problem clogging the courts that they start dropping copyright cases.

Copyright has a purpose. I don't mean a purpose in the Disney Hegemony sense but a real, actual purpose.

Dropping copyright cases is great if you hate proprietary code. That's fine. Open source is also powered by copyright. If we start dumping copyright cases we don't get the "well bob we may as well open source it!" We get large companies like Google just completely ignoring copyright and using open source code without attribution. What you've described (causing enough a problem in the courts) is the exact purpose of GPL. If we start dropping copyright cases altogether the open source movement may as well be dead in the water.

Re: GitHub Copi­lot inves­ti­ga­tion

#934
In most of the controversies posted on HN they usually end up with a feeling that nothing would change, because only us, the tech community knows about details of an issue and we are too few to have an impact.

But this is solely affecting a product where we are the target audience, where if we oppose, thing should change. Now I wonder if it will show that we are actually caring that much to action or we are just as regular consumers as non-tech people in all other cases.

Re: GitHub Copi­lot inves­ti­ga­tion

#936
post #642

Earlier quoted context omitted.

> which is what copilot has been show to sometimes do In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd bet rare things that are representative of some conceptual axis). The main examples presented so far, Davis and Carmack, have the property of having been copied a lot. The generative model…

> not just for code This is quite important, actually, and I don't think enough people realize this. I am a photographer sometimes and it would be really cool if I could share my photos online under a copyright license that forbids their use in training AI.

The problem seems to be that someone could just copy your photo and repost it without that license and we're back to the same spot.

Re: GitHub Copi­lot inves­ti­ga­tion

#938
post #810

While the moral and legal discussions here are interesting and worth exploring, I find this text hyperbolic. Its premise is that the main way that people currently interact with open-source projects is by digging into their source code, copy-pasting away a snippet of code that solves a particular problem, and then of course giving the authors the required attribution. This is far from the truth. The main usage of mos…

The whole point of open-source is about re-using and modifying the source though. Sure it allows using the product, but that's hardly the defining factor of open-source.

The point is access to the source. If that access is mediated through a system that doesn't tell you where the code comes from or how it's licensed, it's failing at open source.

It's not like they don't get it. Even Microsoft will provide the source of its products under certain circumstances for this exact reason.

https://www.microsoft.com/en-us/sharedsource/

Re: GitHub Copi­lot inves­ti­ga­tion

#939

I get the impression that many peoples' grievance with generative AI (text, code, images etc.) isn't _really_ about the data provenance. Or at least, it feels secondary, compared to the general disruptive nature of the tech. If tomorrow someone released a StableDiffusion, CoPilot etc with the same functionality, but respecting the provenance of the data (i.e. licensing etc), what concrete difference would this make?…

Sometimes when working on software the goal is not to get a competitive advantage, but to promote some ideas. Copyleft license is a tool that aims to help with granting, that the work derivatives are available to the public to read and modify (i.e to prevent closing the source code of the program that is commercially sold).

The concrete difference you ask about is that the work derived from copyleft code retains the license and the source code can't be closed. If you scrap the license, then the code created by someone who had clear goal in mind when writing it for not making improvements over it closed source, ends up with possibility of being closed source.

Re: GitHub Copi­lot inves­ti­ga­tion

#940

Earlier quoted context omitted.

I really appreciated the argument in the book "The New Breed," which is that we should adapt ideas around the governance of animals to governance of ML: You can train your dog to attack random passersby, but if you do, you're a monster and ultimately responsible for the dog's actions. Likewise, you can tell Copilot to crank out code specific algorithms written by specific people, but if you do so, you're still creati…

This just sounds like blaming the researchers to me. How would i ever know if my "boring code completion" was actually copyright infringement? Your argument just disallows discussing the problem while doing absolutely nothing about it. If you train your dog to NOT attack random passersby and it still does, that dog is euthanized no matter your intentions.

"Not knowing" doesn't free you from responsibility.

If you took a bunch of copyrighted and non-copyrighted books, cut them into pieces, shuffled them all together, then picked a passage at random from a hat; "not knowing" what you are going to get doesn't mean you aren't violating copyright.

That's essentially what copilot is doing: it's taking a bunch of code - some of it copyrighted without license - and using it as a dataset. The ML algorithm then tries to pattern match against that data to provide the user with something they want. That's just copyright violation lottery with extra steps.

Post reply on HN