Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

221–230 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#221

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

I think I agree with this comment[1] from the other thread; never previously thought that a process being transformative means input and output datatypes do not coincide, but maybe that is it. 1: https://news.ycombinator.com/item?id=33240681

That's an interesting take, the whole level of indirection thing between Microsoft-OpenAI and StabilityAI and that research group is certainly another dimension to this that sort of muddies the waters.

Re: GitHub Copi­lot inves­ti­ga­tion

#222
It seems that Copilot could address this issue by searching for matches in its source repositories for the strings it generates, with appropriate criteria, and give the user a link describing the origin of the code, who wrote it, and what the license is for cases where a match length exceeds a threshold. So, you wouldn't just get the Quake fast integer square root routine, you'd get a pointer to the Quake repository and license info from which it came. A separate model could be trained up that would find the closest match in source code repositories. A user could then use Copilot safely, attribute code correctly, and avoid code with incompatible licenses.

This would be a better approach than "shut it down".

Re: GitHub Copi­lot inves­ti­ga­tion

#223
post #108

Earlier quoted context omitted.

It’s not illegal, it’s at worst a fancy code search tool that Github has the right to show you the results via the license you grant them when you upload and make public code on Github which is way stronger than other search engines like Sourcegraph have to show public code. It doesn’t mean you have the right to use any of the code it generates but Copilot itself isn’t illegal in any meaningful sense.

This is definitely not true. When your license requires you bundle said license with any reproductions of the code, and Copilot spits out said code sans license, they are breaking the law.

> We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.

I don’t think it’s accidental that this product is specifically Github Copilot.

But even then I think this is legal overkill. If you use the search box on Github they will display snippets of code from public repositories without the license. Same as what Sourcegraph does same as Copilot does. Nobody here is arguing ripgrep is violating the license by displaying matches without the corresponding license.

Re: GitHub Copi­lot inves­ti­ga­tion

#224
post #185

Earlier quoted context omitted.

> It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus. I wonder how many people on HN would be on the side of the creators if we were talking about content created by Walt Disney and whether pirating was ethical?

It'd probably turn towards Disney's history of bribing congress for copyright extensions whenever the mouse is about to enter the public domain.

The same copyright laws that the open source proponents are complaining about MS breaking?

Re: GitHub Copi­lot inves­ti­ga­tion

#225
I wonder if people realize that letting GitHub train Copilot on their open source contributions is effectively de-valuing your own time, which (if repeated at a larger scale) devalues your experience, and that eventually has the effect of reducing the correlation between your experience and your salary.

For example, if an overseas firm can just as easily use Copilot as I can write original code (or use Copilot myself), why would any company hire me locally?

Re: GitHub Copi­lot inves­ti­ga­tion

#226
post #83
post #52

I'd be rather saddened if Copilot was shut down or neutered because of a few vocal few protesting against it. It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.

I’d claim that it’s more than a “few vocal” protestors. If the system is illegal, it needs to become legal or disappear. If I’m writing code for a query optimizer, the SQL Server solution isn’t going to magically show up.

Is there any evidence that the PostgreSQL or MariaDB solution will though?

Re: GitHub Copi­lot inves­ti­ga­tion

#227
post #150

What's with the default to "if it's not explicitly legal, it must be illegal"? Imagine if every new piece of software your wrote had to be tested for legality because you don't know that it's explicitly legal. Oh there aren't laws for this new thing, so I guess you should challenge yourself all the way to the supreme court? I get the author not liking Copilot, but I don't see that GitHub/Microsoft have any kind of ob…

Its better to know, even if you like Copilot and want it to continue.

> I get the author not liking Copilot, but I don't see that GitHub/Microsoft have any kind of obligation to figure this out just because they're GitHub/Microsoft.

Because its a trillion dollar company with an infinite amount of lawyers and legal resources?

Re: GitHub Copi­lot inves­ti­ga­tion

#228
post #52

I'd be rather saddened if Copilot was shut down or neutered because of a few vocal few protesting against it. It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.

Careful though, you are trading yours (and their) muscle memory and brainpower to be locked into a proprietary solution.

Reread your post. Doesn't it sound scary? You are blocked from even thinking and crafting because a specific web service is down.

Even if Google is down you can go direct to Stackoverflow and MDN, and have a choice of information sources.

Also what is "productivity" ... as in features built / month or lines of code / month?

Re: GitHub Copi­lot inves­ti­ga­tion

#229

Earlier quoted context omitted.

You aren't allowed to just read code and regurgitate it in order to claim it as your own. That is, just because you memorized this great new novel you read, it doesn't mean you can go and sit down and hammer it out and sell new copies. People go to great lengths to do this sort of things (see: clean room reverse engineering [1]) in order to try and wash themselves of liability. [1] https://en.wikipedia.org/wiki/Clean…

If you think most people pay any attention to licenses or respect them you better think again. Snippets get copied verbatim with no regard to their source all the time. Licenses have no power and are routinely ignored.

The point is not if the law is actually upheld or not. Its if it is legal or not.

Re: GitHub Copi­lot inves­ti­ga­tion

#230

My view is the copilot is not stealing open source code. It is learning from it just as a human reader would. People's disguste is based on the assimilation of what they thought was a human trait being machine derived from their work. The copilot service backed by an army of actual humans wouldn’t be a story at all. Nor would anyone be angry, if an individual offered coding skills as a service, and had gone through t…

You aren't allowed to just read code and regurgitate it in order to claim it as your own. That is, just because you memorized this great new novel you read, it doesn't mean you can go and sit down and hammer it out and sell new copies. People go to great lengths to do this sort of things (see: clean room reverse engineering [1]) in order to try and wash themselves of liability. [1] https://en.wikipedia.org/wiki/Clean…

If the code was purely utilitarian in nature, such as something that was optimized for execution time, there is plenty of precedent stating that the code in question is not covered by copyright.

Do an internet search for “copyright utilitarian” and read up on it if you don’t believe me!

Copyright is about protecting artistic expression which is held in contrast to the useful nature of a work.

Post reply on HN