Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

951–960 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#951

Earlier quoted context omitted.

If this is MS trying to pull off EEE, what does the extinguish phase look like? That they try to make it so that any codebase the uses copilot is owned by them, and that there's no way to turn it off because all other editors or code hosting sites will exist? Plausible I suppose if they play the game for several decades and somehow no one else produces any innovation in the space.

Extinguish might look like Microsoft or their customers/partners writing proprietary replacements for open source products with the help of copilot. I don't know how likely this is, but what co-pilot provides is a plausible path for leveraging open source code to create closed-source products. Over time this allows the proprietary software industry to contribute back less code while still benefitting enormously.

That would mean that Co-pilot is, or at least in big part, a front, a false flag operation to test the legal system's tolerance of what they are doing, to determine whether they can get away with what they are doing, right?

Re: GitHub Copi­lot inves­ti­ga­tion

#952
post #950

Earlier quoted context omitted.

If this is MS trying to pull off EEE, what does the extinguish phase look like? That they try to make it so that any codebase the uses copilot is owned by them, and that there's no way to turn it off because all other editors or code hosting sites will exist? Plausible I suppose if they play the game for several decades and somehow no one else produces any innovation in the space.

I'm not too sure yet, but I wouldn't be surprised if we wake up one day, and just like how it went with Facebook buying oculus, we will all of a sudden require some "microsoft account" to log into Github. Then more layers, and more, until there is nothing left. Either that, or you wake up one day to see that Microsoft have stole your open source software and Microsoft says "but muh AI".

It's a new approach compared to Amazon taking open-source code and building AWS services that kill off attempts at self-funding (dual licensing/support services) by the people who made it. I hope it's not as successful. Amazon at least abided by the letter of the licenses.

>> "I'm not too sure yet, but I wouldn't be surprised if we wake up one day, and just like how it went with Facebook buying oculus, we will all of a sudden require some "microsoft account" to log into Github."

See: Minecraft. It was already a goldmine when they bought it, but they built it into an even bigger one before forcing millions to have a foot into their ecosystem. Copilot might be their way of making everyone dependent on GitHub before "moving on" from git and offering a Community Edition of their own source control system.

It'll be easy. A lot of people hate Git.

Re: GitHub Copi­lot inves­ti­ga­tion

#953

Earlier quoted context omitted.

Being a useful tool doesn't make it legal.

Technical progress takes precedence over pitiful intelectual property discussions. If you don't believe that, I am not sure what you are doing in a community like this.

You are not the authority on what this (or any other) community is about. Fetishizing "progress" (whatever you think this means) is next level idiocy.

Re: GitHub Copi­lot inves­ti­ga­tion

#954
post #486

Here are a few thoughts I haven't formulated before: It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human culture (g…

A human being doesn't violate copyright in learning from a copyrighted work, including when that human being is later more able to produce other works based on that learning (e.g. reading fantasy novels and learning concepts, tropes, or vocabulary that one uses to produce other fantasy novels;

Yes, but a human being isn't allowed to copy that work before learning from it, even if they destroy the copy afterward. AIs don't watch Youtube or browse Github. People download copies of content stored there, analyze and categorize it, and then feed it into AIs. Copyright is broken at step 1.

Re: GitHub Copi­lot inves­ti­ga­tion

#955

So happy to learn of this and I wish them best of luck in their efforts. And I'm surprised to find so many people klinging to Copilot. We shouldn't shed any tears for a megacorporation which shows such blatant disregard for the licensed works of people's labour. Yes, AI is here to stay but we should be able to build AI that respects copyright. Yes, it's easier to just steal data and call it fair use. Whether or not t…

Foolish take. If ML training is not fair use then all ML progress is dead in the water. ML training is akin to reading or learning, and licenses do not apply to that. You’re not thinking past “megacorp = bad”.

AI progress won't be dead in the water if it respects copyright laws. Yes, being free to just freely grab any data is infinitely easier. But having to rely on properly licensed datasets or asking users for consent should be the norm for ML development IMHO.

Also, If we had trained some A.I. on the Windows codebase and started freely using suggestions given by it I bet Microsoft would scream copyright infringement in a heartbeat.

Re: GitHub Copi­lot inves­ti­ga­tion

#956
post #872

Earlier quoted context omitted.

Like posts have said, whether training is fair use is not a matter of opinion, it is a matter of law. You can't use an appeal to your authority to make grand statements like this. Frankly, I don't care that ML/AI _needs_ this to work. That's not my problem. You don't get to circumvent existing agreements (and law) because you believe that ML learning is the same as a human reading a piece of code and then typing it u…

There isn’t law any law yet. And yes, it is the same thing as learning. If you have a robot that learns like a human does … you think it should be illegal for that machine to look at GitHub? To watch a Hollywood movie?

A human that watches a Hollywood movie and then goes on to recreate it frame-by-frame with, idk, everyone has cat ears and go "nah this is all my original creation" is an idiot. A human that watches a Hollywood movie and then goes on to create existing works within the genre, with some homage (say, a specific hat, or a specific framing of a pivotal scene, or a specific lighting choice) to the original movie that inspired them, is learning.

Re: GitHub Copi­lot inves­ti­ga­tion

#957

Earlier quoted context omitted.

when you put your code on GitHub.com, you grant GitHub the right to show that code to others. https://docs.github.com/en/site-policy/github-terms/github-t... this is separate from the license you specify in the repository and you can't revoke it without removing your code from github.com.

From the text: >We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you an…

I would say they've used code they host to train an AI and they charge for a fraction the GPU time required to train and customize their model. they're selling you their model, not the code it produces.

if this is indeed how they charge for Copilot, and I don't know if it is or is not, then they will need to show that they have done their due diligence in making sure that code is not reproduced verbatim when a user requests that it not reproduce code verbatim.

I'm quite sure that GitHub can defend Copilot in court. That's part of the process of offering a new feature to customers; making sure that it is legal and defensible to do so.

All of the armchair attorneys here who think they know better than GitHub's attorneys when operation of the service puts GitHub's ass on the line is ... I wish I had 1 percent of that confidence. I would be a thousand times more confident than I am now.

Re: GitHub Copi­lot inves­ti­ga­tion

#958
post #343

Earlier quoted context omitted.

Yes, we’d all find our work easier if we could just steal other people’s work.

You're probably stealing other people's code daily, willingly or not. Should we run a plagiarism scanner on all your code?

Yes, I want repositories / libraries that steal code taken down, just as GitHub Copilot should be, so I don't unwillingly steal code.

Re: GitHub Copi­lot inves­ti­ga­tion

#959
post #642

Earlier quoted context omitted.

> which is what copilot has been show to sometimes do In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd bet rare things that are representative of some conceptual axis). The main examples presented so far, Davis and Carmack, have the property of having been copied a lot. The generative model…

That seems like a weak defense: "sure, we violated copyright but only because many other people do, too". Kinda the same problem as YouTube. Lots of people copy movies on the high seas, but if you are as big as yt you cannot easily get away with it.

Without a sample case I can't say for certain, but couldn't it also be a defense that some code is generic enough that it shouldn't be copyrighted?

Re: GitHub Copi­lot inves­ti­ga­tion

#960
post #811

Although I'm aware that this tool is a boon to many, particularly those with impediments like RSI, I still have to echo what a number of other comments say: There really is a very large proportion of adult software developers in the market who are simply too young to have lived through the EEE Microsoft era. Add on to that the proportion of old-enough Microsoft-brand "dotnetter" software developers who simply don't c…

If this is MS trying to pull off EEE, what does the extinguish phase look like? That they try to make it so that any codebase the uses copilot is owned by them, and that there's no way to turn it off because all other editors or code hosting sites will exist? Plausible I suppose if they play the game for several decades and somehow no one else produces any innovation in the space.

Embrace: How nice that we don't have to worry about GitHub closing for lack of funds. Thanks, Microsoft!

Extend: Make people dependent on GitHub Copilot. Require a Microsoft account (soon).

Extinguish: Sunset git and transition to Microsoft's own source control system.

Post reply on HN