Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

481–490 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#481
post #122
post #79

Earlier quoted context omitted.

why use the API? why not just use git to get the code? All you need the API for is repository discovery

I'm gonna guess that Microsoft GitHub (tm) would shut you down pretty quickly if you tried to clone tens or hundreds of thousands of repos in a short window of time, b/c of course that's sketchy/abusive use of their infrastructure, right? But of course if the data is already sitting in object storage inside your cloud environment and all you have to do is run some MapReduce jobs to get at it... Hence: unfair, anticom…

Actually they don’t. I’ve cloned thousands of repos before (tried to archive conda-forge org for a project).

I’ve also built many parallel repo downloaders for CI reasons. You can clone repos all day pretty much with little rate limiting. I haven’t pushed parallelism past 64 per host though

Re: GitHub Copi­lot inves­ti­ga­tion

#482
post #453
post #432

Earlier quoted context omitted.

It says why in the linked post. People aren't doing open source for free; they do it for the community. But Copilot is there to extract value from it, giving nothing back, not even credit.

Can't open-source programmers improve their own open-source code with Copilot? Does the inherent improvements that all Copilot offers just not apply to people who write open-source code? I understand that there is a balance, but as an open-source advocate who would love better tools to make their open-source projects better I'm lost as to why this point doesn't counter the "giving nothing back" we hear so often.

Since there's no way to know how code generated by Copilot might be licensed without expensive code-scanning tools, I don't think OSS can safely derive any substantial improvements from it.

Re: GitHub Copi­lot inves­ti­ga­tion

#483
post #302
post #190

Earlier quoted context omitted.

You're not supposed to be able to use dominance in one market (git hosting) to gain dominance in another (AI powered code suggestions).

You might be confused but that's literally what you do as a business. You leverage your domain area expertise to expand into new areas of business. For instance, Apple already knew about the portable hardware market and extended their reach into the portable music market via iPod. It used the iPod to reach the music marketplace via iTunes. Used its market dominance to create iPhone and the rest is history. Maybe that…

I think the parent comment is referring to https://en.wikipedia.org/wiki/Tying_(commerce).

TLDR: If you have a monopoly in a market, you can't use your position to get a heads up in another distinct market by bundling products together. Ex: Microsoft can't use their monopoly position in the OS market to get an advantage in browsers by bundling IE with Windows.

It's my understanding that this doesn't apply if you don't have a monopoly, and it also doesn't apply unless you're actually bundling the sale of multiple things together.

Doesn't seem to be relevant here IMO.

Re: GitHub Copi­lot inves­ti­ga­tion

#484
post #185

Earlier quoted context omitted.

> It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus. I wonder how many people on HN would be on the side of the creators if we were talking about content created by Walt Disney and whether pirating was ethical?

It'd probably turn towards Disney's history of bribing congress for copyright extensions whenever the mouse is about to enter the public domain.

That's a bit of an exaggeration. There have only been two copyright term extensions since Mickey Mouse was created, and only one of those can even remotely be attributed to Disney lobbying.

Re: GitHub Copi­lot inves­ti­ga­tion

#485

Earlier quoted context omitted.

You aren't allowed to just read code and regurgitate it in order to claim it as your own. That is, just because you memorized this great new novel you read, it doesn't mean you can go and sit down and hammer it out and sell new copies. People go to great lengths to do this sort of things (see: clean room reverse engineering [1]) in order to try and wash themselves of liability. [1] https://en.wikipedia.org/wiki/Clean…

If the code was purely utilitarian in nature, such as something that was optimized for execution time, there is plenty of precedent stating that the code in question is not covered by copyright. Do an internet search for “copyright utilitarian” and read up on it if you don’t believe me! Copyright is about protecting artistic expression which is held in contrast to the useful nature of a work.

Note: In the US, this concept is explicitly in the Copyright Act:

"In no case does copyright protection for an original work of authorship extend to any idea, procedure, process, system, method of operation, concept, principle, or discovery, regardless of the form in which it is described, explained, illustrated, or embodied in such work." (17 USC 102(b) [0]).

See also the "Useful Articles" doctrine. [1]

[0] https://www.law.cornell.edu/uscode/text/17/102

[1] https://en.wikipedia.org/wiki/Copyright_law_of_the_United_St...

Re: GitHub Copi­lot inves­ti­ga­tion

#486
Here are a few thoughts I haven't formulated before:

It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human culture (grammar, concepts, facts, cultural tropes, and many others).

A human being doesn't violate copyright in learning from a copyrighted work, including when that human being is later more able to produce other works based on that learning (e.g. reading fantasy novels and learning concepts, tropes, or vocabulary that one uses to produce other fantasy novels; reading a newspaper and learning facts that one incorporates into an essay; learning artistic techniques or stylistic conventions from studying existing artworks and using them when producing new artworks). Current AI systems are (amazingly) becoming capable of all of these things and may do them in ways that are somewhat akin to how human beings do them. (although I guess Jaron Lanier would object "that's what they want you to think")

But there are also examples in existing copyright doctrine where people accidentally repeat enough of a prior work to get in trouble for infringement -- most often with song composition (like George Harrison's "My Sweet Lord") because relatively small pieces of melody (which a person might easily memorize) may be considered copyrightable.

If human beings had much more accurate memories, copyright would be quite a bit more intrusive (and/or quite a bit less effective) because, following any exposure to some kinds of works, we could use our own memories to reproduce those entire works from scratch for our own use or pleasure without obtaining authorized copies from elsewhere.

Computers do have such accurate memories, and machine learning systems, which are optimized for things like maximum likelihood estimation, can and do reproduce both copyrightable and non-copyrightable elements of works that they've been trained on. After all, the maximum likelihood continuation of a fragment of a text or a song is ... the complete original work. And the ability to reproduce the complete original work would, other things being equal, reduce loss in training. After all, that's something someone might specifically ask for, and if the system could oblige, it would be doing a better job of providing what the user wanted.

It's relatively foreseeable that machine learning systems would potentially be able to reproduce both copyrightable and non-copyrightable elements of various works, because the distinction between the two isn't especially clear from an algorithmic or mechanical point of view. (For instance, facts aren't copyrightable, but the notion of what constitutes a "fact" for this purpose is a culturally-bound legal notion and not at all straightforward to make precise.)

But if you had a human author or artist or scholar or programmer who was "trained on" exposure to an enormous body of works, and that person had an exceptional eidetic memory, you could imagine that he or she would be perfectly capable of recreating many of those works from memory (and that other people might request such recreations). (Again, in music in particular, it's already routine that someone could have unambiguously copyrightable material memorized and be subject to copyright restrictions on performing songs. Like if a singer or band performs a cover from memory.)

If you wanted to avoid this ability then you might need to build in an explicit notion of copyright that limits the accuracy or level of detail inside of the model in some way. This is tricky because (1) I don't think people have really tried to do this much so far, (2) copyright applies very differently to different categories of work, (3) it obviously wouldn't satisfy critics even if it mitigated the most extreme examples of "regurgitation", and (4) it would be kind of weird because you would be intentionally limiting the quality and extent of learning that the system was allowed to do. (I imagine Jaron Lanier getting mad again about my repeated comparison between human learning and machine learning, and between human memory and machine memory)

Some of the weirdness in point (4) is that accurate prediction is usually cool / great / impressive / accepted as an appropriate goal or capability, but if it's too accurate in certain contexts, it may be deemed a copyright infringement. Like if you said "what word comes next? FOUR SCORE AND SEVEN YEARS AGO OUR FATHERS", there's a clear correct answer and knowing it requires having a certain text memorized. OK, if you said "what word comes next? MR. AND MRS. DURSLEY OF NUMBER FOUR PRIVET DRIVE WERE PROUD TO SAY" ... same thing, but Bloomsbury Publishing may be unhappy if you have a system that can get all such questions right.

Re: GitHub Copi­lot inves­ti­ga­tion

#487
post #180

Earlier quoted context omitted.

Not a supporter of Copilot, but I think it's pretty easy to access the same data through BigQuery: > The Google BigQuery Public Datasets program now offers a full snapshot of the content of more than 2.8 million open source GitHub repositories in BigQuery. Thanks to our new collaboration with GitHub, you'll have access to analyze the source code of almost 2 billion files with a simple (or complex) SQL query. https://…

There is a distinction between being able to access the source code, and a tool giving it to you without any context of the underlying license it is governed by.

GP was saying that GitHub has an unfair advantage in that they have instant access to all GitHub code, whereas everyone else is rate limited.

I'm pointing out that this limitation is not meaningful because everyone can access all GitHub hosted source code through BigQuery, where they won't be rate limited.

I'm not comparing BigQuery to Copilot.

Re: GitHub Copi­lot inves­ti­ga­tion

#488

One issue I see with Copilot is that they get free access to all open-source data on GitHub, but using GitHub APIs to download the data yourself isn't possible (rate limiting). This is an unfair advantage. Copilot is not only making money off of open-source, they are making money off of open-source in a way others can't. I would love to see a lawsuit which requires GitHub to provide their full Copilot dataset.

That’s the problem with big tech. When they own all the roads, it makes organic competition almost impossible, as they will always be 2 steps ahead.

Not to be anti-capitalist, but you’re describing capitalism lol

Re: GitHub Copi­lot inves­ti­ga­tion

#489
post #163

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

> "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" https://twitter.com/docsparse/status/1581461734665367554 An english description plus three characters of a function name is enough to coax CoPilot into distributing LGPL-licensed code out of context, without a proper license. That's neither "emotional" nor "cherry-picked", it's a clear-cut license violation.

It’s a contrived example. This is not how people use copilot in the real world. Copilot is designed to be used in context. The completions it provides in context are specialized to your project.

Artificially starving copilot of context and then showing that it recites parts of its training dataset is mundane.

Re: GitHub Copi­lot inves­ti­ga­tion

#490
post #247

There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out . The latter is obviously a violation of copyright, full stop. The former, to me, is obviously not a violation. If it were, that would massively tilt the playing field in favor of large corporations. It would become very hard to independently train your own models. Philosophically, I go by the principl…

> There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out.

> The latter is obviously a violation of copyright, full stop.

It's not obvious to me that (2) is a violation of copyright. Unlike patents, copyright violation is not as simple to prove. My understanding is that, at least in the US, independent creation is a valid defense against copyright infringement. For example if 2 people independently write the same story and can prove that they did, they can both hold copyright over that story.

The analogue to this does exist without AI, when creating something that looks like copyright infringement, clean room design (don't look at similar things) is often done to ensure that "independent creation" can be used as a valid defense in court. Given that, I think (1) is probably not safe to do at all if you can't prevent (2).

Post reply on HN