Live data from Hacker News

All public GitHub code was used in training Copilot

twitter.com

691–700 of 734 posts

Re: All public GitHub code was used in training Copilot

#691
post #551

To me, the particular use case and whether it is fair use or not, is of minor interest. A far more pressing matter is at hand: AI centralization and monopolization. Take Google as an example, running Google Photos for free for several years. And now that this has sucked in a trillion photos, the AI job is done, and they likely have the best image recognition AI in existence. Which is of course still peanuts compared…

Well of course only the huge companies can develop products that require enormous resources.

But I'm not too worried here because everyone gets access to larger datasets every year, and it gets cheaper to process every year, so whatever Microsoft or Google is capable of doing now, smaller companies will be capable of doing in a few years.

Re: All public GitHub code was used in training Copilot

#692
post #548

Earlier quoted context omitted.

That already exists though? SongSmith and other similar tools are used by musicians a lot.

At what point is it not a derivative work?

afaik, chord progressions aren't copyrightable, and even some lyrical things aren't. Melodies are the main thing, I believe. (I could be wrong, this is just what I have been told in the past)

Re: All public GitHub code was used in training Copilot

#693
post #302

Earlier quoted context omitted.

If it's large sections, that can be fixed by either licence attribution or result filtering. That's at best a technical issue. What way too many people claim, however, is that the machine isn't even allowed to look at GPL'ed code for some reason, while humans are. I'd like to learn the reasoning behind that.

> What way too many people claim, however, is that the machine isn't even allowed to look at GPL'ed code for some reason, while humans are. Why would those be the same thing? It's a matter of scale. Just like how people are allowed to read websites, but scraping is often disallowed.

> Just like how people are allowed to read websites, but scraping is often disallowed.

Hosting code on Github explicitly allows this type of usage (scraping) according to their TOS so I have to ask again - why the sudden complains?

Are we still talking about a shortcoming of the ML model, which very occasionally spits out a few lines of copied code or should we include search engines into this, because they do the exact same thing by design?

robots.txt, for example, has a non-binding, purely advisory character as well and Common Crawl [0] (also used for training GPT-3) publishes a dataset that by definition contains GPL'ed code as well, no matter where it's hosted. So is that off-limits now, too?

[0] http://commoncrawl.org

Re: All public GitHub code was used in training Copilot

#694
post #676

Earlier quoted context omitted.

> If we predict that ultimately AI will change virtually every aspect of society, these companies will become omnipresent, "everything companies". God companies. What we currently call AI is very from AGI, and it's not clear that sitting on piles of proprietary data gives an edge towards AGI. If the goal is human-level intelligence, that has been demonstrably achieved with the far lesser resources of the public schoo…

> If the goal is human-level intelligence, that has been demonstrably achieved with the far lesser resources of the public school system. Pretending that the scientifically managed public school system, that attempts to manufacture uniform educated humans on a conveyer belt, is responsible for human education is fairly ridiculous. Children have a remarkable capacity to learn, and do so automatically through free play…

Yeah everyone knows children will innately learn calculus from flinging mud at eachother.

Re: All public GitHub code was used in training Copilot

#695
post #676

Earlier quoted context omitted.

> If we predict that ultimately AI will change virtually every aspect of society, these companies will become omnipresent, "everything companies". God companies. What we currently call AI is very from AGI, and it's not clear that sitting on piles of proprietary data gives an edge towards AGI. If the goal is human-level intelligence, that has been demonstrably achieved with the far lesser resources of the public schoo…

> If the goal is human-level intelligence, that has been demonstrably achieved with the far lesser resources of the public school system. Pretending that the scientifically managed public school system, that attempts to manufacture uniform educated humans on a conveyer belt, is responsible for human education is fairly ridiculous. Children have a remarkable capacity to learn, and do so automatically through free play…

> Pretending that the scientifically managed public school system,

Say what now? There may be places on Earth that practice scientific management, there are definitely some that pretend to, but IME public school systems are neither.

Re: All public GitHub code was used in training Copilot

#696

Earlier quoted context omitted.

If you do get sued, the Copilot page is written in a way that would make Github legally responsible for it, not you. "Just like with a compiler, the output of your use of GitHub Copilot belongs to you."

Yeah, right... This isn't going to fly in court any more than if the Pirate Bay page was written in a way that says that it's solely responsible for what you do with the magnet links that they share.

The pirate bay is very clear to not claim any responsibility for what people post on their site. That's how they get away with it.

Re: All public GitHub code was used in training Copilot

#697

Earlier quoted context omitted.

> If the goal is human-level intelligence, that has been demonstrably achieved with the far lesser resources of the public school system. Pretending that the scientifically managed public school system, that attempts to manufacture uniform educated humans on a conveyer belt, is responsible for human education is fairly ridiculous. Children have a remarkable capacity to learn, and do so automatically through free play…

> Pretending that the scientifically managed public school system, Say what now? There may be places on Earth that practice scientific management, there are definitely some that pretend to, but IME public school systems are neither.

Schools (at least American public schools) are one of the last bastions of Taylorism in the west. They treat students like uniform widgets on an assembly line.

You can read for yourself: https://files.eric.ed.gov/fulltext/ED566616.pdf https://radicalpedagogy.icaap.org/content/issue3_2/rees.html

Re: All public GitHub code was used in training Copilot

#698

Earlier quoted context omitted.

You don't encrypt your data before uploading to backblaze?

Oh heck no, I never encrypt data. I run windows. It can't ever be secure, anyone who wanted to hack me could. Scrambling the data really makes things worse as any accident requiring recovery of my data is also probably going to lose the encryption key. The only time I ever lost any significant chunk of data (a persons lifetime set of photos!) was because Windows encrypted data at rest, and thus it couldn't be recover…

Agreed. Everybody talks about encrypting backups like it's common sense, but almost nobody talks about the risks involved with failing to back up the encryption key itself properly. The entire integrity of the backup then depends on that sensitive piece of data, and it's not something that can be openly shared by its nature, or included in the encrypted backup itself. It's even deceptive if your measure for success is restoring the backup to make sure it works properly, because there is now an implicit assumption that the encryption key is still valid and undamaged the next time you restore.

I wish backup tools like Duplicity would warn you about the risks of encrypting backups instead of warning the user if they disable encryption, because encryption has the possibility of rendering all those backups useless when the moment to use them finally comes.

I have a similar feeling that large swathes of my digital life would be rendered permanently inaccessible if 2FA was enabled and my device was rendered inoperable. (That's why I keep meticulous physical backups of emergency keys.) I think 2FA and the like should be considered a tradeoff with its own inherit risks and benefits, instead of a universally better option than randomly generated 80-character passwords alone.

Re: All public GitHub code was used in training Copilot

#699

Earlier quoted context omitted.

> Pretending that the scientifically managed public school system, Say what now? There may be places on Earth that practice scientific management, there are definitely some that pretend to, but IME public school systems are neither.

Schools (at least American public schools) are one of the last bastions of Taylorism in the west. They treat students like uniform widgets on an assembly line. You can read for yourself: https://files.eric.ed.gov/fulltext/ED566616.pdf https://radicalpedagogy.icaap.org/content/issue3_2/rees.html

“Treating X as uniform widgets” (where X are not uniform widgets) and “scientific management” are not only not the same thing, they are anticorrelated.

Re: All public GitHub code was used in training Copilot

#700

I am really confused by HN's response to copilot. It seems like before the twitter thread on it went viral, the only people who cared about programmers copying (verbatim!) short snippets of code like this would be lawyers and executives. Suddenly everyone is coming out of the woodworks as copyright maximalists? I know HN loves a good "well actually" and Microsoft is always suspect, but let's leave the idea of code la…

Perhaps people on HN start sensing that successors of Github Copilot will take their programming job. Rightly so. Personally, I think that in the age of AI programming any notions of code licensing should be abolished. There is no copyright for genes in nature or memes in culture; similarly, these shouldn't be copyright for code.

People aren't happy because Microsoft is exploiting open source. They're training it on open source code and keeping the service for themselves.

If they made the trained model public (and also trained it on private code) the response would be completely different.

Post reply on HN