Live data from Hacker News

GitHub Copilot as open source code laundering?

twitter.com

441–450 of 473 posts

Re: GitHub Copilot as open source code laundering?

#441

Earlier quoted context omitted.

this complicated copyright problem shows we're still using last century concepts on new and emerging technology that surpassed it; it's time to think hard about it because we need neural nets and they need training data

Some are more equal than others though, aren't they? I mean, if MS throws out licensed code from others, as if to say: "Ahh, software licensing, such an outdated concept ..." but then keeps its own code out of that loop. "Yeah, but that's our own code, no one is allowed to copy that!"

I doubt they will corner the market for AI code assistants. ML models are replicated or surpassed in a few months by the competition. We will all benefit from them, it won't remain concentrated in a few hands.

Re: GitHub Copilot as open source code laundering?

#442

Earlier quoted context omitted.

Don't forget settlements paid in Ai-generated crypto-currencies backed by Gold mined in Australia fully automated mine. Run it all on solar and humans can just fuck right off.

The market ultimately obeys customer demand, so all these problems will be sorted out... until customer AI.

That's the next step for Amazon Prime: 0-click shopping. Just buys stuff from your recommendations every month and sends it to you.

Re: GitHub Copilot as open source code laundering?

#443

Attempts to litigate any license violation are going to get precisely nowhere I bet, but I find the actual license violation argument persuasive. This is an excellent example of how the AI singularity/revolution/whatever is a total distraction and that a much bigger and more serious issue is how AI is becoming so effective at turning the output of cheap/free human mental labour into capital. If AI keeps getting bette…

Do we need an update of free software licenses to specifically address this?

Unlikely. If this use counts as a derivative work, then it's already a violation, and no update is needed.

OTOH if laundering through machine learning is a fair use, then licenses can't do anything about this. Licenses can't override the copyright law, so the law would have to change.

Re: GitHub Copilot as open source code laundering?

#444

For years people have warned about hosting the majority of world's open source code in a proprietary platform that belongs to a for profit company. These people were called lunatics, fundamentalists, radicals, conspiracy theorists, and many other names. Well, they were ignored and this is the result. A for profit company built a proprietary system using every code hosted in its platform without respecting the code li…

If we feed the entirety of a library to an AI and have it generate new books, is it an exploitation of people's work? If we read a book and use its instructions to build a bicycle, is it an exploitation of people's work? No, no it's not.

I think you are abstracting the matter by taking out the humanity. It's one thing to learn to do something by hand after purchasing the book. It's a totally different thing to read every single book in the world (humanly impossible) and then absorb some knowledge and train yourself to write exceptional books because you (the AI in this scenario) have learned that some words and sentence structures have lead to books having higher ratings than others. It's not humanly possible.

Of course we generate the world around us and its rules but I get angry every time we compare people to machines and say that it's the same thing. No it's not. We are constrained by time and space. I can't add more brain or more eyes to my body so I read more books can I? Microsoft can have a small city of servers somewhere and that could replace lots of people's jobs.

Re: GitHub Copilot as open source code laundering?

#446

Earlier quoted context omitted.

Oh man, that got meta super fast. Its like a mobius strip!

The nice thing about co-pilot is that it will suggest to do the same mistakes as in other software. If you accept all autosuggestions in C++ you might end up with Windows.

This is such a ridiculous statement to me. If this were a real problem we would have noticed by now with stackoverflow. I truly believe the vast majority of capable developers read, understand and test code they copy from somewhere. This is even more obvious with an AI that will never suggest 100% correct code all the time.

Re: GitHub Copilot as open source code laundering?

#447
post #246

Earlier quoted context omitted.

> 100 lines in 100k program The intention is autocomplete boilerplate, not write a kernel.

This is not a difference in kind. Autocomplete, do you have anything to say to the commenter ? “This isn’t the best thing to say.”

How is designing a very large system even close to the same thing as writing a few small functions? That's like saying an architect designing a building is doing the same thing as a brick layer putting down cement.

Re: GitHub Copilot as open source code laundering?

#448
post #101

One interesting aspect, that I thing will make it difficult for GitHub to argue and justify its not a a license violation would be the answer to the following question: Was Copilot trained using Microsoft internal source code or will it be in the future ? As GitHub is a Microsoft company and OpenAI although a non-profit just got a massive one billion investment from Microsoft (presumably not for free), will it start…

Late to the thread but: OpenAI is not a non-profit since 2019 (technically they call it a capped profit company [1], but until the singularity you can ignore the cap). I guess this does impact the dynamic with Microsoft

[1] https://openai.com/blog/openai-lp/

Re: GitHub Copilot as open source code laundering?

#449

Earlier quoted context omitted.

I'm not arguing that machines will be more efficient than human brains. A airplane isn't more efficient than a goose. But airplanes do fly faster, higher and with more cargo than any flock of geese could ever carry. Similarly, there is no contradiction between AI being less efficient than a human brain, and AI being preferable to humans because it can deal with data sets that are two or three orders of magnitude too…

Even so, such AI doesn’t exist. All the AIs that exist today operate by fitting data. And to be able to perform a useful task it has to have well defined parameters and fit the data according to them. I’m not sure an AI that operates outside of these confinements have even been conceived of. To make an AI that outperforms humans in any task has not been proven to be possible (to my knowledge) not even in theory. An a…

> All the AIs that exist today operate by fitting data. And to be able to perform a useful task it has to have well defined parameters and fit the data according to them. I’m not sure an AI that operates outside of these confinements have even been conceived of.

Such an AI has absolutely been conceived of. In Superintelligence: Paths, Dangers, Strategies, Nick Bostrom goes over the ways such an AI could exist, and poses some scenarios about how a recursively self-improving AI could "take off" and exceed human intellectual capacity on its own.

Moreover, we're already building such AIs (in a limited fashion). Deepmind recently made an AI that can beat all Atari games [1]. The AI wasn't given "well defined parameters". It was just shown the game, and it figured out, on its own, how to map inputs to actions on the screen, and which actions resulted in progress towards winning the game. Then, the same AI went on to do this over and over again, eventually beating all 57 Atari games.

Yes, you can argue that this is still a limited example. However it is an example that shows that AIs are capable of generalized learning. There's nothing, in principle, that prevents a domain-specific AI from learning and improving at other problem domains. The AI that I'm conceiving of is a supersonic jet. This AI is closer to the Wright Flyer. However, once you have a Wright Flyer, supersonic jets aren't that far away.

> To make an AI that outperforms humans in any task has not been proven to be possible (to my knowledge) not even in theory. An airplane will fly faster, higher and with more cargo then a flock of geese, but a flock of geese reproduce, communicate with each other, digest grass, etc. An airplane will not outperform a flock of geese in any task, just the tasks which the airplane is optimized for.

That's fair, but besides the point. The AI doesn't have to be better than humans at everything that humans can do. The AI just has to beat humans at everything that's economically valuable. When all jobs get eaten by the AI, it's cold comfort to me that the AI is still worse than humans at, say, enjoying a nice cup of tea.

[1]: https://www.technologyreview.com/2020/04/01/974997/deepminds...

Re: GitHub Copilot as open source code laundering?

#450

For years people have warned about hosting the majority of world's open source code in a proprietary platform that belongs to a for profit company. These people were called lunatics, fundamentalists, radicals, conspiracy theorists, and many other names. Well, they were ignored and this is the result. A for profit company built a proprietary system using every code hosted in its platform without respecting the code li…

If we feed the entirety of a library to an AI and have it generate new books, is it an exploitation of people's work? If we read a book and use its instructions to build a bicycle, is it an exploitation of people's work? No, no it's not.

If you read a book and use the instructions to build a bicycle you are learning a new skill and this is obviously not exploitation of people's work.

When you read a book and copy this book partially or entirely to create a new book or create a derivative work using this book without citation it's called plagiarism and copyright infringement. It is not only exploitation, it is against the law.

If you feed an entire library to an AI to generate new books without source citation and copyright agreements it is not only exploitation, it is against the law. We can call this automated plagiarism and copyright infringement, and automated or not, it is against the law. Except if you use public domain books. It wouldn't be illegal but highly unethical considering there are powerful companies with big pockets bending public domain's laws to avoid their assets to be public available (I'm looking at you Disney), but that is another story.

Post reply on HN