Live data from Hacker News

Ownership of AI-Generated Code Hotly Disputed

spectrum.ieee.org

121–130 of 205 posts

Re: Ownership of AI-Generated Code Hotly Disputed

#121
post #119

Earlier quoted context omitted.

Sorry but terms of service don't give you a blanket license to re-purpose someone else's copyrighted work at your pleasure. By that same token any hosting provider could make a small change to their terms of service and suddenly all of the data of all of the customers would be theirs. Copyright does not work that way, you need to actively sign away your rights.

> Sorry but terms of service don't give you a blanket license to re-purpose someone else's copyrighted work at your pleasure. That’s correct. It’s also not what happened here. I don’t violate your copyright when I scrape your web page. I don’t violate your copyright when I train a model on a web page that I scraped. I violate your copyright only if I use that tool to produce code that violates your copyright. This is…

First you wrote:

> In my opinion it was covered under the GitHub terms of service and is clearly transformative

Now you change the argument completely and let's just say that even with that change that is not how I understand that copyright works and leave it at that. You're very welcome to your own interpretation. Best of luck.

Re: Ownership of AI-Generated Code Hotly Disputed

#122

Earlier quoted context omitted.

> It's even being seen in image generation, where NatGeo cover images are reproduced in their entirety or where stock photo watermarks are emitted on finished images. Can you cite sources? I’ve heard this claim repeatedly but have yet to see a good example.

https://youtu.be/kqPKNksl9hk?t=80 https://news.ycombinator.com/item?id=33061707 <- Watermarks (there's been a few linked here on HN)

Neither of those are verbatim copies, do you have an example of a verbatim (i.e. exact, non-transformative) replica?

Re: Ownership of AI-Generated Code Hotly Disputed

#123

Earlier quoted context omitted.

In my opinion using all of the code on GitHub without respecting the licenses was a capital mistake. It should have been opt-in, maybe with some incentive but to just take it all without so much as a by-your-leave is not going to play well in court.

It would have been awesome if Copilot had been written so that you tell it what your license is and it emits code trained only on code with compatible licenses. That would have created an incentive for people to publish their own code under more open licenses, instead of facilitating IP theft and thus creating a dis incentive. That GH didn't choose this option is rather telling IMO.

Yes, that would have been much, much better. Better yet if you could trace it back to the authors so that in the case of attribution licenses you could have properly attributed the code.

Re: Ownership of AI-Generated Code Hotly Disputed

#124

To me AI code generators are the equivalent of crypto tumblers or mixers for digital coins. You can pretend all you want that the output is 'clean' but we all know it came from somewhere else and wasn't actually generated by the software, just endless little snippets that other people made.

I don't think that's entirely true. When I've tested codepilot it also knows something about the context in which it's working, so it can use relevant variable names in it's suggestions. To me that's a step beyond rote regurgitation of other code.

Re: Ownership of AI-Generated Code Hotly Disputed

#125

Earlier quoted context omitted.

Agreed. You don't mind if I borrow your car do you? And your house? Let's not quibble about the particulars of me ever giving them back.

I don't want you taking my car because I need to use my car and only one of us can use it at a time. Intellectual property can be endlessly copied without anyone losing anything. This is not a defense of IP theft (and it's pretty reasonable to think that IP law serves an important purpose protecting creatives' ability to get compensated for their work) but it means that they require a fundamentally different approach…

> ownership is a concept in dire need of revision

GP did not limit the scope of their opinion on ownership to IP.

Re: Ownership of AI-Generated Code Hotly Disputed

#126
post #63

by reducing credit, copilot reduces incentive to create and publish free code. biting the hand that feeds it. exact same problem exists with GPT3 and others. big tech slashing and burning, ruthlessly exploiting the least empowered people in the tech economy. neat hack.

How does copilot reduce the incentive to publish free code?

If it writes what would have been open source dependencies based on open source code you no longer have to license that work or your own. Also coders might not want to open source original work for fear they are just going to feed that beast that will ultimately kill them. Sort of a chilling effect if they cannot chose to contribute to that cause or not while also sharing code as open source.

Re: Ownership of AI-Generated Code Hotly Disputed

#127

Earlier quoted context omitted.

I propose that information naturally wants to degrade. Paper decomposes. Bits flip. File formats are replaced and lost. Storage mediums degrade. It's all an extension of the universe trending towards entropy. It actually takes quite a bit of effort to store, then distribute information precisely and broadly. There's a lot of infrastructure, effort, and money involved, and still information degrades and disappears ove…

I cede your larger point that information availability requires upkeep, but I don’t think the tendency for information to degrade is usually attributable to entropy. It’s probably true enough in some cases, but more often than not it’s going to be general interest by the population that determines the accessibility of information—and that crosses boundaries and regulatory efforts like copyright. The more intense the…

> It’s why you can easily find literature written hundreds or thousands of years ago in multiple languages with little effort

Little effort, because someone else put in the effort to keep and disseminate it. And we only know of the surviving material, we have no idea how much was lost.

Take the bible for example. It's a classic example of stories from millennia ago making it to the modern world. But its journey took the concerted effort of millions of scribes, translators, and so forth to make the journey. And it has not done so intact. We can even see how the Bible's been changed by comparing just the last two major revisions of it: NIV and KJ. Language aside, there are some pretty major ideological changes based on interpretation. And thanks to some miraculous archeological finds in the dead sea scrolls, we can see how the bible has drifted even further from its original roots.

That's not information being free, it's information degrading and being re-purposed - often re-created - for specific ideologies and beliefs.

Re: Ownership of AI-Generated Code Hotly Disputed

#128
post #16

“…modify[ing] its AI model so that it traces attribution and gives credit to the original authors of the code, adding the associated copyright notices and license terms in the process…Biderman says is technologically feasible.” Is it really feasible? What does “traces attribution” even mean here? It’s not emitting “code”, it’s emitting individual tokens that each were found throughout the input corpus. The “code” is…

> It’s not emitting “code”, it’s emitting individual tokens ... The “code” is the arrangement of those tokens, but that is determined by the weighting of the whole network This theory of operation is not borne out in reality. It's been clearly displayed that these tools are emitting verbatim copies of existing code (and its comments) in their input. It's even being seen in image generation, where NatGeo cover images…

The quote does not make that assertion…

Re: Ownership of AI-Generated Code Hotly Disputed

#129
post #73
post #40

Earlier quoted context omitted.

In a word scale. Scale matters. How is one locust different to a million? How is hand copying a manuscript (highly controlled in medieval Europe) different to a printing press. These models will and are already having a profound economic impact on the creative sectors they mimic. Whether existing law applies is a specialist question. But they're not remotely comparable from a pragmatic perspective with pieces manuall…

>A closer parallel would be the industrialisation of painting duplication But they directly reproduce the source material. AI art they clearly does not. The luddites seem a better parallel when it comes to scale. Where a machine comes along capable of producing in much higher quantities and in much greater efficiencies. Or perhaps photography? Also a fear of scale. For a long time photographers were not considered ar…

photography is actually a very good example of how this talking point breaks down. there are a great deal of laws and social codes around where you can and cannot film and take photographs. photography is explicitly disallowed in most museums and art galleries. filming in public requires permits and waivers in many places, especially if the filming is done with a commercial interest.

'but how is capturing an image with a camera lens any different from capturing it with my retina? both are simply methods of refracting light and translating that information into a storage medium. the camera is just remembering* like a human remembers and we mustn't let luddites regulate this revolutionary technology or we'll impede the progress of the camera becoming an even more better rememberer'*

the social ramifications of the camera are still being reckoned with to this day and yes, cameras very highly regulated compared to pretty much any other human method of expression or transcription, in fact we are still coming up with new things you cannot film eg the recent crop of revenge porn laws.

Re: Ownership of AI-Generated Code Hotly Disputed

#130
post #100

Earlier quoted context omitted.

In my opinion using all of the code on GitHub without respecting the licenses was a capital mistake. It should have been opt-in, maybe with some incentive but to just take it all without so much as a by-your-leave is not going to play well in court.

In my opinion it was covered under the GitHub terms of service and is clearly transformative. I am optimistic that the courts will find it so and we can put these debates to rest similar to how we’ve done for web scraping.

Apple made that interpretation regarding their app-store. It is also the reason why GPL software is not accepted on the store. Authors would need to give apple a full copyright transfer agreement (or exemptions of the GPL terms), including any dependencies, and that was simply more costly than just banning GPL from the store.
Post reply on HN