Live data from Hacker News

Ownership of AI-Generated Code Hotly Disputed

spectrum.ieee.org

151–160 of 205 posts

Re: Ownership of AI-Generated Code Hotly Disputed

#151

Earlier quoted context omitted.

People have a strange tendency to enjoy being rewarded for work they've done. I suspect that the folks who invest hundreds of millions of dollars into production costs for a movie, rather enjoy the ability to recoup those costs by restricting access to only those who are willing to pay for the privilege.

but what about people getting paid for the invested money of movies made by their ancestors, dead over 50 years ago (and up to more than 100 years ago, whatever is the current copyright term for movies) do they have a right to enjoy being rewarded for work they didn't do but their ancestors did? if the intention of those laws was to encourage creativity, how come they're stiffing it more than encouraging without any…

If your argument is about the absurdity of current copyright terms, you'll find no disagreement here.

We should have stuck with the original 14 + 14 as laid out in the constitution.

Re: Ownership of AI-Generated Code Hotly Disputed

#152
post #143

Earlier quoted context omitted.

And yet it's a major part of the overall concept of being responsible with our use of AIs. Throwing our hands up in the air and prematurely declaring defeat is not an option long term. It's a non-starter for no other reason than potential copyright infringement means the government becomes involved, and they will stomp on the AI mouse with the force of an elephant - the opinions of amateurs and the anti-copyright mov…

Those companies are not solving the problem you are describing

"No they're not" and "no it's not" (simplified from the actual response) are conversation enders, so I'll follow the lead and let this conversation end.

Have a great new years!

Re: Ownership of AI-Generated Code Hotly Disputed

#153

Earlier quoted context omitted.

ChatGPT output is very clean and precise especially when variables provided in detail. I doubt you can trace it back if you make the prompt very elaborate.

Why can't you trace it back? Remove all the input, fill the model with noise. Prompt all you want it will produce just noise. Train the model on your chosen input text. Suddenly it starts to provide 'clean and precise text'. Absent an argument that we're looking at AGI I see no other possible conclusion than that this is a mechanical transformation of two inputs (yours + the training data) into some output. What happ…

What kind of platform can detect what variables or instructions I added? It can be very difficult to reverse engineer the prompt from the output.

Re: Ownership of AI-Generated Code Hotly Disputed

#154
post #100

Earlier quoted context omitted.

In my opinion using all of the code on GitHub without respecting the licenses was a capital mistake. It should have been opt-in, maybe with some incentive but to just take it all without so much as a by-your-leave is not going to play well in court.

In my opinion it was covered under the GitHub terms of service and is clearly transformative. I am optimistic that the courts will find it so and we can put these debates to rest similar to how we’ve done for web scraping.

The GitHub ToS only applies a blanket license to copy as a fallback if you forget to stick a license file in your code - i.e. it's there so that if you post code to GitHub you can't sue GitHub for doing the thing you told it to do. Furthermore, that only covers stuff intentionally uploaded to GitHub by the author; people mirroring repos hosted elsewhere do not have the legal capacity to license attribution-free, copyleft-free AI training.

Re: Ownership of AI-Generated Code Hotly Disputed

#155
post #118
post #54

Earlier quoted context omitted.

Why wouldn’t it be feasible? (Maybe this depends on what you mean by ‘feasible’.) There’s no technical reason you can’t back-track the weights and make a list of which tokens from which training data were sampled. The list might be long, it could be impractical, but that has little bearing on whether it’s technically possible, right? The problem here happens when the same source is sampled for many tokens in a row be…

Fair use is a bit more context dependent, especially when applied on other form of media. Unstable diffusion is a big example where many seems to feel that no fair use should be allowed regardless of how small of a token is taken by the AI. A small amount also depend on the original work and how it fit into society at large. Having fair use to be the pillar that all AI training stand on is going to take a while.

Yeah good point, TBH I feel like the most practical way forward right now isn’t to wait for Fair Use to clarify AI, it’s probably going to take training data that’s better curated and partitioned into groups that explicitly allow certain kinds of ML use, and models that aren’t leaky across partitions (which sounds like what the problem is with Copilot in this case). This might mean new kinds of code licenses that specify whether it can be used for ML training or not, and companies are going to need to start respecting those licenses rather than assuming they can train on everything reachable from the internet.

Otherwise, if we are going to let companies ignore copyright and licenses because of their pinky promise that their models won’t be recognizable copies of any single input, then we really do need a legal criteria for how much of a single input is fair game, right?

Re: Ownership of AI-Generated Code Hotly Disputed

#156
post #143

Earlier quoted context omitted.

Those companies are not solving the problem you are describing

"No they're not" and "no it's not" (simplified from the actual response) are conversation enders, so I'll follow the lead and let this conversation end. Have a great new years!

Well it helps to be accurate in your statements

Re: Ownership of AI-Generated Code Hotly Disputed

#157

To me AI code generators are the equivalent of crypto tumblers or mixers for digital coins. You can pretend all you want that the output is 'clean' but we all know it came from somewhere else and wasn't actually generated by the software, just endless little snippets that other people made.

This is definitely not true. In theory a generator could reproduce a distribution of data that includes more than the data it was able to sample from. For example, you could create the distribution of all possible faces from a subset of sample faces. That means you could create the faces of people who do not yet exist. Now good luck getting that perfect generator (black swans, bias, etc), but it would be naive to believe that they (well... that there can be no generator that) are just copy input data and mashing them together. At least if you're going to believe that that also isn't happening in the real world, but then what's the difference?

Re: Ownership of AI-Generated Code Hotly Disputed

#158
post #23

Earlier quoted context omitted.

Algorithms aren't copyrightable as far as my understanding goes and this is where it gets murky. The expression is copyrightable, the algorithm isn't (except when it is...) My personal feeling is that we've gone too far down the road of assigning IP to code and we should be rolling it back. I'd hate for the current AI boom to trigger an extension of current IP law.

Leaving aside what I think IP law ought to look like, LLMs literally train on (and reproduce!) the fixed, copyrighted expression. They aren't producing a new work based on abstract knowledge of algorithms that happens to share a particular expression. It's not at all obvious to me why the training set copyright wouldn't nominally flow through the model, even if that seems impossible to actually implement in practice.

Why is everyone fixating on 1% of the outputs. 99% are sufficiently distinct. It replicates code as an unintended side effect, but there are mitigations.

Re: Ownership of AI-Generated Code Hotly Disputed

#159

Code cannot be owned. A creative expression may be copyrighted. Purely functional expressions may not be copyrighted. The output of a trained AI is insufficiently creative to be copyrighted. Only humans can hold a copyright. Now with all that, there really isn’t anything here to get worked up over.

> The output of a trained AI is insufficiently creative to be copyrighted. Only humans can hold a copyright.

That's overgeneralisation. A language model alone, yes, is just derivative. But a language model trained on solving problems with reinforcement learning can surpass humans. For example AlphaGo and AlphaTensor are models that learned from running simulations.

Re: Ownership of AI-Generated Code Hotly Disputed

#160
post #59

Ideas cannot be owned. People can be owned, if you have slavery. Everyone must unlearn the term "Intellectual Property". These laws are anti-property rights. They are Intellectual Slavery laws ( https://breckyunits.com/an-unpopular-phrase.html ). The United States government employs more knowledge workers than all other companies (see NIH, DoD, CDC, NASA, NOAA, NWS, et cetera). Everything they produce is public domai…

People have a strange tendency to enjoy being rewarded for work they've done. I suspect that the folks who invest hundreds of millions of dollars into production costs for a movie, rather enjoy the ability to recoup those costs by restricting access to only those who are willing to pay for the privilege.

> People have a strange tendency to enjoy being rewarded for work they've done.

Did you not read: "The United States government employs more knowledge workers than all other companies (see NIH, DoD, CDC, NASA, NOAA, NWS, et cetera). Everything they produce is public domain, by law. And yet, the people producing these information products still get paid!"

These laws are counterproductive and have terrible second order effects, and are cruel to everyone except a tiny % of the population, just like slavery laws.

Post reply on HN