Live data from Hacker News

Ownership of AI-Generated Code Hotly Disputed

spectrum.ieee.org

161–170 of 205 posts

Re: Ownership of AI-Generated Code Hotly Disputed

#161

Earlier quoted context omitted.

It seems like a very blurry line for any non-trivial piece of code.

> It seems like a very blurry line Opinion me, this statement could apply to the vast majority of copyright law in the US.

Haha, imagine they make all the noise now and get language models to solve attribution, but as a side effect all human plagiarism that is hiding in plain sight will be revealed as well. Be careful what you wish for.

Re: Ownership of AI-Generated Code Hotly Disputed

#162
post #55

Earlier quoted context omitted.

In the ideal case the next token is determined by the local context (the prefix string) and the entire corpus of trained code. In this case the prefix string has not been seen before and so the generator must do some interpretation/extrapolation to determine the likely continuation. But in some cases, perhaps many cases, the prefix string has been seen before, or is similar enough to what has been seen before, that t…

That “presumably” is exactly what I’m trying to get past. How would it actually work? Once the input corpus has been chewed up and reduced to tokens, is there actually a representation of “the similar string in the corpus” any more? I don’t think people are even clear on the problem statement here. If I fed in three very similar functions from different sources, and I got a fourth, also very similar, output, what is…

There's a clear difference between an output string that matches 10's of tokens from training data and one that is synthesized from the statistical regularities of the training corpus as a whole and so will not match any input data past a few highly informative tokens (ie ignoring structural tokens). The computational dynamic of doing synthesis should be noticeably different from doing verbatim copying. For one, the statistically relevant context for doing synthesis is much larger than doing verbatim copying. The prior 5 tokens may be sufficient to determine that the continuation should be the verbatim copy (eg quake's fast inverse square root). This difference must be represented in the network activation patterns in some form or another. I don't know what this difference is, but I know it's there.

The case of genuine synthesis is the best case for code generators and aren't the cases where people are expecting attribution.

Re: Ownership of AI-Generated Code Hotly Disputed

#163

Earlier quoted context omitted.

"Purely functional expressions may not be copyrighted." Is this more than a opinion? Because you can have whole programs as a long functional expression (not that I am a fan of such a coding style, but it exists).

It's not worded great (code can be owned under current law). But surprisingly enough, there is a core here that is more than opinion, it just needs some elaboration. It's also going to be hotly debated in the future, because right now most of the commercial AI-generators are just kind of ignoring this and at some point I think they're going to make the argument, "this law/interpretation can't hold or else it would be…

Someone demonstrated a Redux backend simulated by GPT-3. Others demonstrated front end design run by GPT-3. Maybe in the future we're going to ask, the model is going to whip up an app, we use it once and throw it away. One time use apps. They are getting comodified.

Re: Ownership of AI-Generated Code Hotly Disputed

#164

Outside of AI models, copy/pasting snippets from the likes of StackOverflow is already on unsteady ground. The threshold to bother with (and win) legal fights is pretty high. AI is catalyzing some kind of slow revolution in what "ownership" is, but there doesn't seem to be any definition that would _always_ satisfy common sense. Even if github yields and adds attribution or filters on license types, there's still a m…

> stolen, commandeered, swiped from unassuming creators

swiped from unassuming creators who put it online for everyone and Google Bot to see? I presume they follow robots.txt when they crawl

Re: Ownership of AI-Generated Code Hotly Disputed

#165
post #33

Outside of AI models, copy/pasting snippets from the likes of StackOverflow is already on unsteady ground. The threshold to bother with (and win) legal fights is pretty high. AI is catalyzing some kind of slow revolution in what "ownership" is, but there doesn't seem to be any definition that would _always_ satisfy common sense. Even if github yields and adds attribution or filters on license types, there's still a m…

>Many artists the world over are rightfully furious about DALL-E and Stable Diffusion I struggle with this one. How are these models different from the typical human creative process of: 1. look at lots of existing art to get inspiration 2. select components from several different styles & add your own flair 3. call the output an "original" painting in your own style Of course I see the other side as well. These mode…

> filtering a bunch of data through a neural network somehow clears the copyright of that original data

There are two kinds of data - the idea and the expression. You can protect expression, but can't stop AI from learning the ideas. This is not "clearing the copyright", it is fair use of the data. Eventually it will even learn how many fingers and heads to draw on a human - the kind of "idea data" you can't copyright.

We should stop mixing together idea and expression. Ideas are not owned by anyone. The real question is what criteria and threshold to use for judging infringement, so AI people can get on to making safe models.

Besides ideas, you can't copyright recipes, APIs and purely functional/obvious code.

Re: Ownership of AI-Generated Code Hotly Disputed

#166
post #79
post #22

Earlier quoted context omitted.

Even that notion is tricky in the context of text AI. It doesn't seem obviously clear to me what the implied stance towards AI training would be for many popular license types. It seems more like the creators of many licenses didn't explicitly consider this use case. And the more meta question is whether the creators' rights should even extend to that realm. Can you specify in a license "This text must not be read by…

Why is it fair use if you train a model on copyrighted material and use its transformative output, despite going against the will of the author? Can we all legally pirate educational books since it's for self training and producing transformative outputs? Can I consume all media (books, movies, music) the same way, and call it fair use? Also for software?

Content that hasn't been legally obtained wouldn't be legal to consume in any way. That's not what I meant, and I think that's a bit of a disingenuous interpretation of my comment.

The debate is about content that is generally legally obtained, but might come with certain restrictions. Where restrictions is a broad term and might also just come in the form of a copyleft license, eg. My main point was that in many situations, eg involving open source licenses, it's really not clear from the terms what the creator's intent regarding AI training was. And the broader question is whether training is fair use, or something else, maybe even a new legal concept that would have to be established. Or, what's the difference between art students going to the museum to be inspired, and Dall-E 'looking' at public domain images?

Re: Ownership of AI-Generated Code Hotly Disputed

#167

Earlier quoted context omitted.

indeed, the first step (IMO) towards revising the notion is to recognize that physical (material, tangible) assets inherently work differently than digital assets. As I understand so far the main reason to seek a revision of the concept of ownership is exactly due to the existence (enabled by internet technology) of digital assets. copy-pasting is HOW computers work. copy-pasting does not do well in society ruled by…

I'm not exactly an IP hawk, from from it. I agree with you that digital "asset" ownership is new, unexplored territory for humans. We've never been able to separate the content from the distribution medium before now, and we're struggling to recreate a model we're familiar with (physical media) by imposing absurdities like DRM. NFTs are also an absurd way of trying to solve the same problem. The issue we wrestle with…

but this essentially restricts the spread of the 'digital boon' (referring the the new possibilities afforded by this new technologies)

I'm trying to point out that this distinction (ownership as distinct from access) leads towards a capture of the digital advantage by people with better leverage.

moreover, I disagree that if I own a book I do not own the contents of it. the mindset that I don't seems too close to saying that I can know things (well understood learned concepts) but still somehow not own them.

This in my view is like a 'hook' which pulls towards the reality that somebody else owns the contents of my own mind, hence that I do not own that part of myself. I hope you see where I'm going with this and why I find your posture troubling.

It's only a few short (conceptual) steps from doing away with individual freedom for the sake of what?

Re: Ownership of AI-Generated Code Hotly Disputed

#168
post #54
post #16

“…modify[ing] its AI model so that it traces attribution and gives credit to the original authors of the code, adding the associated copyright notices and license terms in the process…Biderman says is technologically feasible.” Is it really feasible? What does “traces attribution” even mean here? It’s not emitting “code”, it’s emitting individual tokens that each were found throughout the input corpus. The “code” is…

Why wouldn’t it be feasible? (Maybe this depends on what you mean by ‘feasible’.) There’s no technical reason you can’t back-track the weights and make a list of which tokens from which training data were sampled. The list might be long, it could be impractical, but that has little bearing on whether it’s technically possible, right? The problem here happens when the same source is sampled for many tokens in a row be…

> There’s no technical reason you can’t back-track the weights and make a list of which tokens from which training data were sampled.

The weights map inputs to outputs, not training samples to outputs. They are adjusted during the training phase and then kept fixed (until there's a new iteration of the model; and let's ignore online learning for simplicity's sake now, the major models we're talking about here don't use it afaik). From the perspective of a user who inputs prompts and receives answers, the weights are constant. We could create some score indicating how much each training sample contributed to the weights overall, but then that score would be the same for all prompts.

It's really not clear to me what you mean by 'sampling from training data' for a specific prompt. The best I could imagine would be creating another metric measuring the distance between the prompt and all training samples, and then somehow combining that with the first score to get an overall contribution metric for this prompt. But that would be a) quite a bit of guesswork with a lot of modelling freedom and b) quite different from what you described.

Re: Ownership of AI-Generated Code Hotly Disputed

#169

Earlier quoted context omitted.

Why can't you trace it back? Remove all the input, fill the model with noise. Prompt all you want it will produce just noise. Train the model on your chosen input text. Suddenly it starts to provide 'clean and precise text'. Absent an argument that we're looking at AGI I see no other possible conclusion than that this is a mechanical transformation of two inputs (yours + the training data) into some output. What happ…

What kind of platform can detect what variables or instructions I added? It can be very difficult to reverse engineer the prompt from the output.

That's not the problem though. The onus would be on you to prove that your machine generated the output and that it wasn't based on the training data but just on your input. Otherwise the generated work isn't yours.

Re: Ownership of AI-Generated Code Hotly Disputed

#170
post #106
post #71

Earlier quoted context omitted.

The token “if” appears in a heck of a lot of inputs, and a heck of a lot of outputs. What you’re describing seems like basically building an inverted index of the input corpus and then doing a search for the output. If the answer comes back with a high relevance then you consider it “traced”. I wonder what the size of that inverted index would be compared to the size of the generative model.

That sounds exactly right to me, or at least this is one specific way to implement a solution to the problem. Good question on size. Speculating… I would guess the size of the index relates not to the size of the model, but to the size of the training data. If you built an extremely naive and straightforward uncompressed index like this, you could imagine building the complete list of known tokens, and for each one t…

I think you'd find that most code matches (some) other existing code.

Some code, like the famous fast inverse square root, is widely shared. Even trivial code is often just copy-pasted from a popular SO answer.

Other code is driven by something like convergent evolution. It ends up similar to other code because of the limits of well known algorithms, language syntax, APIs, boilerplate, common coding styles, etc.

In other words, I doubt most human programmers would pass a "copyright check" against all existing open source code.

Post reply on HN