Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

241–250 of 274 posts

Re: AI weights are not open “source”

#241
post #235

Earlier quoted context omitted.

> Outputs from LLMs, machine generated art, and machine generated music probably are not copyrightable either. I don't have a strong sense of whether this is reasonable (I see arguments both ways) but I do think it's pretty strongly at odds with how we treat photographs. There are a bunch of photos on my phone where I unquestionably own the copyright, despite putting in much less creativity than I did for some AI ima…

Your prompt for the AI image generation is copyrightable. The output is not. The photo you take involved choices of composition and timing and equipment choice. Just because you don't feel you put in a lot of consideration does not mean at a fundamental level that you still put in creative choices that give the resulting product copyright protection. But if you took that photo and put it into software which made a de…

> The photo you take involved choices of composition and timing and equipment choice.

And the AI image involved choices of prompt and model, and subsequent selection from among several generated images.

I recognize that what you said here:

> Your prompt for the AI image generation is copyrightable.

> The output is not.

... probably represents the state of the law at the moment (with meaningful amounts of uncertainty), but I don't think there's a principled difference based on the amount or nature of creativity involved. IMO the equivalent would be "you own the specification of (position, equipment, relevant world state) but not the photo" which obviously doesn't do anything we want for photography. And I guess that's a part of my point. We should pick the policy we want to make sure we capture the incentives we want. Maybe it is best that AI assisted art (past some point?) not be copyrightable. But I don't think basing the distinction on the amount or nature or... propagation (I guess?) of creativity makes any sense in distinguishing flippant and bullshit photographs (at least a third of my photos, although I would hesitate to apply the labels to any particular photo by someone else) from prompt-driven generative works.

Re: AI weights are not open “source”

#242
post #226
post #223

Earlier quoted context omitted.

That sounds like wishful thinking, individual training items have significant impact on the result. Anyway, suppose you’re building an AI to walk, there’s nothing creative about selecting 9.8m/s/s for gravity that’s simply the ideal value to achieve a desired goal. Labeling an elephant as “Elephant” rather than “coat hanger” is similarly a functional choice. Just because a person is holding a camera and taking a phot…

> Anyway, suppose you’re building an AI to walk, there’s nothing creative about selecting 9.8m/s/s for gravity that’s simply the ideal value to achieve a desired goal. Suppose you're not building a strawman, but instead building an AI to be an LLM. The exact sequence of what you choose to do for instruction tuning, and the metrics and labels that you choose, the prompt/response pairs you write, and the loss functions…

> not simple mechanical steps and are the result of creative choice

Creative choices requires intentional control over the output across a meaningfully different range of viable possibilities. A brick layer has a huge range of viable options in the specific brick and its alignment in a wall but none of those choices are artistically meaningful.

The coefficients are also not in any meaningful sense chosen based on instruction tuning. It’s no more under direct control than the specific arrangements of atoms in the brick wall and is instead the output of a purely mechanical process.

> We are nowhere near a point where they are an uncreative, mechanical recipe to follow.

Thus: We are nowhere near the point where the output is under creative control rather than being the result of a poorly understood mechanical recipe.

Re: AI weights are not open “source”

#243

Earlier quoted context omitted.

> Outputs from LLMs, machine generated art, and machine generated music probably are not copyrightable either. Let me put a straw man, and try to find a middle point, when the copyright argument stops being applicable: 1. A painting was done by an artist. 2. On a computer. 3. With a help from an image processor software. 4. Using some advanced filters, like super-resolution, that utilize computer vision techniques. L…

The distinction, defining your straw man, is simply that the image itself is generated by the “commissioned artist” that is the AI. Even non-generative-AI inside Photoshop only mutates images. Generative AI is the source of images.

I'm not sure I understand the "source" term here. All the AI images I've seen so far were generated by humans using software tools like neural networks.

Re: AI weights are not open “source”

#244

Earlier quoted context omitted.

This is mostly right - It depends on what the weights represent and how they were generated so I would not go as far as the initial claim. A collection of numbers is copyrightable if it's the encoded result of a creative process. Just because it's represented as a bunch of numbers does not make it non copyrightable. That's why it says " original works of authorship fixed in any tangible medium of expression, now know…

I'm not a lawyer, but it seems like you stood up a straw man there. >Just because it's represented as a bunch of numbers does not make it non copyrightable. Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted…

> Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted work.

There are random number books and I know one of them has a copyright registration [1].

[1]: https://publicrecords.copyright.gov/detailed-record/7060844

Re: AI weights are not open “source”

#245
post #235

Earlier quoted context omitted.

Your prompt for the AI image generation is copyrightable. The output is not. The photo you take involved choices of composition and timing and equipment choice. Just because you don't feel you put in a lot of consideration does not mean at a fundamental level that you still put in creative choices that give the resulting product copyright protection. But if you took that photo and put it into software which made a de…

If the prompt is copyrightable, then why wouldn't that copyright flow through to the output? I can't legally pirate Windows just because the source code was run through a compiler. Even though the compiler itself adds no additional creativity, the underlying source code is still a creative work[0], so pirating the binaries still infringes a copyright. Just one that's in a slightly different place than what we're norm…

In general if you pay someone to paint a picture they own the copyright and it needs to be assigned back to you even if you give them quite specific instructions. The instructions lacked sufficient control over the outcome to give some form of dual copyright.

Presumably that general rule would also prevent your instructions to DALLE from giving you copyright ownership of the output either. The AI isn’t getting ownership, so it’s either in the public domain or a derivative work from the artists creating the training data.

Re: AI weights are not open “source”

#246
post #134
post #64

Earlier quoted context omitted.

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below

> you can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. They don't necessarily do. Think about that. You can take some copyrighted material and transform the information contained in it (for instance a fictional book). You can then write a summary. The summary contains information that was present in the original but it has been transfo…

I'm not saying that "they do or don't objectively" because that doesn't matter as much as people think it does. I'm thinking of what a "jury" COULD decide. I think average joe on a jury is very likely to see that process as "feeding them in."

Re: AI weights are not open “source”

#247
post #198
post #182

Earlier quoted context omitted.

I like this argument a lot; but again -- how does this play out in the real world? It's pretty easy to refute what will happen in real life. Think, e.g Batman. I could write a very new and original "Batman" comic that doesn't strongly resemble anything -- movie, toy, comic, whatever -- that exists, but would be recognizable to fans. Once it starts doing well, will DC come after me? You bet.

These models can definitely be used to intentionally store and recall content that is copyrighted in a way that's not subject to fair use. (eg: trivially, I can very easily train a large model that has a small subnetwork which encodes a compressed or even lossless copy of a picture, and if I were to intentionally train a model is that way then this would be no less a copyright violation than distributing a JPEG of th…

Honestly, your second to last sentence is literally the kind of thing I hate hearing most from non-lawyers; the whole "if the legislature were just smarter" thing is just a weird pie-in-the-sky concept that is more-or-less like saying "the world would be better if CEOs were less greedy."

Like, yes, but it's not very likely to happen and it's not a particularly horrible thing if it doesn't; the law is slow and little-c conservative and you're just expecting it to be something it MOST often just ain't.

Re: AI weights are not open “source”

#248
post #242
post #226

Earlier quoted context omitted.

> Anyway, suppose you’re building an AI to walk, there’s nothing creative about selecting 9.8m/s/s for gravity that’s simply the ideal value to achieve a desired goal. Suppose you're not building a strawman, but instead building an AI to be an LLM. The exact sequence of what you choose to do for instruction tuning, and the metrics and labels that you choose, the prompt/response pairs you write, and the loss functions…

> not simple mechanical steps and are the result of creative choice Creative choices requires intentional control over the output across a meaningfully different range of viable possibilities. A brick layer has a huge range of viable options in the specific brick and its alignment in a wall but none of those choices are artistically meaningful. The coefficients are also not in any meaningful sense chosen based on ins…

I don't really agree.

Here's what the supremes said in Feist V. Rural:

> Factual compilations, on the other hand, may possess the requisite originality. The compilation author typically chooses which facts to include, in what order to place them, and how to arrange the collected data so that they may be used effectively by readers. These choices as to selection and arrangement, so long as they are made independently by the compiler and entail a minimal degree of creativity, are sufficiently original that Congress may protect such compilations through the copyright laws. Nimmer ss 2.11[D], 3.03; Denicola 523, n. 38. Thus, even a directory that contains absolutely no protectible written expression, only facts, meets the constitutional minimum for copyright protection if it features an original selection or arrangement.

Alphabetical order wasn't quite enough. But people directing the work that produces the coefficients are doing considerably more creative work than that.

> Thus: We are nowhere near the point where the output is under creative control rather than being the result of a poorly understood mechanical recipe.

No one requires complete creative control of the output. I can spatter paint and have relatively poor control of what's happening, but I am certainly generating a copyrightable work when I engage in creative choices as part of this.

Re: AI weights are not open “source”

#249
post #209
post #202

Earlier quoted context omitted.

Some thought experiments: What happens if we train a neural network on a single, copyrighted work? Say it has one input node (or even zero, if you like), and regardless of this input, its output is always exactly the copyrighted work it was trained on. What do its weights represent? Clearly, its weights represent a direct encoding of the original work. Those weights are copyrightable, but not by the person who traine…

In fact, the "factoring out" process shouldn't even be that hard: find the input vector that forces the ANN to output the copyrighted work verbatim. There should be some simple method of "baking in" the first step of the feedforward algorithm, applying that vector to the first layer of weights, and then considering the input layer as the first hidden layer of a network with 0 input nodes. It is now equivalent to a ne…

It’s also politically hard, in that the organisations best positioned to build these attribution tools have large incentives not to do so.

Re: AI weights are not open “source”

#250
post #245

Earlier quoted context omitted.

If the prompt is copyrightable, then why wouldn't that copyright flow through to the output? I can't legally pirate Windows just because the source code was run through a compiler. Even though the compiler itself adds no additional creativity, the underlying source code is still a creative work[0], so pirating the binaries still infringes a copyright. Just one that's in a slightly different place than what we're norm…

In general if you pay someone to paint a picture they own the copyright and it needs to be assigned back to you even if you give them quite specific instructions. The instructions lacked sufficient control over the outcome to give some form of dual copyright. Presumably that general rule would also prevent your instructions to DALLE from giving you copyright ownership of the output either. The AI isn’t getting owners…

Right. And furthermore, I don't actually think prompts alone are copyrightable in most cases - I just wanted to propose an argumentum ad absurdum. Certainly, you can't argue creativity when you're also keyword-stuffing your prompts to call up various feature sets that the model just so happens to associate with them.
Post reply on HN