Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

181–190 of 274 posts

Re: AI weights are not open “source”

#181
post #123
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Is there an official ruling? Or is it just a Reddit-style over exaggeration?

There is no official ruling... yet. We are very early in this rapid public development. Laws and rulings take years or decades.

They are trained on a lot of text. News sites, comments, books etc. Most books and news sites fall under copyright. Is this fair use? Who knows. Fair use is also an American thing. ChatGPT can be used in the EU, which doesn't have such a broad view of fair use.

If you make a game only out of a lot of copyrighted assets without paying it isn't fair use. Are LLMs different?

What about image generation, which you can prompt the models for specific styles of artists, which works are all copyrighted, but still used for training?

Re: AI weights are not open “source”

#182
post #134
post #64

Earlier quoted context omitted.

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below

> you can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. They don't necessarily do. Think about that. You can take some copyrighted material and transform the information contained in it (for instance a fictional book). You can then write a summary. The summary contains information that was present in the original but it has been transfo…

I like this argument a lot; but again -- how does this play out in the real world? It's pretty easy to refute what will happen in real life. Think, e.g Batman. I could write a very new and original "Batman" comic that doesn't strongly resemble anything -- movie, toy, comic, whatever -- that exists, but would be recognizable to fans.

Once it starts doing well, will DC come after me? You bet.

Re: AI weights are not open “source”

#183
post #66

Earlier quoted context omitted.

Interesting. I wonder what they mean by "Ethical" -- instead of e.g. saying "definitely free and open." I'm willing to bet "stuff they gathered from likely unwitting Adobe users."

It means trained on data from stock photo sites they own, for which all uploaders agreed to terms of service which state that uploaded materials can be fed into AI training.

So yes, exactly what I said. :)

Re: AI weights are not open “source”

#184
post #66

Earlier quoted context omitted.

Interesting. I wonder what they mean by "Ethical" -- instead of e.g. saying "definitely free and open." I'm willing to bet "stuff they gathered from likely unwitting Adobe users."

Adobe also happens to own Adobe stock, so maybe they simply trained on their own corpus. Who am I kidding this is Adobe of course they're fucking over their users

They at least say they did and other copyright free artworks. They are a big company and know that they would get sued, so it should be in their interest to do it this way.

Re: AI weights are not open “source”

#185
post #126

Earlier quoted context omitted.

I think there could be an argument that it's copyrightable but not a derivative work. If I read a few books about a subject as research, and then I write an article about the subject, it's my own copyright. The fact that I did research doesn't make it derivative of those books (correct me if I'm wrong, IANAL). Perhaps a model created from copyrighted material be treated in the same way?

That's because you are human and have rights that a computer program doesn't

A fertile subject for sf stories.

Re: AI weights are not open “source”

#186

Earlier quoted context omitted.

The answer would be if the weights are transformative enough, and the copyright would come from the person who decided what images to include in the training set. The act of choosing to place images in a certain arrangement, such as a collage, can be copyrightable. The same could be said for the "act" of choosing what images to include in a training set and which parameters to use to train the model.

Does that mean if someone copies a phone book but leaves out some numbers and adds some other numbers then it's a creative work?

The legal system is not like a computer program. The line between what is "creative" and what is not concrete, but is instead up to the interpretation of the judge who rules on it.

So your phonebook modifications may or may not be considered "creative" depending on the judge and your ability to convince them. The more your modify it, the more likely you are to convince a judge it is a creative work, though.

Re: AI weights are not open “source”

#187
post #163
post #134

Earlier quoted context omitted.

> you can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. They don't necessarily do. Think about that. You can take some copyrighted material and transform the information contained in it (for instance a fictional book). You can then write a summary. The summary contains information that was present in the original but it has been transfo…

what is a 'copy'? byte accurate, or 'something with general resemblance'? would a badly compressed "copy" image of a copyrighted material still be 'a copy' or would it be some other thing? would low quality image compression be enough to skirt around copyright claims? image formats and viewers just 'reproduce' an impression of original data from derive compressed data. it is also just 'information that's been general…

Badly compressed still counts. I think if the data allows you to reconstruct a recognizable recreation of the original work, you have a good chance of it being considered a derivative copy.

A mono audio version of Star Wars, compressed down to 320x240, filmed from the back of a theater on a VHS camera, converted to Video CD, would under any reasonable interpretation be just a copy of the original.

I assume it starts getting murky when there's some sort of transformation done it it. What if I run motion capture on it, and use that motion capture data to create a cartoon version of Star Paws (my puppies in space epic)? What if I do a scene for scene recreation as the animated cartoon (removing any mentions to copyrighted names -- Luke Skywalker is now Duke Dogwalker, for example)? In this case, there's been no actual data transfer -- all the sprites are hand drawn, backgrounds etc.

What would be an interesting exercise would be to try and create a series of artifacts that each on their own are considered non-derivatives, but can be used together to reconstitute the original. For example, create a compression method that relies heavily on transforms / macroblocks, but strip out any of the actual pixel data from the film. That info might be supplied as palette files which are themselves not really copyrighted data, but together with the compressed transform stream can be used to recreate the original video.

Re: AI weights are not open “source”

#188

The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like…

> The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. Yes. Weights probably aren't copyrightable in the US. See Feist vs. Rural Telephone, in which the Supreme Court ruled that telephone directories are not copyrightable. The copyright clause in the Constitution ("To promote the Progress of Science and useful Arts, by securing for limited Times to…

> Outputs from LLMs, machine generated art, and machine generated music probably are not copyrightable either.

I don't have a strong sense of whether this is reasonable (I see arguments both ways) but I do think it's pretty strongly at odds with how we treat photographs. There are a bunch of photos on my phone where I unquestionably own the copyright, despite putting in much less creativity than I did for some AI images I've generated.

I don't think it's clear how to resolve this, but I do think that if we are going to protect photos and not prompted AI images, the distinction needs to turn on something other than whether "sufficient creativity" was applied to the input of the mechanical system.

Edited to add: It's probably also worth calling out that the question of whether we protect the work produced by a person's use of mechanical system is a separate one from whether we protect the work of others when it is (in various ways, to various degrees, with various likelihoods) reproduced by use of those mechanical systems.

Re: AI weights are not open “source”

#189
post #179

Earlier quoted context omitted.

I'm not a lawyer, but it seems like you stood up a straw man there. >Just because it's represented as a bunch of numbers does not make it non copyrightable. Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted…

The key factor of Feist v Rural is whether there was any original or creative process in the way the facts were arranged. Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights, so it's reasonable to think it might be copyrightable. That is, the numbers are a whole lot more original than the issuance of phone numbers or part numbers.

IANAL, but I’d wonder whether ‘creativity’ is really present in labelling - and indeed, mightn’t it be the last thing you want? I’d argue labelling should be strictly factual and reproducible, and ideally following a logical structure… maybe akin to how addresses of buildings might appear in a phone directory…

(Agree that the skill in knowing how to code and guide the training of a model is probably very different though. It’s not just access to compute time that separates me from OpenAI :) )

Re: AI weights are not open “source”

#190
post #179

Earlier quoted context omitted.

I'm not a lawyer, but it seems like you stood up a straw man there. >Just because it's represented as a bunch of numbers does not make it non copyrightable. Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted…

The key factor of Feist v Rural is whether there was any original or creative process in the way the facts were arranged. Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights, so it's reasonable to think it might be copyrightable. That is, the numbers are a whole lot more original than the issuance of phone numbers or part numbers.

> Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights

Often, labelling is part of large public datasets that are chosen for use for that exact reason, and/or is otherwise not the work of the party claiming copyright in the model.

Post reply on HN