Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

171–180 of 274 posts

Re: AI weights are not open “source”

#172

If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pr…

Fully agree that the existing copyright and intellectual property systems are dire need of deep reform. But to get people on board, you can't just propose burning it all down, you need to point to a viable alternative. Say, limiting copyrights to something sane like 15 or 30 years. Or making it easier to invalidate obvious or trivial patents. Or do you really want to do away with notions of intellectual property alto…

> You still need some way to encourage the creation of new content.

Do you? What's the argument for this? Is there some sort of extreme shortage of creative work that the state should find it necessary to encourage it? How about we end copyright, and if there's ever a problem, we offer copyrights for a short period to fluff the commons up again. A copyright anti-holiday, as it were.

Instead we do the opposite: automatically copyright everything anyone produces, and make it very difficult to surrender your copyright (unless Google or Microsoft want it, then if you object you're literally a Luddite caveman who is trying to turn back the clock on modernity because you're old, stupid, and afraid of fire.)

Re: AI weights are not open “source”

#173

I'm disappointed that the article is only making the (somewhat pedantic) distinction between source code and weights. From the quotation marks in the headline I hoped that it would instead be making the distinction between human-readable source code and machine-readable compiled form. For example, IMHO (IANAL) an AI code-completion tool that had been trained on GPL software is (or should be) only be legal to distribu…

This is an interesting point. If you read the OSI open source definition, specifically on source code (quoted below) I'm inclined to treat the training data as part of the source code for the purpose of determining whether to consider any model open source.

  2. Source Code
  The program must include source code, and must allow distribution in source code as well as compiled form. Where some form of a product is not distributed with source code, there must be a well-publicized means of obtaining the source code for no more than a reasonable reproduction cost, preferably downloading via the Internet without charge. The source code must be the preferred form in which a programmer would modify the program. Deliberately obfuscated source code is not allowed. Intermediate forms such as the output of a preprocessor or translator are not allowed.
https://opensource.org/osd/

Re: AI weights are not open “source”

#174

The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like…

Weights are equivalent to compiled object code IMO. All else follows from there.

Compiled object code of a bunch of code you didn't write. I don't know why programmers are so eager to forget that copyright is not at all about what something is, and all about where it came from. It'd be hard to assert that you hold copyright over object code compiled from code you didn't write!

Re: AI weights are not open “source”

#176

Earlier quoted context omitted.

The answer would be if the weights are transformative enough, and the copyright would come from the person who decided what images to include in the training set. The act of choosing to place images in a certain arrangement, such as a collage, can be copyrightable. The same could be said for the "act" of choosing what images to include in a training set and which parameters to use to train the model.

Does that mean if someone copies a phone book but leaves out some numbers and adds some other numbers then it's a creative work?

It would depend on how transformative the work is.

There is in fact a whole art form where people cut out words from different newspapers and books, for example, and re-arrange those words to form new and interesting art.

So there are ways in which such a work would be a creative work, and ways it which it would not, and it would depend on the particular instance and example.

Re: AI weights are not open “source”

#177
post #154

Earlier quoted context omitted.

> The summary contains information that was present in the original but it has been transformed and hence it's not a copy. The summary also contains original thought, something is added to it by a human to make it unique. AI models are primarily deriviative. A better example would be: if I take 1,000 different copyrighted works and put them into a ZIP file, does that resulting file violate copyright?

Let's say you take the harry potter books and create a spreadsheet with each word in it as a column, and the number of times that word appears. Would that violate the copyright? I'd be interested in the rationale if someone thinks it would.

If your table was the number of times a word was followed by a chain of other words, that would be a closer comparison to AI weights. In that case it would be possible with reasonable accuracy to reconstruct passages from the harry potter books (see GitHub Copilot).

The copyright aspect makes more sense when you start thinking of AI training models as lossy compression for the original works. Is a downsampled copy of the new Star Wars movie still protected under copyright?

Just tabulating the word counts would not violate copyright as it is considered facts and figures.

Re: AI weights are not open “source”

#178

Earlier quoted context omitted.

> The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. Yes. Weights probably aren't copyrightable in the US. See Feist vs. Rural Telephone, in which the Supreme Court ruled that telephone directories are not copyrightable. The copyright clause in the Constitution ("To promote the Progress of Science and useful Arts, by securing for limited Times to…

This is mostly right - It depends on what the weights represent and how they were generated so I would not go as far as the initial claim. A collection of numbers is copyrightable if it's the encoded result of a creative process. Just because it's represented as a bunch of numbers does not make it non copyrightable. That's why it says " original works of authorship fixed in any tangible medium of expression, now know…

I would say that is uncertain. Model weights are always going to effectively be a huge collection of statistics about the training corpus. Unless you are envisioning artisanal, hand-crafted, free-range model weights where a person used a non-mathematical method to purposely and creatively choose each one?

Re: AI weights are not open “source”

#179

Earlier quoted context omitted.

This is mostly right - It depends on what the weights represent and how they were generated so I would not go as far as the initial claim. A collection of numbers is copyrightable if it's the encoded result of a creative process. Just because it's represented as a bunch of numbers does not make it non copyrightable. That's why it says " original works of authorship fixed in any tangible medium of expression, now know…

I'm not a lawyer, but it seems like you stood up a straw man there. >Just because it's represented as a bunch of numbers does not make it non copyrightable. Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted…

The key factor of Feist v Rural is whether there was any original or creative process in the way the facts were arranged.

Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights, so it's reasonable to think it might be copyrightable.

That is, the numbers are a whole lot more original than the issuance of phone numbers or part numbers.

Re: AI weights are not open “source”

#180
post #134
post #64

Earlier quoted context omitted.

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below

> you can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. They don't necessarily do. Think about that. You can take some copyrighted material and transform the information contained in it (for instance a fictional book). You can then write a summary. The summary contains information that was present in the original but it has been transfo…

More over, you are clearly not in violation of copyright if you are talking about statistics about the material. In your example, printing out a "there were 7000 instances of the word 'the'" is certainly not a violation. A ML model is just a huge pile of these statistics.

However, saying "the first word of the book is 'The'" would not be a violation, while repeating that for every word in the book, as a whole, would be one.

Post reply on HN