Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

61–70 of 274 posts

Re: AI weights are not open “source”

#61
post #15

Model weights are not source code, but data. Arguably because of how they are generated, they are not even copyrightable at all.

Corporate data is of course protect-able. Otherwise why don't you just open up all your databases so anyone can access them?

Re: AI weights are not open “source”

#63
post #45

Earlier quoted context omitted.

No, a document's contents aren't inherently copyrightable. They have to be a creative work or a method of production, and part of that basically means it has to be human generated content (as opposed to computer or animal generated). AI weights might be considered a method of production, but that isn't clear yet.

[flagged]

> Was there no work put into their creation by someone?

This is the “sweat of the brow” theory of copyrightability, which courts have rejected (for good reason based on the statute.)

“Someone did work to enable this thing to exist” is not sufficient to make a thing copyright-protected.

> There's no fundamental difference between an image, code, or weights.

And neither images, code, nor weights that are mechanically produced with no creative input by a particular author are subject to copyright in their own right (depending on their relation to the source material on which the mechanical process rests, they may be covered by the copyright on the source material.)

The best argument for weights being copyrightable (and it probably applies better to some models than others) is that the assembly of source material is a creative work subject to a compilers copyright, and that the model weights themselves are just a mechanical translation of that compilation subject to its copyright.

Re: AI weights are not open “source”

#64
post #49

Earlier quoted context omitted.

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

A lot of people are just upset because their local equilibrium has been disrupted and they think that means they lost a natural right. "You wouldn't look at a car and then remember what that looked like when someone asks you to draw another"

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal."

*this is not legal advice, dangit commenter person below

Re: AI weights are not open “source”

#65
post #49

Earlier quoted context omitted.

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

A lot of people are just upset because their local equilibrium has been disrupted and they think that means they lost a natural right. "You wouldn't look at a car and then remember what that looked like when someone asks you to draw another"

IP itself violates a natural right

(Yes the idea of rights is also unnatural and absent from visions such as anarchy)

Re: AI weights are not open “source”

#66
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

> Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Maybe the popular and free ones. Adobe has a product in beta that uses "ethical training data" as a selling point.

Interesting. I wonder what they mean by "Ethical" -- instead of e.g. saying "definitely free and open." I'm willing to bet "stuff they gathered from likely unwitting Adobe users."

Re: AI weights are not open “source”

#67
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

If I collect a set of copyright free data or public domain data would we conclude that the weights are also public domain?

That seems fair. I was under the impression that there weren't too many out there like this.

Re: AI weights are not open “source”

#68

Earlier quoted context omitted.

A lot of people are just upset because their local equilibrium has been disrupted and they think that means they lost a natural right. "You wouldn't look at a car and then remember what that looked like when someone asks you to draw another"

IP itself violates a natural right (Yes the idea of rights is also unnatural and absent from visions such as anarchy)

Natural rights are a fiction to pretend that someone’s moral code is a privileged aspect of physical reality in a way every competing moral code is not.

Re: AI weights are not open “source”

#69
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

If I collect a set of copyright free data or public domain data would we conclude that the weights are also public domain?

No, that doesn’t follow at all. The argument is that either the training or the expression violated existing cooyrights through the making of unlicensed copies. It’s not based on open source licensing. Although OSS viral licensing may well apply if fair use is not a successful defense.

Re: AI weights are not open “source”

#70
post #49
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

It seems very difficult to ensure that a model will never output any of the copyrighted content that it was trained on. I can only think of three ways, but perhaps there are others

1. Evaluate every output from the model to ensure that none of the outputs are copyrighted

2. Evaluate every input to a model to ensure that the inputs are either not copyrighted or properly licensed

3. Change the definition of copyright so that ML models can do whatever they want

Nobody is doing #1, because that makes the business models not work. Established brands (like Adobe) are doing #2. I get the feeling that there are a lot of ML startups that are hoping that #3 will happen, but it seems unlikely

Post reply on HN