Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

91–100 of 274 posts

Re: AI weights are not open “source”

#91
post #49
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

My issue with this take is that machines are not people. We only have lax rules for humans precisely because they are humans, not on the basis that they can learn. Copyrighted works are produced for people and given how human learning works, applying the derivative works rule to humans would be completely impractical and destroy the point of works with copyright. The same cannot be said for AI companies treating everything on the internet as fair use for training AI.

Re: AI weights are not open “source”

#92

Earlier quoted context omitted.

IP itself violates a natural right (Yes the idea of rights is also unnatural and absent from visions such as anarchy)

Yeah I didn't even think that was controversial. I'd always been taught that copyright and patents exist to explicitly restrict what people can do by granting a monopoly to the owners in order to encourage invention and creative work. Edit to add I'm not saying I agree with the justification or am trying to argue for it, only that the point above is commonly raised as the justification, implying that the intrusion on…

[flagged]

Re: AI weights are not open “source”

#93
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

If I collect a set of copyright free data or public domain data would we conclude that the weights are also public domain?

IANAL, I rather think the weight is not copyrightable anyway, and, if I build a model on copyrighted data, I would conclude that the inseparable but reproducible parts of weight retains copyrights, despite the whole weight not having its own.

Re: AI weights are not open “source”

#94
post #52

Earlier quoted context omitted.

Why is it unestablished? Is a document not copyrightable based on its contents? Weights are just a different kind of a document.

Copyright is not for "documents", it is for works that have creativity in them. The legal bar for that level of creativity is low, so low that it is easy to come away thinking that anything that can be cast as a "document" must be copyrightable, but the bar is in fact not zero. In particular, taking other documents and shoving them through a process that generates a lot of other numbers with no human or creative inte…

IANAL, but I suspect that the "novel legal theory" in your last paragraph would fail. It might succeed if you gave GPT a hand-curated list of materials; hoovering up the entire internet is not that.

Re: AI weights are not open “source”

#95

Earlier quoted context omitted.

IP itself violates a natural right (Yes the idea of rights is also unnatural and absent from visions such as anarchy)

Natural rights are a fiction to pretend that someone’s moral code is a privileged aspect of physical reality in a way every competing moral code is not.

Even so, you can ask whether a given moral code is more principled than another (e.g. in the sense of having some algebraic structure), and use that to investigate what might be considered "more natural". For example, one might argue that if a "natural" right exists, then it ought to be symmetric under exchange of humans (or sentient beings or whatever). It's then "more natural" to conclude that you have a right to perform actions that have no interaction or consequences for other humans (e.g. to sing a copyrighted song to yourself in an empty room or downloading a song that you already have on CD but don't feel like ripping yourself) than those that do (e.g. taking food from someone because you'd otherwise starve).

Re: AI weights are not open “source”

#96
post #66

Earlier quoted context omitted.

> Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Maybe the popular and free ones. Adobe has a product in beta that uses "ethical training data" as a selling point.

Interesting. I wonder what they mean by "Ethical" -- instead of e.g. saying "definitely free and open." I'm willing to bet "stuff they gathered from likely unwitting Adobe users."

It means trained on data from stock photo sites they own, for which all uploaders agreed to terms of service which state that uploaded materials can be fed into AI training.

Re: AI weights are not open “source”

#97
post #64

Earlier quoted context omitted.

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below

It's just a race for which test case gets to the supreme court first really...

...and the Supreme Court could rule however it likes. It doesn't matter what anyone else says, or what any law says, what any lawyer or other judge says.

They could be completely biased, could completely ignore everyone and everything else and rule however they want.

I'm almost surprised they still bother to write any kind of "legal reasoning" in their ruling and don't simply focus on what the ruling is rather than why they ruled that way. But I guess such "reasoning" still serves a propaganda purpose and still provides a fig leaf for those who still believe in the quaint absurdity that "we are a nation of laws, not men."

Re: AI weights are not open “source”

#98

Earlier quoted context omitted.

That's also my understanding, either the weights are copyrightable and then all the models need explicit agreements for any work they include in it because models become derivatives or they are not copyrightable being just machine data (the most likely scenario in my opinion), they can't have it both ways.

I think there could be an argument that it's copyrightable but not a derivative work. If I read a few books about a subject as research, and then I write an article about the subject, it's my own copyright. The fact that I did research doesn't make it derivative of those books (correct me if I'm wrong, IANAL). Perhaps a model created from copyrighted material be treated in the same way?

The difference is you are person and have many more rights than a machine.

Re: AI weights are not open “source”

#99
post #69

Earlier quoted context omitted.

No, that doesn’t follow at all. The argument is that either the training or the expression violated existing cooyrights through the making of unlicensed copies. It’s not based on open source licensing. Although OSS viral licensing may well apply if fair use is not a successful defense.

What copyright is violated by training on public domain data?

DMCA works for free data too. It doesn’t matter if your gains are in the form of fiat or crypto or social currency.

Re: AI weights are not open “source”

#100
post #49
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

When the web was young, there was a lot of information considered "public" like criminal record, marriage records, birth certificates, property records, etc. But those were still fairly veiled because of the amount of effort required to see them. Suddenly these were getting blasted all over the internet because now that was an easy thing to do, and everyone had to rethink what "public" meant.

I suspect we're going to see the same kind of rethink about intellectual property in the age of AI.

Post reply on HN