Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

101–110 of 274 posts

Re: AI weights are not open “source”

#101
post #69

Earlier quoted context omitted.

No, that doesn’t follow at all. The argument is that either the training or the expression violated existing cooyrights through the making of unlicensed copies. It’s not based on open source licensing. Although OSS viral licensing may well apply if fair use is not a successful defense.

What copyright is violated by training on public domain data?

If it's public domain, then no copyright is violated. I'm not talking about public-domain data; the G-G-GP specifically mentioned the possible legal interpretation that training on large amounts of publicly visible (but not public domain) data is itself a copyright violation.

Re: AI weights are not open “source”

#102
Just like you can't de-compile a binary without loss of information, "source" means that you can reconstruct it, so the training data should be available as well as the code that was used to train it, and the build script that invoked it.

Re: AI weights are not open “source”

#103

Earlier quoted context omitted.

That's also my understanding, either the weights are copyrightable and then all the models need explicit agreements for any work they include in it because models become derivatives or they are not copyrightable being just machine data (the most likely scenario in my opinion), they can't have it both ways.

"Transformative use". The inputs could be copyrighted and the weights could be copyrighted if creating the weights from the inputs is (legally) regarded as a transformative use. And I think it could reasonably be considered to be transformative - the weights don't look anything like the input data. Disclaimer: IANAL. So far as I know, no court has ruled on whether this qualifies as a transformative use. I take no pos…

Another way to look at it is if a thing reproduced a data subjectively resembling originals and then you used it anyhow, then its non-transformative use, and methods used is just extra details.

Re: AI weights are not open “source”

#104
post #39

The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like…

Furthermore, if weights are copyrightable, wouldn't this make the issue of training data licenses even more urgent? IANAL, but if weights are IP, wouldn't they constitute a "derived work" of the training data?

In a sane legal system a new copyright law would be passed to clarify all of this. In ours, the poor copyright office needs to make things up on the fly.

Their recent decision that implies that anything that AI is used to produce is non-copyrightable is silly, sad, and not sustainable.

Re: AI weights are not open “source”

#105
post #49
post #43

Earlier quoted context omitted.

Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

Copyright is for things that are the result of human creativity. If the weights come from running an algorithm on a training set (that one does not have a copyright to) then how can the weights then be copyrightable? They might be a derivative work, but that just means they infringe copyright, not that they are copyrightable themselves.

Re: AI weights are not open “source”

#106

Earlier quoted context omitted.

That's also my understanding, either the weights are copyrightable and then all the models need explicit agreements for any work they include in it because models become derivatives or they are not copyrightable being just machine data (the most likely scenario in my opinion), they can't have it both ways.

"Transformative use". The inputs could be copyrighted and the weights could be copyrighted if creating the weights from the inputs is (legally) regarded as a transformative use. And I think it could reasonably be considered to be transformative - the weights don't look anything like the input data. Disclaimer: IANAL. So far as I know, no court has ruled on whether this qualifies as a transformative use. I take no pos…

Transformative use doesn't necessarily mean copyrightable.

Google's thumbnails are a purely mathematical transformation on images (no copyright themselves), and yet are considered a transformative use.

I believe that trained models are similarly a purely mathematical transformation of {data}, but is transformative in what that can be used for going forward.

"Can" bearing a lot of weight in that sentence.

It's how the human, with agency, uses the model that may be a derivative or copyright infringing use - not the model itself nor necessarily the output.

The output of a generative AI may be similar enough to an existing work that it is derivative of that work. It is possible to construct a prompt that infringes on an existing work even if that work wasn't part of the training data.

For that case, consider you drew a picture. That picture that you just drew isn't part of any training data. I could presumably look at it and describe it with sufficient detail that something similar enough would be generated... and that may be considered a derivative work. The same test could be applied to me describing it to someone on Fiverr with the same outcome.

If I were to publish that work by the generative AI or Fiverr - who would be infringing on copyright? me? or the black box that may be AI or Fiverr that created a picture based on my prompts?

Re: AI weights are not open “source”

#107
post #64

Earlier quoted context omitted.

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below

This is the first time I ever saw a comment including the text "I am a lawyer". Does that mean the comment technically contains "legal advice"?

As always, a lawyer is not necessarily your lawyer.

Re: AI weights are not open “source”

#108
post #49

Earlier quoted context omitted.

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

It seems very difficult to ensure that a model will never output any of the copyrighted content that it was trained on. I can only think of three ways, but perhaps there are others 1. Evaluate every output from the model to ensure that none of the outputs are copyrighted 2. Evaluate every input to a model to ensure that the inputs are either not copyrighted or properly licensed 3. Change the definition of copyright s…

Ensuring a model never outputs copyrighted content is unimportant and tangential. It's irrelevant. You don't look for a way to make humans output no copyrighted content, you address each time they do case by case.

A model training being rendered fair use doesn't mean any of its output can be used for whatever regardless.

Re: AI weights are not open “source”

#109
If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are.

The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights.

The express intent of copyright is now a sad joke.

Personally - I'm pretty over the entire show. This system is generating an incredible amount of inequality. New and novel content is absolutely NOT getting made, and these laws are creating vicious infights that drain resources from well intentioned companies & individuals and pass them along to complete scam corporations.

We are told stories as children that we cannot retell in our own voices decades later to our own children.

I am firmly ready to burn this copyright system to the fucking ground. It's been 300 years since the Statute of Anne - I'm ready for a different game.

Re: AI weights are not open “source”

#110

The article makes a good point: we should prevent “open-washing” and draw a distinction between well-intentioned restrictive licenses like “Open”RAIL and true open source. However, I worry the name “ethical source” is itself a bit question-begging. While outfits like Bloom may believe in good-faith ethical principles, their definition of ethics isn’t necessarily everyone’s. If restricted models are “ethical”, is rele…

I think the more interesting aspect of all this is that the confusion created by this new business model ( not sure to classify it so business model had to do ) appears to be largely intentional. The subject matter is complicated to begin with experts being niche of a niche of a niche and the assumption that the general public can even understand it ( and whether it can even dumbed down to digestible sound bites ) is…

"Now, courts are not typically stacked with dummies, but again how many are well versed in issues of technology?"

Even if they are well versed in issues of technology that does not mean they'll make what any given one of would consider a good decision, as plenty of people well versed in issues of technology disagree with each other on these issues.

Nothing guarantees that on, on any issue, really, as you can always find people who disagree.. and if they happen to be judges, they get to decide unless another higher judge overrule them.. and that judge has the same problem as the first.

Post reply on HN