Earlier quoted context omitted.
No, that doesn’t follow at all. The argument is that either the training or the expression violated existing cooyrights through the making of unlicensed copies. It’s not based on open source licensing. Although OSS viral licensing may well apply if fair use is not a successful defense.
What copyright is violated by training on public domain data?
AI weights are not open “source”
101–110 of 274 posts
Re: AI weights are not open “source”
#102Re: AI weights are not open “source”
#103Earlier quoted context omitted.
That's also my understanding, either the weights are copyrightable and then all the models need explicit agreements for any work they include in it because models become derivatives or they are not copyrightable being just machine data (the most likely scenario in my opinion), they can't have it both ways.
"Transformative use". The inputs could be copyrighted and the weights could be copyrighted if creating the weights from the inputs is (legally) regarded as a transformative use. And I think it could reasonably be considered to be transformative - the weights don't look anything like the input data. Disclaimer: IANAL. So far as I know, no court has ruled on whether this qualifies as a transformative use. I take no pos…
Re: AI weights are not open “source”
#104The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like…
Furthermore, if weights are copyrightable, wouldn't this make the issue of training data licenses even more urgent? IANAL, but if weights are IP, wouldn't they constitute a "derived work" of the training data?
Their recent decision that implies that anything that AI is used to produce is non-copyrightable is silly, sad, and not sustainable.
Re: AI weights are not open “source”
#105Earlier quoted context omitted.
Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…
> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.
Re: AI weights are not open “source”
#106Earlier quoted context omitted.
That's also my understanding, either the weights are copyrightable and then all the models need explicit agreements for any work they include in it because models become derivatives or they are not copyrightable being just machine data (the most likely scenario in my opinion), they can't have it both ways.
"Transformative use". The inputs could be copyrighted and the weights could be copyrighted if creating the weights from the inputs is (legally) regarded as a transformative use. And I think it could reasonably be considered to be transformative - the weights don't look anything like the input data. Disclaimer: IANAL. So far as I know, no court has ruled on whether this qualifies as a transformative use. I take no pos…
Google's thumbnails are a purely mathematical transformation on images (no copyright themselves), and yet are considered a transformative use.
I believe that trained models are similarly a purely mathematical transformation of {data}, but is transformative in what that can be used for going forward.
"Can" bearing a lot of weight in that sentence.
It's how the human, with agency, uses the model that may be a derivative or copyright infringing use - not the model itself nor necessarily the output.
The output of a generative AI may be similar enough to an existing work that it is derivative of that work. It is possible to construct a prompt that infringes on an existing work even if that work wasn't part of the training data.
For that case, consider you drew a picture. That picture that you just drew isn't part of any training data. I could presumably look at it and describe it with sufficient detail that something similar enough would be generated... and that may be considered a derivative work. The same test could be applied to me describing it to someone on Fiverr with the same outcome.
If I were to publish that work by the generative AI or Fiverr - who would be infringing on copyright? me? or the black box that may be AI or Fiverr that created a picture based on my prompts?
Re: AI weights are not open “source”
#107Earlier quoted context omitted.
These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below
This is the first time I ever saw a comment including the text "I am a lawyer". Does that mean the comment technically contains "legal advice"?
Re: AI weights are not open “source”
#108Earlier quoted context omitted.
> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.
It seems very difficult to ensure that a model will never output any of the copyrighted content that it was trained on. I can only think of three ways, but perhaps there are others 1. Evaluate every output from the model to ensure that none of the outputs are copyrighted 2. Evaluate every input to a model to ensure that the inputs are either not copyrighted or properly licensed 3. Change the definition of copyright s…
A model training being rendered fair use doesn't mean any of its output can be used for whatever regardless.
Re: AI weights are not open “source”
#109The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights.
The express intent of copyright is now a sad joke.
Personally - I'm pretty over the entire show. This system is generating an incredible amount of inequality. New and novel content is absolutely NOT getting made, and these laws are creating vicious infights that drain resources from well intentioned companies & individuals and pass them along to complete scam corporations.
We are told stories as children that we cannot retell in our own voices decades later to our own children.
I am firmly ready to burn this copyright system to the fucking ground. It's been 300 years since the Statute of Anne - I'm ready for a different game.
Re: AI weights are not open “source”
#110The article makes a good point: we should prevent “open-washing” and draw a distinction between well-intentioned restrictive licenses like “Open”RAIL and true open source. However, I worry the name “ethical source” is itself a bit question-begging. While outfits like Bloom may believe in good-faith ethical principles, their definition of ethics isn’t necessarily everyone’s. If restricted models are “ethical”, is rele…
I think the more interesting aspect of all this is that the confusion created by this new business model ( not sure to classify it so business model had to do ) appears to be largely intentional. The subject matter is complicated to begin with experts being niche of a niche of a niche and the assumption that the general public can even understand it ( and whether it can even dumbed down to digestible sound bites ) is…
Even if they are well versed in issues of technology that does not mean they'll make what any given one of would consider a good decision, as plenty of people well versed in issues of technology disagree with each other on these issues.
Nothing guarantees that on, on any issue, really, as you can always find people who disagree.. and if they happen to be judges, they get to decide unless another higher judge overrule them.. and that judge has the same problem as the first.