Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

131–140 of 274 posts

Re: AI weights are not open “source”

#131

Earlier quoted context omitted.

Weights are data, not a type of program. A computer program is a set of instructions that may be executed. Weights are values that may be loaded by a program, but are not a program in and of themselves.

But not "raw" data. They are derived from other data and a program. If this was a collaboration where one collaborator did the processing and one sourced the data, they would likely both claim some amount of ownership of the trained weights. At a minimum, it would be an active area of negotiation that the attorneys would take notice of. Source: have negotiated these agreements.

A curated data set is still a data set.

I imagine it is not settled law, but there's a clear argument to be made that regardless of the difficulty in curating the data set, it's still a data set.

Can it be licensed and sold. Yes, surely. Is it proper to pretend an open source license is sufficient protection, probably not.

Re: AI weights are not open “source”

#132

If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pr…

Fully agree that the existing copyright and intellectual property systems are dire need of deep reform. But to get people on board, you can't just propose burning it all down, you need to point to a viable alternative. Say, limiting copyrights to something sane like 15 or 30 years. Or making it easier to invalidate obvious or trivial patents.

Or do you really want to do away with notions of intellectual property altogether? You can make an argument for that, but there would lead to deep economic changes, and you need to anticipate what the end result would look like. You still need some way to encourage the creation of new content.

Pointing out that our copyright/IP system is broken is easy. And you're right, it's totally broken! Coming up with a fix is hard work.

Re: AI weights are not open “source”

#133

Earlier quoted context omitted.

Weights are data, not a type of program. A computer program is a set of instructions that may be executed. Weights are values that may be loaded by a program, but are not a program in and of themselves.

Weights are data in the same way that instruction codes in memory is data.

Values for the variables do not the function make.

Re: AI weights are not open “source”

#134
post #64

Earlier quoted context omitted.

A lot of people are just upset because their local equilibrium has been disrupted and they think that means they lost a natural right. "You wouldn't look at a car and then remember what that looked like when someone asks you to draw another"

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below

> you can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy.

They don't necessarily do. Think about that. You can take some copyrighted material and transform the information contained in it (for instance a fictional book). You can then write a summary. The summary contains information that was present in the original but it has been transformed and hence it's not a copy. The ML model contains information that has been generalized by some degree. So it's just a grey area IMO.

Re: AI weights are not open “source”

#135
post #49

Earlier quoted context omitted.

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

Copyright is for things that are the result of human creativity. If the weights come from running an algorithm on a training set (that one does not have a copyright to) then how can the weights then be copyrightable? They might be a derivative work, but that just means they infringe copyright, not that they are copyrightable themselves.

The answer would be if the weights are transformative enough, and the copyright would come from the person who decided what images to include in the training set.

The act of choosing to place images in a certain arrangement, such as a collage, can be copyrightable. The same could be said for the "act" of choosing what images to include in a training set and which parameters to use to train the model.

Re: AI weights are not open “source”

#136
Agreed, output weighs are target code, and no one would argue the contrary. Companies pretending to publish source code is nothing new.

Stallman defines source code as "the preferred way in which developers modify the program"

I wrote for wikipedia once that

"Stallman's definition thus contemplates JavaScript and HTML's source-target ambivalence, as well as contemplating possible future forms of software production, like visual programming languages, or datasets in Machine Learning."

So the datasets could be a form or source code, but the most appropriate source code would be the code that crawls or downloads the dataset and modifies it.

Clear as water

Re: AI weights are not open “source”

#137

Earlier quoted context omitted.

There is also the whole patent / copyright trolling issue too. The fact that $BIG_CORP can hire armies of lawyers to freeze competitors and beat them to market by filing frivolous lawsuits is yet another example of insanity in the whole system.

It's a problem with legal system (not unique to any specific country, mind you, the problem is global), not patent or copyright system specifically. It grew incredible amounts of complexity so pro se became a sad joke in all but simplest cases, and there's no incentive to fix it - quite the opposite, everyone in the system is all for keeping the status quo, because it generates money.

But there are specific problems with copyright and patent law that could be improved without a global systemic overhaul that may never happen.

We have to take some small wins even in the presence of big problems.

Re: AI weights are not open “source”

#138

If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pr…

Fully agree that the existing copyright and intellectual property systems are dire need of deep reform. But to get people on board, you can't just propose burning it all down, you need to point to a viable alternative. Say, limiting copyrights to something sane like 15 or 30 years. Or making it easier to invalidate obvious or trivial patents. Or do you really want to do away with notions of intellectual property alto…

I don’t think we need to use the legal system to encourage creation of new content! That’s a natural thing people do. In fact there’s a lot of artistic remixing that is illegal or ambiguously legal under the current copyright regime that can be a powerful form of expression.

I really don’t think we need government policy to encourage artists to create art. (At least not of this sort - I am all for art grants.)

Re: AI weights are not open “source”

#139
The lack of freedom to modification makes it not "open" either.

Comparing to traditional software, weights are actually worse than binary. You can't "decompile" the weights into the training source code so there is no way for the community to make useful changes to them.

Re: AI weights are not open “source”

#140

If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pr…

> I am firmly ready to burn this copyright system to the fucking ground.

Same, but the issue is not copyright, which is simply an effort to wield the state to control intellectual property in the same way the state is wielded to control physical property.

The compounding problem arises when property is capital, defined as the means to convert labor into new value. Capitalism is specifically a system in which one can wield control of capital (intellectual or otherwise) to extract profit from labor then trade that profit for more capital. As a result, capital accumulates infinitely, independent of the value produced by the labor which is provided to society.

Artists require capital to convert their labor into value just as any other worker would, so where should that capital come from if not from control of the value they produce? Society must solve this problem or we will not have art to begin with. Only looking at the demand side obfuscates such issues that arise on the supply side, and the only reason we're talking about them now is that digital technology has solved the scarcity problem on the supply side. It has not solved the scarcity problem on the demand side, however.

Finally, art, just like all technological progress, is always the product of entire societies and the history of all mankind that came before it. For this reason, all copyright and patents have no rational basis and are merely bandaids for the ill side effects of controlling capital to extract profit from labor to begin with.

Post reply on HN