Live data from Hacker News

AI weights are not open “source”

opencoreventures.com

141–150 of 274 posts

Re: AI weights are not open “source”

#141
post #49

Earlier quoted context omitted.

> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.

Copyright is for things that are the result of human creativity. If the weights come from running an algorithm on a training set (that one does not have a copyright to) then how can the weights then be copyrightable? They might be a derivative work, but that just means they infringe copyright, not that they are copyrightable themselves.

Note that the requirements for copyright are not consistent between nations.

The US has the "threshold of originality" as its principle. Under that doctrine, it requires some human (and this has been emphasized many times over the years) originality in order for something to be copyrighted. It's a low bar for how original it needs to be, but it must be human (monkeys taking selfies are not human).

https://en.wikipedia.org/wiki/Threshold_of_originality

In England, the doctrine is "sweat of the brow" instead.

https://en.wikipedia.org/wiki/Sweat_of_the_brow

> Under a "sweat of the brow" doctrine, the creator of a work, even if it is completely unoriginal, is entitled to have that effort and expense protected; no one else may use such a work without permission, but must instead recreate the work by independent research or effort.

The definitive case for this in the US that set the two apart is Feist Publications, Inc., v. Rural Telephone Service Co. ( https://en.wikipedia.org/wiki/Feist_Publications,_Inc.,_v._R.... ) where it was deemed that a telephone directory is not copyrightable in the US as there is no originality in it... but under the sweat of the brow doctrine it would have been.

So the "[c]opyright is for things that are the result of human creativity" gets an "it depends" and it would be curious to see if companies that are firmly in the "models are valuable" camp go to the UK for what I believe would be a more favorable copyright protection.

... However there are other IP laws around trade secrets that may be better for it in the US (I'm not as familiar in that domain - I would be curious to find out).

Re: AI weights are not open “source”

#143

If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pr…

Fully agree that the existing copyright and intellectual property systems are dire need of deep reform. But to get people on board, you can't just propose burning it all down, you need to point to a viable alternative. Say, limiting copyrights to something sane like 15 or 30 years. Or making it easier to invalidate obvious or trivial patents. Or do you really want to do away with notions of intellectual property alto…

> But to get people on board, you can't just propose burning it all down, you need to point to a viable alternative.

This may be true for most people, I don't know. However, I personally am fully on board the "burn it all down" train and have been for a while.

Re: AI weights are not open “source”

#144

If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pr…

Fully agree that the existing copyright and intellectual property systems are dire need of deep reform. But to get people on board, you can't just propose burning it all down, you need to point to a viable alternative. Say, limiting copyrights to something sane like 15 or 30 years. Or making it easier to invalidate obvious or trivial patents. Or do you really want to do away with notions of intellectual property alto…

Pointing out that IP is broken is _not_ easy because most people believe in the contradictory notion of intellectual property, you included, not knowing the legal history of IP, the legal and economic history of the concept of property, and so on. If it were easy, it would be obvious to everybody that 1) IP law is immoral and 2) nothing bad would happen if it's abolished outright.

Here's a free ebook on the subject, written by a patent lawyer no less: https://mises.org/library/against-intellectual-property-0

Re: AI weights are not open “source”

#145

Earlier quoted context omitted.

There is also the whole patent / copyright trolling issue too. The fact that $BIG_CORP can hire armies of lawyers to freeze competitors and beat them to market by filing frivolous lawsuits is yet another example of insanity in the whole system.

It's a problem with legal system (not unique to any specific country, mind you, the problem is global), not patent or copyright system specifically. It grew incredible amounts of complexity so pro se became a sad joke in all but simplest cases, and there's no incentive to fix it - quite the opposite, everyone in the system is all for keeping the status quo, because it generates money.

Personally I don’t think patents do what people believe they do (encourage innovation). It’s a bigger discussion but briefly, the only literal function of a patent is to discourage innovation by legally barring anyone from using a patented idea as part of a new innovation. The idea we have is that the secondary effects of this will be increased profits for inventors and therefore more innovation. But actually there’s loads of secondary effects and often many of them outweigh the effect of increased profit. For every one inventor that gets a patent there might be 100 prevented from using that idea in a different and innovative way.

A classic example is 3D printers. Stratasys spent 15 years selling printers that cost tens of thousands of dollars. It wasn’t until the patent expired that people figured out how to make them for $250. Those cheaper printers are enabling mechanical engineers and designers to accelerate their process and make other new innovations faster. Stratasys had such a powerful patent they never bothered innovating down in price, instead rested on their laurels selling $25k printers to big customers.

So how many inventions were delayed or shelved because the inventors couldn’t afford a $25,000 3D printer, and $250 printers didn’t exist yet? Both Stratasys and IBM held patents related to 3D printing and they had to cross license to go in to production, so how many others would have come up with 3D printing in the 1990’s if they had not been patented? Would first mover advantage in a free market have been enough to stimulate development of 3D printers? Could we have had $2000 3D printers in the early 2000’s (Stratasys sold theirs for $30k) instead of ten years later? How many engineers would have invented new gadgets faster if they had a 3D printer ten years earlier?

Re: AI weights are not open “source”

#146

Earlier quoted context omitted.

That's also my understanding, either the weights are copyrightable and then all the models need explicit agreements for any work they include in it because models become derivatives or they are not copyrightable being just machine data (the most likely scenario in my opinion), they can't have it both ways.

There is also a (IMO less likely, but still conceivable) scenario where weights ARE copyrightable, but represent fair use of the training data on grounds of being "sufficiently transformative".

I consider that one super likely, but then using the model to make competing works with one the artists in their own style is a non-fair use derivative work

Re: AI weights are not open “source”

#147
post #134
post #64

Earlier quoted context omitted.

These are not bad arguments, but I don't think they're conclusive. I am a lawyer, and I could absolutely see this going the other way. "You can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. You can see this by when they reproduce, e.g. the "getty images" deal." *this is not legal advice, dangit commenter person below

> you can't make these machine things without literally feeding this copyrighted information into them, therefore they do contain a copy. They don't necessarily do. Think about that. You can take some copyrighted material and transform the information contained in it (for instance a fictional book). You can then write a summary. The summary contains information that was present in the original but it has been transfo…

This is a great example. Summarizing or paraphrasing copyrighted content, or simply using it as a seed to generate input-output pairs - this kind of data transformation prior to training could solve the issues with copyright. It cleanly separates form from content.

Re: AI weights are not open “source”

#148
post #31

One thing I don't see discussed enough is that, ok let's say the weights are unencumbered, and the source is under an OSI license: the point of open source licenses and free software was to expose the *human understandable* meaning of the final program. That's why distributing binaries isn't allowed even though technically all of the functionality is present in the machine code. AI weights are basically binary blobs.…

>> AI weights are basically binary blobs. We don't know what they mean, there is really no source code for them. No. You can do further training on them. If they are something less than code I don't think it's going to warrant all this talk about licensing. GPL, MIT, or some proprietary should cover it.

You can do further training on them, just like you can patch a binary blob. There are some surgeries you can do to the weights, and there are analyses you can do to poke at them and try to understand them, but ultimately they weren't created from a human understandable spec, and without a ton of reverse engineering work the weights by themselves aren't human understandable: hence the "source" component is missing.

The source code that generated the weights is one step removed from the kind of source code we'd need to interpret a bunch of AI weights. It's really meta-source code

Re: AI weights are not open “source”

#149

If anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pr…

Fully agree that the existing copyright and intellectual property systems are dire need of deep reform. But to get people on board, you can't just propose burning it all down, you need to point to a viable alternative. Say, limiting copyrights to something sane like 15 or 30 years. Or making it easier to invalidate obvious or trivial patents. Or do you really want to do away with notions of intellectual property alto…

>Pointing out that our copyright/IP system is broken is easy. And you're right, it's totally broken! Coming up with a fix is hard work.

The problem I have with these arguments is they ultimately tend to boil down to the devil you know or the devil you don't know.

We keep claiming when something is broken we must provide a "fix" and the assumption is that fix has to be better than the current approach. There's pretty much no way to guarantee this because the systems in place are the only systems with evidence. So, because we have other ideas, we dare not try them because they have to "fix" the problem. The amount of inertia that keeps corruption in motion bothers me and at a fundamental level most of the inertia comes down, ironically, back to property ownership. If we abolish copyright or change it we have to make sure things are fair/equitable. Well sure, that's ideal, but what we have isn't even remotely fair and equitable anymore, so even something broken is likely an improvement.

We have no willingness as a society to try some modifications and be willing to accept failure, then shift to the next modification and iterate around until we get something sane in place. As such, the systems in place remain in place and more and more holes are found to exploit as time progress.

Our systems need to be more adaptable. Founders of the country understood that which is why they made the legal system a legal adaptable system. The question has always been though, what is the threshold? We've played it safe so long that much of the entire system designed to adapt to fix these issues has itself been targeted and gummed up intentionally to prevent that.

Re: AI weights are not open “source”

#150

> The ethical license category applies to licenses that allow commercial use of the component but includes field of endeavor and/or behavioral use restrictions set by the licensor. I don’t love the name, “ethical license” sounds like a description of the license: this license is ethical. Really this sort of license imposes a particular ethical framework on the user. Not to throw shade, though. It is actually hard to…

"Opinionated" is how I think about it.

That might be a good pick, IMO the word has negative connotations elsewhere, but in tech circles seems basically neutral.
Post reply on HN