Earlier quoted context omitted.
I don't even care about it being illegal. As said above, US can completely ban such tech and China will keep on trucking along. I just want artists affected to be paid, and actually paid. Not "paid" the way spotify artists are. if you're making a billion off of 100 artists' work, you better be making each affected artist a millionaire in royalties.
I'm a bit confused - you draw parallels between artists whose work gets used in a training dataset and musicians who sign contracts with Spotify, but the comparison is very strained. The financial link with Spotify is obvious - the company pays some minuscule amount of money per use to directly use a product of an artist, in part for their own profit. With AI, the situation is far from being that clear-cut. For one,…
Random means there can be copyrighted content in it. It shouldn't be "random". You can't do such things when making your own piece of media, so I don't see why it's okay when it's a bot scraping massive pieces of the internet for use in some for-profit venture. Scraping was already this gray area in the internet and practices like these argue against such techniques.
>Then, when a model is trained on that dataset and that model is used to make other outputs, are those outputs actually derivative with a direct link to some origin?
That's the million dollar question (literally. Or more like 10/100 million). I imagine like existing copyright it depends on how far out you derive. Arguments above that suggest the ability to regurgitate the exact or very close to piece of data would certainly lean towards "yes". In which case it implies they have that piece of data in their database, instead of just being a "reference".
That feels too close hosting off of copyrighted material in my eyes, but I don't have the full picture of how and what these LLMs contain.
>If I generate a landscape using an AI, to whom exactly would I even owe money?
Like any media creator/producer would tell you: depends on thr license. For example, I Google "mountain" and my first result is this :
https://unsplash.com/photos/aerial-photography-of-mountain-r...
License says "free to use under the unsplash license", which is surprisingly lax. But there are two stipulations
Photos cannot be sold without significant modification.
Compiling photos from Unsplash to replicate a similar or competing service.
So we come back to the whole issue in the beginning with the first point. The 2nd point is a much more nebulous one to consider but not to be completely dismissed (maybe an AI service can be argued to generate a competitor).And that's one of the more generous results. Another result came from CNN but sourced from Shuttershock. Here's their license :
https://www.shutterstock.com/license
I'm not going to read nor summarize the license, but I'm sure CNN at the bare minimum needed to ask permission and likely negotiate to use it in their news post.
So yeah, an "ethical" training model will be doing this for every single image they use, if the aren't creating nor taking the pictures themselves.
>Equally to all contributors in a dataset?
Don't know. That would be for the courts to decide. If it's anything like streaming music and movies, you would be compensated proportionally to the amount of "usages" your work gets when generating pieces. Doesn't necessarily have to be equal, but it may surprisingly equalize out since algorithms are at the helm and not brands trying to stand out.
It's certainly be it's own rabbit hole to explore though. If we ever get that far.