Live data from Hacker News

Database of artists used to train Midjourney AI garners criticism

artnews.com

111–120 of 144 posts

Re: Database of artists used to train Midjourney AI garners criticism

#111
post #62

Earlier quoted context omitted.

I don't even care about it being illegal. As said above, US can completely ban such tech and China will keep on trucking along. I just want artists affected to be paid, and actually paid. Not "paid" the way spotify artists are. if you're making a billion off of 100 artists' work, you better be making each affected artist a millionaire in royalties.

I'm a bit confused - you draw parallels between artists whose work gets used in a training dataset and musicians who sign contracts with Spotify, but the comparison is very strained. The financial link with Spotify is obvious - the company pays some minuscule amount of money per use to directly use a product of an artist, in part for their own profit. With AI, the situation is far from being that clear-cut. For one,…

>For one, the majority of datasets aren't comprised of capital-A Art with specific authors and attributions - a ton of it is just random information from around the internet, or just downright junk data.

Random means there can be copyrighted content in it. It shouldn't be "random". You can't do such things when making your own piece of media, so I don't see why it's okay when it's a bot scraping massive pieces of the internet for use in some for-profit venture. Scraping was already this gray area in the internet and practices like these argue against such techniques.

>Then, when a model is trained on that dataset and that model is used to make other outputs, are those outputs actually derivative with a direct link to some origin?

That's the million dollar question (literally. Or more like 10/100 million). I imagine like existing copyright it depends on how far out you derive. Arguments above that suggest the ability to regurgitate the exact or very close to piece of data would certainly lean towards "yes". In which case it implies they have that piece of data in their database, instead of just being a "reference".

That feels too close hosting off of copyrighted material in my eyes, but I don't have the full picture of how and what these LLMs contain.

>If I generate a landscape using an AI, to whom exactly would I even owe money?

Like any media creator/producer would tell you: depends on thr license. For example, I Google "mountain" and my first result is this :

https://unsplash.com/photos/aerial-photography-of-mountain-r...

License says "free to use under the unsplash license", which is surprisingly lax. But there are two stipulations

    Photos cannot be sold without significant modification.
    Compiling photos from Unsplash to replicate a similar or competing service. 
So we come back to the whole issue in the beginning with the first point. The 2nd point is a much more nebulous one to consider but not to be completely dismissed (maybe an AI service can be argued to generate a competitor).

And that's one of the more generous results. Another result came from CNN but sourced from Shuttershock. Here's their license :

https://www.shutterstock.com/license

I'm not going to read nor summarize the license, but I'm sure CNN at the bare minimum needed to ask permission and likely negotiate to use it in their news post.

So yeah, an "ethical" training model will be doing this for every single image they use, if the aren't creating nor taking the pictures themselves.

>Equally to all contributors in a dataset?

Don't know. That would be for the courts to decide. If it's anything like streaming music and movies, you would be compensated proportionally to the amount of "usages" your work gets when generating pieces. Doesn't necessarily have to be equal, but it may surprisingly equalize out since algorithms are at the helm and not brands trying to stand out.

It's certainly be it's own rabbit hole to explore though. If we ever get that far.

Re: Database of artists used to train Midjourney AI garners criticism

#112
post #46

Earlier quoted context omitted.

I used to agree with this, but there's been research coming out of Google that has altered my opinion. Specifically, Google's gotten rather good at making AI spit out unaltered training set data[0]. This is only possible if the AI is remembering large portions of the original trained-on works, which would make the weights infringing. [0] In the most egregious case, they found that just asking ChatGPT to repeat a word…

I'm kind of confused - you've claimed that Google's AI was broken, but cite an anecdote over ChatGPT? Regardless, even if these cases did happen often enough, it's erroneous to assume that this is something universal (i.e. can be applied to all generative AI models) or intentional. Said models are vastly smaller than the sizes of their training datasets, so it's more or less impossible for all the data to be stored v…

Google does research on both their own and other people's models all the time. Here's the blog post and paper if you want to know more: https://not-just-memorization.github.io/extracting-training-...

You appear to be refuting a slightly different point, though. When OpenAI was making Dall-E 2, they found that duplicates in the training set would incentivize memorization of specific images, as you said. This is more like finding a secret cheat code that would make the model regurgitate everything it had remembered, regardless of how "incentivized" it is to do so or even if it had been aligned to not do that.

My personal argument is this: the primary metric that the training process attempts to minimize is perplexity. This is how good the model is at guessing the training set data. Base models are specifically being designed to compress huge amounts of text and we just so happen to accidentally get a decent word calculator out of it. The alignment fine-tuning that happens later adjusts the model to prefer answering questions, but the underlying memorized data is still there.

>The way these algorithms work is public knowledge, there's not really any black boxes in the hands of OpenAI that would be relevant here.

Nope. Modern GPT is entirely a black box. OpenAI stopped publishing model weights the moment Elon Musk stopped writing the checks. Hell, they don't even publish the model architecture anymore. How GPT-4 works, even on a basic "this is how many transformer layers and attention heads we're using" basis, is a trade secret.

Even if you have model weights, the actual meaning of the learned model parameters has never been known; there is active research on figuring them out. One particular problem is polysemanticity. If you look inside a particular hidden layer, you won't see a single "dog" or "cat" neuron in its hidden layers. You'll have a 512-dimension concept bouillabaisse with "dog", "cat", "bird", "guinea pig", "kangaroo", and so on all floating around whatever shape made sense at training time (even if it implies absurdities like "desk is the opposite of loin cloth"). To untangle this, you have to train another AI to pick out monosemantic clusters of neurons that can then be inspected, and that requires extreme amounts of GPU resources.

Re: Database of artists used to train Midjourney AI garners criticism

#113
post #107
post #67

Earlier quoted context omitted.

"I claim this art was made by this person" "Who?" "OK . did you work on this?" "Where are related work products? Are there any? What about invoices? Simultaneous employment?" The reality is that most legal things are determined by _convincing people of a truth_. Perhaps you can set up a whole scheme to "launder" AI art and attach names to them. And all the papertrail you generate doing this will show up in discovery…

The way that AI will be laundered into art is by including it into things like Photoshop. There'll still be a human touch just with "smart brushes" and "smart auto fill" that paints 90% of what you want. Art will then take less skill to produce, and be produced faster for lower prices. An 80% price reduction on art (because artists can now produce it 5 times faster thanks to AI) is 80% as good as getting it for free.

Art will take more skill to produce, not less, at faster speed by select artists. GenAI will become another tool that artists and clients alike must understand and use effectively within unspoken guild rules, that is, if it stays.

Re: Database of artists used to train Midjourney AI garners criticism

#114

Earlier quoted context omitted.

> for one, this requirement is a complete departure from how copyright systems work now. A complete departure? Here is the current form used to register an artistic visual work for copyright. Its more elaborate than you might think. https://www.copyright.gov/forms/formva.pdf Registration is not a rubber-stamp, it is increasingly refused because of indicia of AI tooling. Why would adding some questions on provenance a…

Nothing in the form seems out of the ordinary to me. It is a lot of fields, but ultimately the main goal is establishing ownership, not discerning the specific methodology in which a person made the work. It's a departure in that the current system is results-based, where you register a final product, while the proposed system also must take into consideration every intricacy of creating the work. > it is increasingl…

What part of generative AI seems ordinary to you? the rest follows from there my friend.

Re: Database of artists used to train Midjourney AI garners criticism

#115
post #86
post #51

Earlier quoted context omitted.

It is for the sake of copyright, if you want society to protect your work, provide evidence for your creative work. It seems rather simple to me. Keep in mind that in the not so far future, producing art will be as cheap as consuming it, this means that the original benefits society got in return for copyright no longer applies, so why should they protect it?

> It is for the sake of copyright, if you want society to protect your work, provide evidence for your creative work. I'm not sure if it's that simple - for one, this requirement is a complete departure from how copyright systems work now. Providing complete history logs isn't normal practice, and expanding law to necessitate it isn't common sense. > Keep in mind that in the not so far future, producing art will be a…

> How long will it take until some advanced multimodal algorithm can make a full game that can measure up to ones that are released today?

We can disagree on how long it will take us to get there, but if you use AI generated content, that is not product by copyright, your game as whole, sure, as long as it is not the result of a simple prompt, you're protected as usual.

Keep in mind that already, in many games, there is a mix of protected and unprotected content, for reasons of trademark, copyright, and licensing.

Re: Database of artists used to train Midjourney AI garners criticism

#116

Earlier quoted context omitted.

If it's physical media, you have the physical media. If it's digital media, the software can keep an encrypted record at the brushstroke level that can be played back to produce a bit-perfect reproduction. Maybe even write it to a public ledger.

All of these things have loopholes. For physical media, depending on the quality of the output, one could pay a sufficiently skilled person to reproduce an AI output on physical media in a fraction of the time it'd take to come up with and draw for real. For digital media - ignoring how overbearing this whole system could be, what prevents someone from taking all that data and making an algorithm that outputs brush s…

> All of these things have loopholes. For physical media, depending on the quality of the output, one could pay a sufficiently skilled person to reproduce an AI output on physical media in a fraction of the time it'd take to come up with and draw for real.

Um, that's a real work, you know? In what way does this differ from people who take a photograph and then, for example, creating an oil painting?

Now, there are some weirdnesses because of the copyright of the source photograph, but the oil painting would be your own work.

Yeah, you might get called into court to demonstrate that you can produce the work. But so did Michael Jackson.

Re: Database of artists used to train Midjourney AI garners criticism

#117
post #3

One positive aspect of the status quo in the United States is that AI-generated images are not currently eligible for copyright. I think this is a great direction to go in, I highly doubt Wizards of the Coast or whoever is going to want their premium products to lose copyright protections, so they'll need to keep paying artists. I'd love for us to lean into this -- you can make all the AI art you want, but it automat…

> AI-generated images are not currently eligible for copyright.

There is another bright side of it. This images can be used for AI training without copyright violation. Does this apply to texts as well?

Re: Database of artists used to train Midjourney AI garners criticism

#118
post #6

Earlier quoted context omitted.

Completely unenforceable. How can you even tell if an image was made by AI? What if AI created an outline that was worked on by a human artist (or vice versa)? Who would the burden of proof be on? Steam has a “no AI art” policy, and it’s rapidly turning into a “no obvious AI art policy”. How could they tell?

Furthermore, what is the threshold for something to still legally be considered AI art once an artist's hand has modified it? What if they change the brightness? Fix a hand? completely replace a character in a scene? Illustrate most of the scene themselves but add an AI figure or background? Use AI to sharpen a hand drawn image?

It's a gray area. But if someone generates AI images at scale without human in the loop it's likely not copyrightable.

Re: Database of artists used to train Midjourney AI garners criticism

#119
post #6

Earlier quoted context omitted.

Completely unenforceable. How can you even tell if an image was made by AI? What if AI created an outline that was worked on by a human artist (or vice versa)? Who would the burden of proof be on? Steam has a “no AI art” policy, and it’s rapidly turning into a “no obvious AI art policy”. How could they tell?

The thing about AI art is that, absent lots of prompt engineering, seed grinding, and touchups, you're likely to have a bunch of images that are obvious tells if your entire project is AI. Anyone trying to hide it would be spending time equivalent to just making the art themselves. There's also another advantage to having a "no obvious AI art" policy; and that's to cut down on spam. AI is extremely useful to people w…

> Anyone trying to hide it would be spending time equivalent to just making the art themselves.

That's basically the argument why Jason Allen should have been allowed to win the art competition, is it?

It's not that he typed "award winning painting" into Midjourney and the image was the result.

He tried hundreds of seeds, selected one that he liked and refined it over countless iterations with infilling until he was satisfied with the result.

I honestly don't see how this is fundamentally different from other art forms.

Re: Database of artists used to train Midjourney AI garners criticism

#120

Earlier quoted context omitted.

Furthermore, what is the threshold for something to still legally be considered AI art once an artist's hand has modified it? What if they change the brightness? Fix a hand? completely replace a character in a scene? Illustrate most of the scene themselves but add an AI figure or background? Use AI to sharpen a hand drawn image?

It's a gray area. But if someone generates AI images at scale without human in the loop it's likely not copyrightable.

What does "at scale" mean here, and how would it be detected or enforced?
Post reply on HN