Live data from Hacker News

Database of artists used to train Midjourney AI garners criticism

artnews.com

91–100 of 144 posts

Re: Database of artists used to train Midjourney AI garners criticism

#91
post #37
post #17

Earlier quoted context omitted.

> Completely unenforceable. Complete wrong. You just flip the defaults--something is AI unless you can prove otherwise. This is done already and has precedent. Producing porn requires that you keep artifacts demonstrating that who the performers were, that they were of age, etc. If you claim a work is not AI generated, you should have to produce some artifacts to back up that claim. In the case of a corporation, that…

I'm really struggling to conceptualize a world where every picture that's drawn must have a full notary log of how exactly it was produced, all for the sake of removing generative AI. Besides, it's not that easy of a problem - a lot of corporate artists are salaried workers, they don't get specified commissions with an attached bill per work, but are paid a salary so the company can ask them to draw whatever they nee…

> drawn must have a full notary log of how exactly it was produced

I love when non-artists talk about art.

Re: Database of artists used to train Midjourney AI garners criticism

#92
post #6
post #3

One positive aspect of the status quo in the United States is that AI-generated images are not currently eligible for copyright. I think this is a great direction to go in, I highly doubt Wizards of the Coast or whoever is going to want their premium products to lose copyright protections, so they'll need to keep paying artists. I'd love for us to lean into this -- you can make all the AI art you want, but it automat…

Completely unenforceable. How can you even tell if an image was made by AI? What if AI created an outline that was worked on by a human artist (or vice versa)? Who would the burden of proof be on? Steam has a “no AI art” policy, and it’s rapidly turning into a “no obvious AI art policy”. How could they tell?

A company today has the burden of proof to demonstrate authorship of a claimed work when they sue for copyright infringement. This isn't a crazy expansion of that concept -- companies do not generally break the law just because there's a low chance of getting caught. Furthermore it's not hard to imagine, for example, a whistleblower calling out their employer for copyrighting AI works.

Re: Database of artists used to train Midjourney AI garners criticism

#93
post #27

Earlier quoted context omitted.

No, it's a Butlerian Jihad[0] argument. The Copyright Office's argument holds even for a fully public domain training set. US copyright law is already speciesist[1] - you can't assign authorship to an animal - so computers are also forbidden from authorship. [0] In the Dune universe, the "Butlerian Jihad" refers to a legal ban on thinking machines. [1] https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...

I'm not really sure how this connects to the argument. No one is trying to grant authorship to an algorithm - it would be a ridiculous effort that was never even in the cards. In these copyright disputes, the authorship on AI outputs would be on the person using the AI. Generative AI takes inputs that are provided by a human and transforms it into certain outputs. Legally speaking, I don't see it as different from me…

This is literally a case of someone making AI art and trying to attribute it to the algorithm.

Re: Database of artists used to train Midjourney AI garners criticism

#94
post #74

I do find it entertaining to watch the creative world struggle against this unstoppable force. It's the Luddites and the mechanical knitting machines of the early 19th century all over again. The human race never changes. It will undoubtedly end the same way.

It's really easy, "unstoppable" even, to send pirated music files over the internet. It sure was fun to have free music. And yet Napster was sued out of existence, and now the vast majority of people who listen to music pay for it. And it's way harder to run a pirate AI datacenter than it is to set up a torrent tracker. Laws will be written, license fees will be paid, and AI will continue to develop but with proper c…

The problem here is that the law has already been written to some extent. Google Books is the most relevant case you can point to that set precedent over what constitutes fair use of copyrighted material and the extent the transformation of those works can be commercially exploited. In that case works were copied without permission with small chunks of those works redistributed verbatim, and that was when it was working as designed. There's an argument to be made that this AI generator stuff is actually less egregious.

Re: Database of artists used to train Midjourney AI garners criticism

#95
post #16

Earlier quoted context omitted.

> Anyone trying to hide it would be spending time equivalent to just making the art themselves. The gap of being distinguishable from manually drawn images is still closing - we don't know if it'll ever reach the threshold, but the amount of effort required to stamp out all the wonkiness from an AI generation has been going down ever since the first viable algorithms appeared. I don't think that this was an anti-spam…

> The gap of being distinguishable from manually drawn images is still closing People have been trumpeting this since day one of Stable Diffusion releasing, but I'm seeing the same output quality as that day and I've been keeping up.

Just because the pace of progress isn't exponential (like what some people would want to believe) doesn't mean it isn't happening. I remember getting an early invite to DALL-E 1 all the way back, and while I don't use it anymore, the modern improvements made seem very substantial. From plain comparisons of different versions where the same inputs produce substantially better outputs, to the mere fact that the latest version can actually generate decent, often discernible text at all (something that people joked would be impossible from AI to achieve) shows that some progress is being made.

The reason why it's not as visible with Stable Diffusion is because a lot of the technologies around it circle the same few foundational SD models - people build on top of them, add new ways of interacting with them, but ultimately, the same thing underlies them all. Community support is seen as more important than cutting-edge tech, which is why something like Stable Diffusion XL hasn't even seen universal adoption yet.

Re: Database of artists used to train Midjourney AI garners criticism

#96

My personal worry is that now that companies can't simply scrape all the data for AI, they will instead start buying up massive amounts of data, turning the companies even more powerful and monolithic. Megacorps, here we come.

Or they'll just make a branch entity in Japan where the government explicitly says that copyright does not apply to training AI.

Re: Database of artists used to train Midjourney AI garners criticism

#97
post #49
post #28

Earlier quoted context omitted.

> This is only possible if the AI is remembering large portions of the original trained-on works which is fine in my books. You could equivalently produce the same works from searching thru the digits of pi. The only enforcement that's needed is on the end user of the model - if they choose to produce the training set, they are violating copyright. The creator of the model _does not_ violate copyright merely by creat…

That's a really weird argument. Copyright is a legal system created by the Constitution and statutes and administrative rules. It cares about whether you are "copying", and it cares about whether you're creating things that compete with the works of the original authors. It doesn't care about potential output spaces. In this context, I don't see a principled difference between the model weights and really good compre…

You can eke out large chunks of works from Google Books with the right queries too.

Re: Database of artists used to train Midjourney AI garners criticism

#98
post #86
post #51

Earlier quoted context omitted.

It is for the sake of copyright, if you want society to protect your work, provide evidence for your creative work. It seems rather simple to me. Keep in mind that in the not so far future, producing art will be as cheap as consuming it, this means that the original benefits society got in return for copyright no longer applies, so why should they protect it?

> It is for the sake of copyright, if you want society to protect your work, provide evidence for your creative work. I'm not sure if it's that simple - for one, this requirement is a complete departure from how copyright systems work now. Providing complete history logs isn't normal practice, and expanding law to necessitate it isn't common sense. > Keep in mind that in the not so far future, producing art will be a…

  > for one, this requirement is a complete departure from how copyright systems work now.
A complete departure? Here is the current form used to register an artistic visual work for copyright. Its more elaborate than you might think.

https://www.copyright.gov/forms/formva.pdf

Registration is not a rubber-stamp, it is increasingly refused because of indicia of AI tooling.

Why would adding some questions on provenance and methodology be beyond the pale?

Re: Database of artists used to train Midjourney AI garners criticism

#99
post #3

One positive aspect of the status quo in the United States is that AI-generated images are not currently eligible for copyright. I think this is a great direction to go in, I highly doubt Wizards of the Coast or whoever is going to want their premium products to lose copyright protections, so they'll need to keep paying artists. I'd love for us to lean into this -- you can make all the AI art you want, but it automat…

One question that is unclear to me is how this works if images are packaged with text or other content. For example; let’s say I write a book and then use AI images to illustrate it. It doesn’t seem logical to me that somehow the book would be copyrighted but the images inside the book wouldn’t be…? At some level, the “package” of images + text would supersede the two things separately. Otherwise you would have a situation where sharing the book is a copyright violation but sharing the images inside of it isn’t.

Re: Database of artists used to train Midjourney AI garners criticism

#100
post #46

Earlier quoted context omitted.

I used to agree with this, but there's been research coming out of Google that has altered my opinion. Specifically, Google's gotten rather good at making AI spit out unaltered training set data[0]. This is only possible if the AI is remembering large portions of the original trained-on works, which would make the weights infringing. [0] In the most egregious case, they found that just asking ChatGPT to repeat a word…

I'm kind of confused - you've claimed that Google's AI was broken, but cite an anecdote over ChatGPT? Regardless, even if these cases did happen often enough, it's erroneous to assume that this is something universal (i.e. can be applied to all generative AI models) or intentional. Said models are vastly smaller than the sizes of their training datasets, so it's more or less impossible for all the data to be stored v…

I think you’re a bit confused—these models are trained by copying works, and the memorization is the incontrovertible proof. Imagine if you download a font and use it to make a logo without paying to license the font.
Post reply on HN