Live data from Hacker News

Database of artists used to train Midjourney AI garners criticism

artnews.com

131–140 of 144 posts

Re: Database of artists used to train Midjourney AI garners criticism

#131
post #16

Earlier quoted context omitted.

> Anyone trying to hide it would be spending time equivalent to just making the art themselves. The gap of being distinguishable from manually drawn images is still closing - we don't know if it'll ever reach the threshold, but the amount of effort required to stamp out all the wonkiness from an AI generation has been going down ever since the first viable algorithms appeared. I don't think that this was an anti-spam…

IME, if GenAI ever reaches human parity, whatever that amounts to, the relevant subgenre of art will just move into surrealism. Invention of paintbrushes didn't kill art.

the fight about AI is using copyrighted stuff for its weights... i wonder which % of artists that wouldn't tweak or use heavily AI that has a transparent/ethical data-base (read it: they didn't added anything proprietary without authorization)

Re: Database of artists used to train Midjourney AI garners criticism

#132
post #37
post #17

Earlier quoted context omitted.

> Completely unenforceable. Complete wrong. You just flip the defaults--something is AI unless you can prove otherwise. This is done already and has precedent. Producing porn requires that you keep artifacts demonstrating that who the performers were, that they were of age, etc. If you claim a work is not AI generated, you should have to produce some artifacts to back up that claim. In the case of a corporation, that…

I'm really struggling to conceptualize a world where every picture that's drawn must have a full notary log of how exactly it was produced, all for the sake of removing generative AI. Besides, it's not that easy of a problem - a lot of corporate artists are salaried workers, they don't get specified commissions with an attached bill per work, but are paid a salary so the company can ask them to draw whatever they nee…

I think there's a lot of confusion about what records are needed presumably due to lack of understanding of what the law requires for proof of ownership.

Generally the party asserting ownership has the burden of proof. The standard is "preponderance of the evidence", which generally is understood to mean "more likely than not" or "> 50%". So basically it means if you can prove to judge or jury that there's a >50% chance you own the work, it's good enough.

Also note that in many cases where there's a dispute over the evidence, witnesses are summoned to testify. So you might not have a "full notary log" of how it was produced or all the "intermediate artifacts", but as long as the artist is able to convincingly explain how the work was created, and the other party's lawyers are not able to poke holes in their story during cross evidence, that's usually enough.

Which is, basically, what happens today, if the authorship or ownership of a work is disputed.

That said, I'm not sure whether "assume work AI (thus uncopyrightable) unless proven otherwise" should be the default for other reasons. For one, most quality "AI art" needs some manual adjustments or touch ups, and arguably the prompt and hyperparameters may be sufficient creativity element. I mean, that's basically how copyrights dealt with photography (the mere fact you decided when and where to point the camera with what settings is sufficient for copyright to subsist in a photo).

Re: Database of artists used to train Midjourney AI garners criticism

#133
post #3

One positive aspect of the status quo in the United States is that AI-generated images are not currently eligible for copyright. I think this is a great direction to go in, I highly doubt Wizards of the Coast or whoever is going to want their premium products to lose copyright protections, so they'll need to keep paying artists. I'd love for us to lean into this -- you can make all the AI art you want, but it automat…

One question that is unclear to me is how this works if images are packaged with text or other content. For example; let’s say I write a book and then use AI images to illustrate it. It doesn’t seem logical to me that somehow the book would be copyrighted but the images inside the book wouldn’t be…? At some level, the “package” of images + text would supersede the two things separately. Otherwise you would have a sit…

There's nothing particularly contradictory about that: there's already situations where that is the case. For example when a book contains images which are public domain, or where elements of the book like the facts within it are not copyrightable. Another interesting one is tabletop game manuals: the layout and presentation of the rules are copyrightable, but the game mechanics generally aren't. So you can make a book which just contains the rules and not be infringing copyright. Using AI-generated images would be exactly the same situation.

Re: Database of artists used to train Midjourney AI garners criticism

#134

Earlier quoted context omitted.

One question that is unclear to me is how this works if images are packaged with text or other content. For example; let’s say I write a book and then use AI images to illustrate it. It doesn’t seem logical to me that somehow the book would be copyrighted but the images inside the book wouldn’t be…? At some level, the “package” of images + text would supersede the two things separately. Otherwise you would have a sit…

There's nothing particularly contradictory about that: there's already situations where that is the case. For example when a book contains images which are public domain, or where elements of the book like the facts within it are not copyrightable. Another interesting one is tabletop game manuals: the layout and presentation of the rules are copyrightable, but the game mechanics generally aren't. So you can make a bo…

I'm envisioning more of a situation where a company adds text directly to AI-generated images, or otherwise somehow modifies them that prevents them from just being generic images, in the way public domain images are. I really don't think companies will just add images straight-from-the-generator without modifying them in such a way that prevents their easy re-use.

Re: Database of artists used to train Midjourney AI garners criticism

#135

Earlier quoted context omitted.

It's a gray area. But if someone generates AI images at scale without human in the loop it's likely not copyrightable.

What does "at scale" mean here, and how would it be detected or enforced?

that means that openai cannot claim copyright on images produced by dalle generator. no can other online services and offline software owners/producers.

this is enough to use them without copyright violation for example for other ai models training.

Re: Database of artists used to train Midjourney AI garners criticism

#136

I do find it entertaining to watch the creative world struggle against this unstoppable force. It's the Luddites and the mechanical knitting machines of the early 19th century all over again. The human race never changes. It will undoubtedly end the same way.

Some said that Napster/KaZaA/etc. was the future of music. They got closed down.

Spotify eventually popped up, and "solved" a big problem in music streaming. But at a cost - the record labels have them by their balls.

You don't see any leading "pirate" Spotify clone dominating the market. Closest thing is YouTube, but monetization goes to the content owners.

So here's what you do, if you're a bunch of creative artists / content owners / etc.:

You find one company that's commercializing the tech, and sue them into oblivion. Force them into signing a deal that ensures they'll never be profitable.

Others will follow.

But, yes, obviously the models are out there. In the same way that people are pirating stuff like they've always done. You can't do much about individuals running models on their own rigs - but that's not the point.

The point is to make it as painful as possible for people that feel the urge to commercialize something at scale. The goal is "Fuck you, pay me" or get dragged through the courts.

Re: Database of artists used to train Midjourney AI garners criticism

#137
post #131

Earlier quoted context omitted.

IME, if GenAI ever reaches human parity, whatever that amounts to, the relevant subgenre of art will just move into surrealism. Invention of paintbrushes didn't kill art.

the fight about AI is using copyrighted stuff for its weights... i wonder which % of artists that wouldn't tweak or use heavily AI that has a transparent/ethical data-base (read it: they didn't added anything proprietary without authorization)

It's been tried. Numerous times. There's a reason why GenAI controversy is stuck at ethics and filled with rage, the generated images just aren't that great and so that part isn't so controversial.

Re: Database of artists used to train Midjourney AI garners criticism

#138
post #28

Earlier quoted context omitted.

I used to agree with this, but there's been research coming out of Google that has altered my opinion. Specifically, Google's gotten rather good at making AI spit out unaltered training set data[0]. This is only possible if the AI is remembering large portions of the original trained-on works, which would make the weights infringing. [0] In the most egregious case, they found that just asking ChatGPT to repeat a word…

> This is only possible if the AI is remembering large portions of the original trained-on works which is fine in my books. You could equivalently produce the same works from searching thru the digits of pi. The only enforcement that's needed is on the end user of the model - if they choose to produce the training set, they are violating copyright. The creator of the model _does not_ violate copyright merely by creat…

Courts won't buy this, at least not completely. Copyright infringement liability accrues to the entire value chain - model users and developers alike. So the only way for a model to see copyrighted data during training and be non-infringing is if there's no conceivable way to reproduce training data.

I don't think fair use saves us either, at least not the model authors, because the whole selling point of these models is to replace artists. Yes, artists can use them as tools, but that is far less lucrative. The valuations and hype being thrown around specifically come from, among other things, being able to cut the creative class out of their own business. No court is going to look at that and say "ok, yeah, sure, it's perfectly fine for you to be using other people's work to train on".

An individual using an AI art generator may still wind up getting novel output that isn't obviously infringing to anything in the dataset. If that's the case then they probably haven't infringed copyright. Or at least, it'd be difficult to make a case around it.

Re: Database of artists used to train Midjourney AI garners criticism

#139
post #95

Earlier quoted context omitted.

> The gap of being distinguishable from manually drawn images is still closing People have been trumpeting this since day one of Stable Diffusion releasing, but I'm seeing the same output quality as that day and I've been keeping up.

Just because the pace of progress isn't exponential (like what some people would want to believe) doesn't mean it isn't happening. I remember getting an early invite to DALL-E 1 all the way back, and while I don't use it anymore, the modern improvements made seem very substantial. From plain comparisons of different versions where the same inputs produce substantially better outputs, to the mere fact that the latest…

I'm telling you the progress isn't happening based on my own consistent observations of various releases across multiple platforms. The only people who don't seem to agree with me are those who have the art literacy of a highschooler and think "discernible text" is a improvement.

As an aside, no one said AI couldn't achieve drawn generated text, that's been possible for years prior to stable diffusion.

Re: Database of artists used to train Midjourney AI garners criticism

#140
post #49

Earlier quoted context omitted.

That's a really weird argument. Copyright is a legal system created by the Constitution and statutes and administrative rules. It cares about whether you are "copying", and it cares about whether you're creating things that compete with the works of the original authors. It doesn't care about potential output spaces. In this context, I don't see a principled difference between the model weights and really good compre…

You can eke out large chunks of works from Google Books with the right queries too.

Yes, but Google Books is a search engine. It doesn't write books, it just tells you where a particular phrase might occur in those books. There's explicit caselaw allowing you to do this, extending back before Google was even a thing. For related reasons, Google Books also does not let you read the whole book - just the page the search match came from.

OpenAI and other large language model developers are claiming they have a machine that can write books, but they also fed it shittons of books, and they can't account for where all that text went. At best they can say "well, it doesn't produce exact, verbatim copies of the training set all the time".

Post reply on HN