Live data from Hacker News

Database of artists used to train Midjourney AI garners criticism

artnews.com

101–110 of 144 posts

Re: Database of artists used to train Midjourney AI garners criticism

#101

My personal worry is that now that companies can't simply scrape all the data for AI, they will instead start buying up massive amounts of data, turning the companies even more powerful and monolithic. Megacorps, here we come.

Remember increased regulations always favor the incumbents. Netflix was a huge proponent of fast lanes crushing net neutrality

Re: Database of artists used to train Midjourney AI garners criticism

#102
post #27

Earlier quoted context omitted.

No, it's a Butlerian Jihad[0] argument. The Copyright Office's argument holds even for a fully public domain training set. US copyright law is already speciesist[1] - you can't assign authorship to an animal - so computers are also forbidden from authorship. [0] In the Dune universe, the "Butlerian Jihad" refers to a legal ban on thinking machines. [1] https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...

I'm not really sure how this connects to the argument. No one is trying to grant authorship to an algorithm - it would be a ridiculous effort that was never even in the cards. In these copyright disputes, the authorship on AI outputs would be on the person using the AI. Generative AI takes inputs that are provided by a human and transforms it into certain outputs. Legally speaking, I don't see it as different from me…

I agree it's not quite the right argument. IANAL, but I think it's more illustrative to remove AI from the example entirely. If you wrote a prompt and gave it to a human artist to draw, would you have joint copyright over the resulting work? If you didn't do anything worthy of copyright, and the AI cannot be granted copyright, then it is not copyrightable.

That said, it seems like a moot point to me. The practical uses of generative AI are not going to be one-and-done prompt-to-image tools. When AI is used more like a brush, the brush strokes the human chooses will still be granted copyright.

Re: Database of artists used to train Midjourney AI garners criticism

#103
post #37
post #17

Earlier quoted context omitted.

> Completely unenforceable. Complete wrong. You just flip the defaults--something is AI unless you can prove otherwise. This is done already and has precedent. Producing porn requires that you keep artifacts demonstrating that who the performers were, that they were of age, etc. If you claim a work is not AI generated, you should have to produce some artifacts to back up that claim. In the case of a corporation, that…

I'm really struggling to conceptualize a world where every picture that's drawn must have a full notary log of how exactly it was produced, all for the sake of removing generative AI. Besides, it's not that easy of a problem - a lot of corporate artists are salaried workers, they don't get specified commissions with an attached bill per work, but are paid a salary so the company can ask them to draw whatever they nee…

If it's physical media, you have the physical media.

If it's digital media, the software can keep an encrypted record at the brushstroke level that can be played back to produce a bit-perfect reproduction. Maybe even write it to a public ledger.

Re: Database of artists used to train Midjourney AI garners criticism

#104
post #84

>The 24-page list of artists’ names used by Midjourney as the training foundation for its AI image generator I don't see how they reached the conclusion that this list is "the training foundation" of Midjourney. I feel like this headline is misleading, perhaps even outright false.

Headline: 16,000 artists.

3rd Para: 24 page list.

Yeah, something is off here.

Re: Database of artists used to train Midjourney AI garners criticism

#105
post #16

Earlier quoted context omitted.

The thing about AI art is that, absent lots of prompt engineering, seed grinding, and touchups, you're likely to have a bunch of images that are obvious tells if your entire project is AI. Anyone trying to hide it would be spending time equivalent to just making the art themselves. There's also another advantage to having a "no obvious AI art" policy; and that's to cut down on spam. AI is extremely useful to people w…

> Anyone trying to hide it would be spending time equivalent to just making the art themselves. The gap of being distinguishable from manually drawn images is still closing - we don't know if it'll ever reach the threshold, but the amount of effort required to stamp out all the wonkiness from an AI generation has been going down ever since the first viable algorithms appeared. I don't think that this was an anti-spam…

IME, if GenAI ever reaches human parity, whatever that amounts to, the relevant subgenre of art will just move into surrealism. Invention of paintbrushes didn't kill art.

Re: Database of artists used to train Midjourney AI garners criticism

#106
post #86

Earlier quoted context omitted.

> It is for the sake of copyright, if you want society to protect your work, provide evidence for your creative work. I'm not sure if it's that simple - for one, this requirement is a complete departure from how copyright systems work now. Providing complete history logs isn't normal practice, and expanding law to necessitate it isn't common sense. > Keep in mind that in the not so far future, producing art will be a…

> for one, this requirement is a complete departure from how copyright systems work now. A complete departure? Here is the current form used to register an artistic visual work for copyright. Its more elaborate than you might think. https://www.copyright.gov/forms/formva.pdf Registration is not a rubber-stamp, it is increasingly refused because of indicia of AI tooling. Why would adding some questions on provenance a…

Nothing in the form seems out of the ordinary to me. It is a lot of fields, but ultimately the main goal is establishing ownership, not discerning the specific methodology in which a person made the work. It's a departure in that the current system is results-based, where you register a final product, while the proposed system also must take into consideration every intricacy of creating the work.

> it is increasingly refused because of indicia of AI tooling.

Do you have a source that a statistically significant number of copyright applications gets refused on account of a work just seeming like AI? On what grounds does it get refused?

Re: Database of artists used to train Midjourney AI garners criticism

#107
post #67
post #6

Earlier quoted context omitted.

Completely unenforceable. How can you even tell if an image was made by AI? What if AI created an outline that was worked on by a human artist (or vice versa)? Who would the burden of proof be on? Steam has a “no AI art” policy, and it’s rapidly turning into a “no obvious AI art policy”. How could they tell?

"I claim this art was made by this person" "Who?" "OK . did you work on this?" "Where are related work products? Are there any? What about invoices? Simultaneous employment?" The reality is that most legal things are determined by _convincing people of a truth_. Perhaps you can set up a whole scheme to "launder" AI art and attach names to them. And all the papertrail you generate doing this will show up in discovery…

The way that AI will be laundered into art is by including it into things like Photoshop. There'll still be a human touch just with "smart brushes" and "smart auto fill" that paints 90% of what you want.

Art will then take less skill to produce, and be produced faster for lower prices.

An 80% price reduction on art (because artists can now produce it 5 times faster thanks to AI) is 80% as good as getting it for free.

Re: Database of artists used to train Midjourney AI garners criticism

#108
post #37

Earlier quoted context omitted.

I'm really struggling to conceptualize a world where every picture that's drawn must have a full notary log of how exactly it was produced, all for the sake of removing generative AI. Besides, it's not that easy of a problem - a lot of corporate artists are salaried workers, they don't get specified commissions with an attached bill per work, but are paid a salary so the company can ask them to draw whatever they nee…

> drawn must have a full notary log of how exactly it was produced I love when non-artists talk about art.

The word choices were kind of on purpose - I meant to highlight the partial absurdity of having to entangle yourself with all these legal considerations and obtaining sufficient legal proof, all for the sake of making some art piece.

Re: Database of artists used to train Midjourney AI garners criticism

#109
What is the difference between AI scanning existing art, and creating new images informed by the prior art, and a person touring galleries, looking at the works of living artists, and creating new works influenced in style, color, framing, etc. by those works? Should the works of a new artist be beholden to all those whose works they have looked upon? As another thing to consider, is there any rationale for copyrights on music (in particular) that wouldn't equally apply to the recipes food dishes created by cooks? Both build upon the works, style, ingredients, etc. of those who preceded them. Seriously interested in the arguments one would make for copyrighting music that wouldn't stand equally well for recipes.

Re: Database of artists used to train Midjourney AI garners criticism

#110
post #37

Earlier quoted context omitted.

I'm really struggling to conceptualize a world where every picture that's drawn must have a full notary log of how exactly it was produced, all for the sake of removing generative AI. Besides, it's not that easy of a problem - a lot of corporate artists are salaried workers, they don't get specified commissions with an attached bill per work, but are paid a salary so the company can ask them to draw whatever they nee…

If it's physical media, you have the physical media. If it's digital media, the software can keep an encrypted record at the brushstroke level that can be played back to produce a bit-perfect reproduction. Maybe even write it to a public ledger.

All of these things have loopholes. For physical media, depending on the quality of the output, one could pay a sufficiently skilled person to reproduce an AI output on physical media in a fraction of the time it'd take to come up with and draw for real.

For digital media - ignoring how overbearing this whole system could be, what prevents someone from taking all that data and making an algorithm that outputs brush stroke parameters instead of pixels? And digital art isn't the only thing we need to concern ourselves with - eventually, we might have AI models that could make 3D models, sounds, vector imagery and other forms of art. The idea of just documenting every workflow would be an ever-growing burden with no perfect solutions.

Post reply on HN