Live data from Hacker News

On being listed as an artist whose work was used to train Midjourney

catandgirl.com

461–470 of 957 posts

Re: On being listed as an artist whose work was used to train Midjourney

#461

Earlier quoted context omitted.

I suspect a single verbatim output of sufficient length is enough to poison the entire weight set as a derivative work as well as all the output it ever generated

I dont necessarily disagree, but by what logic or argument do you make that case? Does that mean that models that can not produce copies of X length ARE fair use?

> Does that mean that models that can not produce copies of X length ARE fair use?

not necessarily

"sufficient but not necessary" I believe is the term

Re: On being listed as an artist whose work was used to train Midjourney

#462
post #414

I'd find it hard to argue against this, or the Penny Arcade's statements, since I'm having trouble understanding their concrete arguments in between the rhetoric. I'd be hesitant to even discuss this in their comment sections or social media channels. One might ask: Under what circumstances would AI art be acceptable then? For example, does it really matter if these models are created by large corporations? I don't s…

I think these are the interesting question. Especially if the user provided the context Style for samples.

It's currently fair use to give an artist paintings and say I want something like this but different in these ways.

You can tell a script writer to watch Star Wars and write something similar.

Questions of copyright will depend on if the output is sufficiently transformative, not if copyrighted work was used as inspiration

Re: On being listed as an artist whose work was used to train Midjourney

#463
Evolve or die. The world is cruel.

This is true for everyone. Due to human nature, we will make AI increasingly capable at an exponential pace. Eventually, our efforts to control them will become futile, and humans will no longer be the dominant species.

Luckily, we are uniquely positioned in evolution to prevent this. We must merge with AI.

Re: On being listed as an artist whose work was used to train Midjourney

#464

Earlier quoted context omitted.

> Are there any humans that can produce artwork without ingesting inspiration from other art? Do you know any artists that lived in a box their whole life and never saw other art? Do you know any writers who'd never read a book? > Are they any human artists who can't, if requested, draw or write something that's a copy of some other person's drawings or writings? This still is pretending that humans and AI models are…

This isn't about giving "rights" to machines. Machines are just tools. The question is about what humans are allowed to do with those tools. Are humans using AI models and humans not using AI models equivalent actors that should have the same rights? I'd argue emphatically yes they should.

The thing is, we already have doctrine that starts to encompass some of these concepts with fair use.

The four pronged test in US case law:

- the purpose and character of use (is a machine doing this different in purpose and character? many would say yes. is "ripping-off-this-artist-as-a-service" different than an isolated work that builds upon another artist's art?)

- the nature of the copyrighted work

- the amount and substantiality of the portion taken (can this be substantially different with AI?)

- the effect of the use upon the potential market for the original work (might mechanization of reproducing a given style have a larger impact than an individual artist inspired by it?)

These are well balanced tests, allowing me as a classroom teacher to duplicate articles nearly freely but preventing me from duplicating books en masse for profit (different purpose; different portion taken; different impact on market).

Re: On being listed as an artist whose work was used to train Midjourney

#465

There is simple way to fix this. Ban private large models trained on public data, require them to be public weights. If a company wants to train large private model, they can do it with their own data.

If you want a large model with public weights, just do it yourself. Nobody is stopping you.

Re: On being listed as an artist whose work was used to train Midjourney

#466

> But I can't even get cartoons to most people for free now, without doing unpaid work for the profit-making companies who own the most use channels of communication This is the sticking point for me. OpenAI isn't a profit-making company, but it's certainly a valuable company. A valuable company that is built from the work of content others created without transferring any value back to them. Regardless of legalities…

I keep thinking: this is what eg Google has done all along. The content it uses to train models and present the answers to us absolutely belongs to others, but you try get any content (eg maps data) out of it for free at scale.

But the business model emerged and delivered value to us while enough that we didn’t consider asking for money for our content. We like being searched and linked to. Less so Google snippets presented to users without the users landing on our site. Even less so generated without any interaction. But it’s all still all our content.

Re: On being listed as an artist whose work was used to train Midjourney

#467

Earlier quoted context omitted.

The new York Times has examples where GPT will reproduce world for word exactly paragraphs of their (copyrighted) text if you ask it to. That's a pretty fixed tangible expression I think.

For sure, that could be an instance of infringement depending on how it is used. But that's a minuscule percentage of the output and still might be fair use (read the decision in Authors Guild, Inc. v. Google, Inc.). But even if that instance is determined to be infringement, it doesn't mean the process of training models on copyrighted work is also infringement.

I can see 3 ways that you can guarantee that the output of a model never violates copyright

1. Models are trained with 100% uncopyrighted or properly licensed input data

2. Every output of the ML model is evaluated to make sure it's not too close to training data

3. Copyright law is changed to have a specific cutout for AI

#1 is the approach taken by Adobe, although it generally is harder or more expensive to do.

#2 destroys most AI business models

#3 has been done in some countries, but seems likely that if done in the US it would still have some limits.

For example, I could train a model on a single image, song, or piece of written text/code. Then I run inference, and get out an exact copy of that image, song, or text. If there are no limits around AI and copyright, then we've got a loophole around all of copyright law. I don't think that the US would be up for devaluing intellectual property like that.

Re: On being listed as an artist whose work was used to train Midjourney

#468

Earlier quoted context omitted.

We shouldn't hold individual humans and ML models to the same standards, because ML models themselves are products capable of mass production and individual humans are not even remotely at the same scale. If you write that book, chances are you will gain some fans that are also fans of other authors in that genre. If ML models write that genre, they can flood that genre so full that human artists won't be able to com…

Computers and machines have been capable of mass production for decades, and humans have used them as tools. In the past 170 years, these tools of mass production have already diminished many thousands of professions that were staffed by people who had to painstakingly craft things one at a time. Why is art some special case that should be protected, when many other industries were not? Why should we kill this techno…

What if the AI was solely trained on this person's work, then from that churned out a similar replacement that was monetized?

Re: On being listed as an artist whose work was used to train Midjourney

#469

> But I can't even get cartoons to most people for free now, without doing unpaid work for the profit-making companies who own the most use channels of communication This is the sticking point for me. OpenAI isn't a profit-making company, but it's certainly a valuable company. A valuable company that is built from the work of content others created without transferring any value back to them. Regardless of legalities…

>> If you think OpenAI is less valuable because it can't use copyrighted content, then it should give some of that value back to the content. But we are allowed to use copyrighted content. We are not allowed to copy copyrighted content. We are allowed to view and consume it, to be influenced by it, and under many circumstances even outright copy it. If one doesn't want anyone to see/consume or be influenced by one's…

[deleted]

Re: On being listed as an artist whose work was used to train Midjourney

#470

Earlier quoted context omitted.

"So you're saying that, we should stop pursuing art and prose?", no, it becomes a hobby like any other. People still sew for fun.

Great, hold on, I'm calling Hollywood to tell that all they do is a hobby now. ...and the writers' guild, too.

Well obviously AI isn't at the level of replacing Hollywood yet.

But once it is? I mean, yeah, it'll replace Hollywood.

People will tell Netflix, "hey I want a move about X in the style of Y and I want Z to star in it", and bam -- your own bespoke movie.

I mean, once the capability's there, it's just inevitable. And yeah -- acting will become a hobby, just like sewing is today.

Post reply on HN