Live data from Hacker News

On being listed as an artist whose work was used to train Midjourney

catandgirl.com

521–530 of 957 posts

Re: On being listed as an artist whose work was used to train Midjourney

#521
post #166

Earlier quoted context omitted.

AI doesn't get inspired. It's not human. It adds everything about it to its endless stream of levers to pull, and if you pull the right ones, it will just give you the source verbatim as proven by the NYT lawsuit filing where it was just outputting unaltered copywritten NYT article text.

If you pull the right levers, you can also copy the NYT article fully.

Yeah and that's copyright infringement. That's why if you're reverse engineering something it needs to be done in a clean room environment, your prior exposure to the copywritten material poisons the well for any derivative you create.

This extends to music as well, if someone hears a song and is inspired by that in their work, the original artist gets credit.

Re: On being listed as an artist whose work was used to train Midjourney

#522

> But I can't even get cartoons to most people for free now, without doing unpaid work for the profit-making companies who own the most use channels of communication This is the sticking point for me. OpenAI isn't a profit-making company, but it's certainly a valuable company. A valuable company that is built from the work of content others created without transferring any value back to them. Regardless of legalities…

>> If you think OpenAI is less valuable because it can't use copyrighted content, then it should give some of that value back to the content. But we are allowed to use copyrighted content. We are not allowed to copy copyrighted content. We are allowed to view and consume it, to be influenced by it, and under many circumstances even outright copy it. If one doesn't want anyone to see/consume or be influenced by one's…

>I have some, but diminishing sympathy for artists screaming about how AI generates images too similar to their work. Yes, the output does look very similar to your work. But if I take your work and compare it to the millions of other people's work, I'd bet I can find some preexisting human-made art that also looks similar to your stuff too.

Just like the rest of AI, if your argument is "humans can already do this by hand, why is it a problem to let machines do it?", its because you are incorrectly valuing the labor that goes into doing it by hand. If doing X that has potentially negative side effect Y, then the human labor to accomplish X is the principle barrier to Y, which can be mitigated via existing structures. Remove the labor barrier, and the existing mitigation structures cease to be effective. The fact that we never deliberately established those barriers is irrelevant to the fact that our society expects them to be there.

Re: On being listed as an artist whose work was used to train Midjourney

#523

There is simple way to fix this. Ban private large models trained on public data, require them to be public weights. If a company wants to train large private model, they can do it with their own data.

I don't see many people suggesting this, but I also quite like this way of thinking about it. The idea that it should be illegal for models to learn from artists, or that artists have a right to extract payment from the model, doesn't make much sense to me, it's too much of a radical departure from the way we treat human learning. But it seems unfair that a company can own such a model. It's not their work, it's a co…

> We should all own it.

Maybe, just maybe better to ask artist, who should own it?

Re: On being listed as an artist whose work was used to train Midjourney

#524
post #41

Earlier quoted context omitted.

> "The bottom line is this," the firm, known as a16z, wrote. "Imposing the cost of actual or potential copyright liability on the creators of AI models will either kill or significantly hamper their development." > The firm said payment for all of the copyrighted material already used in LLMs would cost the companies that built them "tens or hundreds of billions of dollars a year in royalty payments." https://www.bus…

Doesn't that kind of demonstrate the value being actively stolen from the creators, more than anything? Copyright law killed Napster, too. That doesn't mean applying copyright law was wrong.

Of course it was wrong. Abolish all copyright.

Re: On being listed as an artist whose work was used to train Midjourney

#525
post #5

The sooner anyone making profit from models trained on creators proprietary content start paying for the content they’re using the better for creators, society and even the AI companies. It’s pretty tiring hearing people argue about whether copyright law applies to AI companies or not. It applies. Just get on and sort out a proper licensing model.

I think it's worth noting that Adobe Firefly and iStock already have generators that use only licensed content.

Highly doubtful.

Yes, I know Adobe said so. No, I don't trust them.

Facts:

1. Adobe Firefly is trained with Adobe Stock assets. [1]

2. Anyone can submit to Adobe Stock.

3. Adobe Stock already has AI-generated assets that are not correctly tagged so. [2]

4. It's hard to remove an image from a trained model.

Unless Adobe carefully scrutinize every image in the training set, the logical conclusion is Adobe Firefly already contains at least second-handed unauthorized images (e.g. those generated by Stable Diffusion). It's just "not Adobe's fault".

[1] https://www.adobe.com/products/firefly.html : "The current Firefly generative AI model is trained on a dataset of licensed content, such as Adobe Stock, and public domain content where copyright has expired."

[2] Famous example: https://twitter.com/destiny_thememe/status/17448423657672255...

Re: On being listed as an artist whose work was used to train Midjourney

#526

Earlier quoted context omitted.

For what it's worth: that's pretty much how Translate works. Translate operates at a large-chunk resolution, and one of the insights in solving the problem was the idea that you can often get a pretty-good-enough translation by swapping a whole sentence for another whole sentence. So they ingest vast amounts of pre-translated content (the UN publications are a great source, because they have to be published in the la…

Google translate is very basic and not even close to something good if you already know both languages. Useful if you're translating to your language (you do the correction when reading), but can lead to confusion the other way.

Interesting distinction.

If you can do the correction when reading, it seems reasonable to assume the reader in the opposite direction has the same correction capability.

I would expect the chance of confusion to be identical. The only difference is a matter of perspective, where in one case you are the reader and in one case you are the author.

Re: On being listed as an artist whose work was used to train Midjourney

#527
post #397
post #177

Earlier quoted context omitted.

I firmly believe that training models qualifies as fair use. I think it falls under research, and is used to push the scientific community forward. I also firmly believe that commercializing models built on top of copyrighted works (which all works start off as) does not qualify as fair use (or at least shouldn't) and that commercializing models build on copyrighted material is nothing more than license laundering. C…

> I firmly believe that training models qualifies as free use. I think it falls under research, and is used to push the scientific community forward. I don't think this is as cut-and-dry and you make it out here. If I train a model on, say, every one of New York Times' and release it for free and it finds use as a way of circumventing their paywall I have difficulty justifying that as fair use/fair dealing. The purpo…

Wouldn't that depend on the use case? If you just had the model regenerate articles that roughly approximate its source material that is much a more clear cut violation of a paywall. But if you use that data as general background knowledge to synthesize aggregative works such a history of the vietnam war, or trends in musical theatre in the 1980s relative the 1970s, or shifts in the language usage of formal honorifics, then that seems to me to be clearly fair-use categories. There are gray areas, such as aggregating the opinions of a certain op-ed writer over a short timeframe that while it might produce a novel work, is basically is mixmash of recent articles. But would that be unfair, especially if not done in the original authors style?

These technical distinctions like these probably will matter in whatever form regulation eventually ends up becoming.

Re: On being listed as an artist whose work was used to train Midjourney

#528
post #139

Earlier quoted context omitted.

Because we are humans and our capability of abusing those rights is limited. The scale and speed at which LLMs can abuse copyrighted work to threaten the livelihoods of the authors of those works is reason enough to consider it unethical.

Are photocopy machines illegal? Are CD-ROM burners illegal? Both allow near-unlimited copies of copyrighted material at a scale much faster than a human could do alone. The tools are not the problem, it's how humans use them.

>CD-ROM burners

They can be used in an illegal way if used to copy copyrighted material, yes.

Re: On being listed as an artist whose work was used to train Midjourney

#529
post #5

The sooner anyone making profit from models trained on creators proprietary content start paying for the content they’re using the better for creators, society and even the AI companies. It’s pretty tiring hearing people argue about whether copyright law applies to AI companies or not. It applies. Just get on and sort out a proper licensing model.

It's pretty tiring hearing people think anything like what you're saying is going to happen. The horse is out of the barn already, deal with it.

[dead]

Re: On being listed as an artist whose work was used to train Midjourney

#530
post #177

Earlier quoted context omitted.

I firmly believe that training models qualifies as fair use. I think it falls under research, and is used to push the scientific community forward. I also firmly believe that commercializing models built on top of copyrighted works (which all works start off as) does not qualify as fair use (or at least shouldn't) and that commercializing models build on copyrighted material is nothing more than license laundering. C…

> The reason that AI models are generating content similar to other people's work is because those models were explicitly trained to do that. Ah, just like humans who train against the output of other humans. AI models are not fundamentally different in kind in this regard, only scope, and even that isn't perfectly obvious to me a priori.

Oh, right. It just reads a million books in a couple of days, removes all the source information, mix and match it the way it sees fit and sells this output $10/month to anyone comes with a credit card.

It's the same thing with GitHub's copilot.

A book publisher would seize everything I have, and shot me at a back alley if I do 0.0001% of this.

Post reply on HN