Live data from Hacker News

On being listed as an artist whose work was used to train Midjourney

catandgirl.com

911–920 of 957 posts

Re: On being listed as an artist whose work was used to train Midjourney

#911
post #588

Earlier quoted context omitted.

Wouldn't that depend on the use case? If you just had the model regenerate articles that roughly approximate its source material that is much a more clear cut violation of a paywall. But if you use that data as general background knowledge to synthesize aggregative works such a history of the vietnam war, or trends in musical theatre in the 1980s relative the 1970s, or shifts in the language usage of formal honorific…

Yes, I think this is a rather fact-specific inquiry. My main point is that the research/commercial distinction is not the only factor (and not even the most important one). > if you use that data as general background knowledge to synthesize aggregative works such a history of the vietnam war, or trends in musical theatre in the 1980s relative the 1970s, or shifts in the language usage of formal honorifics, then that…

> And if they changed the dataset to include a plurality of additional books which happen to not be about Vietnam, I don't think that changes the analysis substantially.

I think the question is if it changes analysis if the dataset DOES include a bunch of books and articles related to Vietnam beyond your specific book.

In the first cast where it just rewriting a single books content, the unfairness is clear.

But in case where it is producing a new synthesis and analysis of the data, derived in part, but not regurgitating the source material, is that unfair?

The latter isn't clear to me.

Re: On being listed as an artist whose work was used to train Midjourney

#912
post #893
post #532

Earlier quoted context omitted.

This is what licensing negotiations are for. One doesn't get to throw up their hands and say "I don't know how to fairly pay you so I won't pay you at all".

Your argument is ridiculous, because it could identically be applied to "every human artist should have to pay a license to every artist whose work they were inspired by". That would obviously be a horrible future, but megacorps like Disney would love it.

I definitely didn't intend such an implication, and I don't think it follows from what I said.

Re: On being listed as an artist whose work was used to train Midjourney

#913

Earlier quoted context omitted.

Googles search engine is not selling derivative works. If you search for a Disney movie on Google search, it does not try to sell you a film derived from the movie.

They sell you ad space on full Disney movies (re)uploaded by random people who are not affiliated with Disney though: https://www.google.com/search?q=finding+nemo+full+movie I can also get Disney coloring book pages directly from Google's cache on Google images: https://www.google.com/search?q=disney+princess+coloring+boo... Authors Guild, Inc. v. Google, Inc. determined that Google's wholesale scanning and uploading…

In your first two examples, Google is still not providing an alternative to the original content, the people who have uploaded the content are and they are doing so illegally.

Also, Google can and will delist stuff that is violating copyright if you report it: https://transparencyreport.google.com/copyright/overview

The Google Books thing is a lot more interesting I think. I guess the idea is that Google is acting as a Library, and they are lending you the book? I'm not sure how I feel about this and I would need to do a lot more research on it before having a strong opinion.

The other big difference here is that you can't use the content linked from Google search as if it were your own. If you Google search "nes emulator in javascript", and you get a link to a Github repo, you can't copy paste the code as if it were your own, and even basing your code on what you have saw could be risky depending on the license of the repo. LLMs are acting as a sort of search engine, pretty similar to how Google search does, but people are using the output from the "search" as if it were their own work that they have full rights to!

Re: On being listed as an artist whose work was used to train Midjourney

#914

Earlier quoted context omitted.

We shouldn't hold individual humans and ML models to the same standards, because ML models themselves are products capable of mass production and individual humans are not even remotely at the same scale. If you write that book, chances are you will gain some fans that are also fans of other authors in that genre. If ML models write that genre, they can flood that genre so full that human artists won't be able to com…

I feel like the issue here, is you are giving AIs agency. AIs are not magic. They are tools. They are not alive, they do not have agency. They do not do things by themselves. Humans do things, some humans use AI to do those things. Agency always rests with a combination of the tool's creator and operator, never the tool itself. Is there really a difference between a human flooding the market using AI and a human floo…

> Is there really a difference between a human flooding the market using AI and a human flooding the market using a printing press?

Yes. A printing press only floods the market with copies. An AI floods the market with new derivative works.

A human producing a single creative work and then flooding the market with copies leaves lots of room for other humans to produce their own novel work. An AI flooding the market with new derivative works leaves no such room.

I work with DNNs a lot professionally and remain a proponent of the technology, but what OpenAI et al are doing is highly exploitative and scummy. It’s also damaging their social licence and may end up setting the field back.

Re: On being listed as an artist whose work was used to train Midjourney

#915

Earlier quoted context omitted.

> The model isn't an "intelligence" because it's not making a choice […] There’s no evidence any of us are making choices either. I don’t think LLMs are intelligences in the way most people mean the word (I think that would be vastly overestimating their abilities) but appealing to choice when it’s so poorly understood — as poorly understood as intelligence, even — is not useful.

>There’s no evidence any of us are making choices either. There is very clear evidence that the humans that operate the computers make the choice on which data the model gets trained on. The discussion is about those choices.

My point is that there’s no scientific way to demonstrate that one has made a choice vs that outcome being determined (perhaps probabilistically) or being random. We can’t “go back” and do it a different way to demonstrate we’ve actually chosen.

We don’t understand what it means to be intelligent in the way people generally mean, and we don’t understand (scientifically) what it means to choose. So we can’t use choice to usefully define intelligence.

Re: On being listed as an artist whose work was used to train Midjourney

#916
post #895

Earlier quoted context omitted.

This is a load of bullshit and I sincerely hope you know that as well as I do. As a thought experiment, let's say I pirate enough ebooks to stock a virtual library roughly equivalent in scope to a large metropolitan library system, then put up a website where you can download these books for free. I make money on the ads I run on this website, etc. This is theft, but as "compensation" I put some percentage of my reve…

I sincerely think you (in the thought experiment) sounds like an incredible hero. Libraries are awesome and great for reducing inequality, using ads to support that cause and also funneling cash to UBI initiatives? Even better

UBI is great! But taking away artists' livelihoods and justifying it with an uncertain promise of UBI years or (far more likely) decades in the future is (hypothetically) a particularly malignant example of putting the cart before the horse.

My observation of the vast majority of people who think intellectual property law is moral anathema to some idealized notion of "freedom" is that they either know nothing about making art themselves, or at the very least they are dilettantes who don't rely on it for any substantial portion of their livelihood. And, as it turns out, the latter group very rarely produce anything of note. Why is that? Because making interesting art is hard—just making art the average person would be willing to spend more than a second or two thinking about is hard enough, let alone art they'd be willing to pay for. It's hard, and it takes time and incredible effort and focus.

By taking (in whatever sense) an artist's work for less than it's worth, you're damaging their ability to extract value from their time investment in art, and so forcing them to choose between their artistic endeavors and other profitless aspects of their lives (e.g. their families). Even if you think that's your right (which I absolutely, categorically disagree with), your theft will have a very real effect of stunting young artists' development and reducing their output, and therefore cultural impact, in aggregate.

The sort of (frankly) idiots who think AI art is "giving power back to the people", or whatever, probably don't care about that because they don't understand art, or its impact on people and culture, at all. All I can say to them is that they have an incredibly juvenile outlook on the whole subject, and that instead of whining about artists gatekeeping art they should spend $10 on a pad of paper and a box of pencils, or $0 on a Google Doc, and go make some fucking art for themselves. Nobody's stopping them, and if they put in enough work maybe they'll start to understand some things.

Or, if you think UBI fixes everything, go ahead and start sending out checks (with a robust guarantee that they will keep coming). I'm sure that will change the tenor of the conversation.

Re: On being listed as an artist whose work was used to train Midjourney

#917

Earlier quoted context omitted.

> "Under what circumstances would AI art be acceptable then?" Easy! Under the circumstances where the artists whose art was used to train the model explicitly consented to that (without coercion), licensed their art for such use, and were fairly compensated for that. Plenty of artists would gladly paint for AI to learn from — just like stock photographers, or clip art designers, or music sample makers. Somehow, "payi…

This only makes it harder for open, non-profit models to compete with large corporations in making AI models. Examples: Adobe Firefly. Getty's AI. Probably Disney's internal AI. Are those acceptable, then? I'd say Stable Diffusion is more acceptable than any of them. Information wants to be free, and you shouldn't need consent to reuse information. Furthermore, this isn't coercion -- I could just as easily make the c…

>This only makes it harder for open, non-profit models to compete with large corporations in making AI models.

No. This only forces open, non-profit models to think a bit harder about incentives to provide artists for contributing their work into the training set.

What you exhibit here is a failure of imagination.

Re: On being listed as an artist whose work was used to train Midjourney

#918
post #783

Earlier quoted context omitted.

> So does a human brain. A major difference between a human training himself by looking at art, and a computer doing it, is that the human ends up working for himself, the computer is owned by some billionaire. One enhances the creative potential of humanity as a whole, the other consolidates it in the hands of the already-powerful. Another major difference is that a human can't use that knowledge to mass-produce tha…

> I see no reason for why a neural network's output should be protected by copyright. Only humans can produce copyrightable works, and a prompt is not a sufficient creative input. Your brain is a neural network, just FYI.

A privileged one, compared to one made of silicon, in the eyes of the law.

Re: On being listed as an artist whose work was used to train Midjourney

#919

Earlier quoted context omitted.

It'd be a good point if it wasn't for the fact that search engines didn't exist until google, because of technology, and that courts didn't need to consider the issue until then. So where does your point get us? We are here now.

"search engines didn't exist until google" - you might want to, uh, google that

The post you are replying to is poorly-worded.

I think the post is making reference to the "thing" which made Google stand out in a sea of existing search engines.

"Using the idea of relevancy, they built Google in a way that - in comparison to other search engines at the time - was simply better at connecting users with more pertinent results.

"A query typed in Google provided more utility and relevancy than did Excite, Yahoo and other search engines.

"To put it simply, Google had a superior product.

https://www.searchenginepeople.com/blog/125-why-google-won.h....

Re: On being listed as an artist whose work was used to train Midjourney

#920

Earlier quoted context omitted.

I don't buy the argument that the models are not sufficiently transformative. Nobody would look at a weight table and confuse it for the original work, and it is different in basically every way. If there is a case to be made, I think it has to be around the original use of the works, the transcription process. Not the weights, or the output

No one would look at a veracrypt archive containing a copyrighted work and confuse it either. They look very different to the original files but both the encrypted file and the learning model's weights allow one to reproduce the copyrighted work.

Not falling back to the argument that the output is infringing, not the archive.

If the archive can't produce the original work, it's not infringing.

If you printed the binary of Harry Potter and sold it as a painting, that would be fair use. It doesn't matter if the data is encoded in it, if it is not used for extraction.

Think of Andy warhol's Campbell Soup. Nobody is going to confuse the art for a can of soup and try to eat it. That's not being sold as a label for other soups. However, the original Campbell Soup data is absolutely encoded there.

Post reply on HN