Live data from Hacker News

On being listed as an artist whose work was used to train Midjourney

catandgirl.com

701–710 of 957 posts

Re: On being listed as an artist whose work was used to train Midjourney

#701
post #607

Earlier quoted context omitted.

Like how Google has parsed webpages of content to develop their page rank algorithm for searching on the web? I'm assuming it does.

No, because crawling the web, ingesting copyrighted content, and ranking them is not a derivative work of that content.

If crawling the web, ingesting copyrighted content, and ranking them is not a derivative work for that content, then using them to change the values of a mathematical expression should also exempt the expression from being a derivative work.

> Does copyright law say you can ingest copyrighted work at very large scale and sell derivates of those works

In that case the OP should have never posed this irrelevant question because access to the expression isn't giving access to a derivative work.

Re: On being listed as an artist whose work was used to train Midjourney

#702

Earlier quoted context omitted.

Thats pretty silly. you can just put a gatekeeper to prevent it from spitting out anything too similar, or prevent a user from forcing it to. it is not an intractable or pervasive problem. it is a fringe case that rarely occurs, and only with a lot of user prompting.

the weights themselves are still a derivative work even if they post-filter legal discovery could almost certainly compel the LLM host to provide access to the output of the weights themselves without the "gatekeeper" present

I don't buy the argument that the models are not sufficiently transformative. Nobody would look at a weight table and confuse it for the original work, and it is different in basically every way.

If there is a case to be made, I think it has to be around the original use of the works, the transcription process. Not the weights, or the output

Re: On being listed as an artist whose work was used to train Midjourney

#703

Earlier quoted context omitted.

No, I bought the books used for 25 cents at a local booksale, and the authors did not benefit from my secondary market transaction.

>> the authors did not benefit from my secondary market transaction. But they did. The presence of a secondary market for used books increased the value of some new books. People buy them knowing that they might one day recoup some costs by selling them. Would people pay more, or less, for a new car if they were told they could never sell or trade it away as a used car?

Gee I don't know, but I'm glad that digital goods do not incur the same material costs as a car. "You wouldn't download a car", we've come full circle.

Re: On being listed as an artist whose work was used to train Midjourney

#704

Earlier quoted context omitted.

No, I bought the books used for 25 cents at a local booksale, and the authors did not benefit from my secondary market transaction.

You got that via the legal "first sale doctrine" which has been killed for digital works.

It's a tough issue to correlate to physical goods, especially when you realize that people sometimes donate books.

Re: On being listed as an artist whose work was used to train Midjourney

#705
post #695

Earlier quoted context omitted.

I tend to fall more on the "training should be fair use" side than most, but your comment seems to be missing the point. Nobody is arguing that models are violating copyright or social norms around credit simply because they consume this information. Nobody ever argued/argues that the traditional text generation in markov models on your phone's keyboard runs afoul of these issues. The argument being made is that thes…

A company hires an artist. That artist has observed a ton of other artists' work over the years. The company instructs that artist to draw, "X but in the style of Y", where Y is some copyrighted artwork. The company then prints the result and puts it on their packaging. A company builds an AI tool. That AI tool is trained on a ton of artists' work over the years. The company opens up the AI tool and asks it to draw,…

Okay, but then that's an an argument subject to the critiques made upthread that you were initially trying to dismiss? You can't claim that AI doesn't need to worry about citing influences because it's just doing a thing humans wouldn't cite influences for, then proceed to cite an example where you would very much be expected to cite your influences, and AI wouldn't, as evidence.

Re: On being listed as an artist whose work was used to train Midjourney

#706

Earlier quoted context omitted.

No, I bought the books used for 25 cents at a local booksale, and the authors did not benefit from my secondary market transaction.

You got that via the legal "first sale doctrine" which has been killed for digital works.

[deleted]

Re: On being listed as an artist whose work was used to train Midjourney

#707

Earlier quoted context omitted.

No, I bought the books used for 25 cents at a local booksale, and the authors did not benefit from my secondary market transaction.

You got that via the legal "first sale doctrine" which has been killed for digital works.

"In 2012, the Court of Justice of the European Union (ECJ) held in UsedSoft GmbH v. Oracle International Corp that the first sale doctrine applies to used copies of [intangible goods] downloaded over the Internet and sold in the European Union." [0]

Arguably the U.S. courts are in the wrong here. We can only hope first sale doctrine is extended to digital goods in the U.S. in the future, as it has been in the EU for over a decade.

[0] https://scholarlycommons.law.northwestern.edu/cgi/viewconten...

Re: On being listed as an artist whose work was used to train Midjourney

#708

Earlier quoted context omitted.

Oh, right. It just reads a million books in a couple of days, removes all the source information, mix and match it the way it sees fit and sells this output $10/month to anyone comes with a credit card. It's the same thing with GitHub's copilot. A book publisher would seize everything I have, and shot me at a back alley if I do 0.0001% of this.

Yeah, fair use implicitly uses the constraints of typical human lifetime and ability to moderate how much damage is done to publishers with it. That wasn’t an issue before recently, as humans were the only ones who could create output based off fair use laws.

> Yeah, fair use implicitly uses the constraints of typical human lifetime and ability

Authors Guild, Inc. v. Google, Inc. strongly disagrees with you on that (the "Google Books case").

Re: On being listed as an artist whose work was used to train Midjourney

#709

Earlier quoted context omitted.

The goal of my post was not to answer what differentiates google search with LLMs and other generative models, it was to respond to the original post above: > Does copyright law say you can ingest copyrighted work at very large scale and sell derivates of those works to gain massive profit / massive market capitalizations The reasons as to why I don't think training on copyrighted data are stated in my other comments…

Googles search engine is not selling derivative works. If you search for a Disney movie on Google search, it does not try to sell you a film derived from the movie.

They sell you ad space on full Disney movies (re)uploaded by random people who are not affiliated with Disney though: https://www.google.com/search?q=finding+nemo+full+movie

I can also get Disney coloring book pages directly from Google's cache on Google images: https://www.google.com/search?q=disney+princess+coloring+boo...

Authors Guild, Inc. v. Google, Inc. determined that Google's wholesale scanning and uploading of books is allowed under the first sale doctrine because the University of Michigan library they borrowed the books to scan from paid for them (or a donor paid for them, at some point). Here's a book of bedtime stories available in its entirety: https://www.google.com/books/edition/Picnics_in_the_Wood_and...

Re: On being listed as an artist whose work was used to train Midjourney

#710
post #366
post #303

Earlier quoted context omitted.

>But we are allowed to use copyrighted content. We are not allowed to copy copyrighted content. We are allowed to view and consume it, to be influenced by it, and under many circumstances even outright copy it. It's important to consider in any legalistic argument over copyright that, unlike conventional property rights which are to some degree prehistoric, copyright is a recent legal construct that was developed for…

Copyright only makes any sense for goods with a high fixed cost of production and low to zero marginal cost. Any further use beyond solving that problem is pure rent seeking behavior Also, with computers being functional copyright has become a tool of social control; any function in a physical object can be taken away from you at a whim with no recourse so long as a computer can be inserted into the object. Absent a…

> Copyright only makes any sense for goods with a high fixed cost of production and low to zero marginal cost. Any further use beyond solving that problem is pure rent seeking behavior

100% agree. But even then it's not very good. Abolish copyright, severely limit patents, and leave trademarks as they are. The IP paradigm needs an overhaul.

Post reply on HN