There is simple way to fix this. Ban private large models trained on public data, require them to be public weights. If a company wants to train large private model, they can do it with their own data.
On being listed as an artist whose work was used to train Midjourney
221–230 of 957 posts
Re: On being listed as an artist whose work was used to train Midjourney
#222Earlier quoted context omitted.
The point is that they no longer have a choice not to enrich profit-seeking companies with their unpaid labor, because by posting to those sites, you're generating the content they place their ads beside. And then pointing out that at least with social media, they had a choice in the matter, while AI scrapers just take it, wherever they might have posted it.
going to take some HN karma hit for this but i just don't get it. the price of anything you sell (or do not wish to give away for free) is not how much it cost to build but how much someone else is willing to pay for it. If you think the work is so great. charge for it. let the market decide. i think twitter/x does this now somewhat. how is this not similar to thinking google owes me money because my tweet appeared i…
Re: On being listed as an artist whose work was used to train Midjourney
#223Earlier quoted context omitted.
>> If you think OpenAI is less valuable because it can't use copyrighted content, then it should give some of that value back to the content. But we are allowed to use copyrighted content. We are not allowed to copy copyrighted content. We are allowed to view and consume it, to be influenced by it, and under many circumstances even outright copy it. If one doesn't want anyone to see/consume or be influenced by one's…
I'm not as interested in making a technical/legal argument, as I'm just sharing my feelings on the topic (and eventually, what I think the law should be), but during training copies are made of copyrighted material, even if the model doesn't contain exact copies of work. Crawling, downloading, storing (temporarily) for training all involve making copies, and thus are subject to copyright law. Maybe those copies are f…
How would this compensation work? Let's say a portion of profits from LLMs that were trained on copyrighted work should be sent to the copyright holders.
How would we allocate which portion of the profits go to which creators? The only "fair" way here would be if we could trace how much a specific work influenced a specific output but this is currently impossible and will likely remain impossible for quite some time.
Re: On being listed as an artist whose work was used to train Midjourney
#224[flagged]
Re: On being listed as an artist whose work was used to train Midjourney
#225Earlier quoted context omitted.
Because we are humans and our capability of abusing those rights is limited. The scale and speed at which LLMs can abuse copyrighted work to threaten the livelihoods of the authors of those works is reason enough to consider it unethical.
I don’t think it is. What you describe is similar to any other industry disruption, and I don’t think those are unethical. I’d actually argue that preventing disruption is often (not always) unethical, because you artificially prolong an inefficient or inferior alternative.
You removed the value and authenticity that artist in 30 minutes, you applauded it, and defended that it should be the norm.
OK then, we can close down all entertainment business, and generate everything with AI, because it can mimic styles, clone sounds, animate things with gaussian splats, and so on.
Maybe we can hire coders to "code" films? Oh sorry. ChatGPT can do that too. So we need a keypad then, only the most wealthy can press. Press 1 for a movie, 2 for a new music album, 3 for a new book, and so on.
We need 10 buttons or so, as far as I can see. Maybe I can ask ChatGPT 4 to code one for me.
Re: On being listed as an artist whose work was used to train Midjourney
#226Earlier quoted context omitted.
I firmly believe that training models qualifies as fair use. I think it falls under research, and is used to push the scientific community forward. I also firmly believe that commercializing models built on top of copyrighted works (which all works start off as) does not qualify as fair use (or at least shouldn't) and that commercializing models build on copyrighted material is nothing more than license laundering. C…
If I read a lot of stories in a certain genre that I like, and I later write my own story, it’s almost by definition going to be a mish-mash of everything I like. Should I pay the authors of the books I read when I sell mine?
Re: On being listed as an artist whose work was used to train Midjourney
#227Earlier quoted context omitted.
going to take some HN karma hit for this but i just don't get it. the price of anything you sell (or do not wish to give away for free) is not how much it cost to build but how much someone else is willing to pay for it. If you think the work is so great. charge for it. let the market decide. i think twitter/x does this now somewhat. how is this not similar to thinking google owes me money because my tweet appeared i…
If they don't link to you, but instead take the content of the link and show it alongside ads, they do owe you money, at least in some jurisdictions.
which is really not dissimilar to what midjourney is giving you which is a result from an array of public(?) works if i understand (not an AI/ML expert)
Re: On being listed as an artist whose work was used to train Midjourney
#228Earlier quoted context omitted.
Fair use is a defence you can use when you infringe copyright [edit for clarity] or in other words the action you take would otherwise infringe. It's not fair use because you want it to be, and it's not at all legally clear if this defence is valid in the case of AI training. But it's not clear it isn't, either. This is basically what all the noise and PR money is about, currently, in hope that shaping the narrative…
Fair use is not copyright infringement! It’s a limitation placed on copyright to balance the interests of copyright holders with the public interest.
It's a bit of a quibble to differentiate between allowable infringement or exception, here, so i've edited original.
[note: 0xcde4c3db points out correctly that it does actually show up since 1976, but in a way the defers much of the definition back to case law, so same effect]
Re: On being listed as an artist whose work was used to train Midjourney
#229Earlier quoted context omitted.
> That's literally built into their corporate rules for how to take investment money, and when those rules were written they were criticised because people didn't think they'd ever grow enough for it to matter. How is OpenAI compensating the owners of IP they trained their models on? Or is that not what you mean? It's certainly how I read the part of the GP comment you quoted.
So far, looks like funding a UBI study. As the IP owners are approximately "everyone" in law, UBI is kinda the only way to compensate all the IP owners. https://openai.com/our-structure
So the researchers, shareholders, and leadership of OpenAI will be happy to give up being ridiculously wealthy so they can be only moderately wealthy, and everyone else gets a basic income?
I'm also just skeptical of UBI in general, I suppose - 'free' money tends to just inflate everything to account for it, and it still won't address scarcity issues for limited physical assets like land/property.
I've love to be wrong about both of these things.
Re: On being listed as an artist whose work was used to train Midjourney
#230Earlier quoted context omitted.
> Put it this way - you remove all the copyrighted, permission-less content from OpenAIs training, what value does OpenAI's products have? If you think OpenAI is less valuable because it can't use copyrighted content, then it should give some of that value back to the content Couldn't we say the same thing about search engines? What value would google have without content to search for? Is the conclusion we should ma…
Search engines don't replicate the content, they index and point to it. When search engines have been caught replicating content they have been sued or had to work out licenses.