Earlier quoted context omitted.
> For transparency, I am an advocate for human made art, If you believe AI tooling is an artform then you categorically are advocating against human made art as far as I am concerned.
This is just gatekeeping. Art is not better because it was made by hand as opposed to with technology. If I use a generative model to make art then I’m an artist.
On being listed as an artist whose work was used to train Midjourney
681–690 of 957 posts
Re: On being listed as an artist whose work was used to train Midjourney
#682Earlier quoted context omitted.
> Does that mean that models that can not produce copies of X length ARE fair use? not necessarily "sufficient but not necessary" I believe is the term
Sure, but then you need an additional line of rationale and logic to cover those other cases.
Re: On being listed as an artist whose work was used to train Midjourney
#683Earlier quoted context omitted.
> I firmly believe that training models qualifies as fair use There's a hell lot of money to be made from this belief so of course the HN crowd will hold it. Some of us here who have been around the copyright hustle for a little longer laugh at this bitterly and pray that the courts and/or Doctorow's activism saves us. But there's so much money to be made from automatized plagiarism and the forces against are so weak…
Generative models are just a tool. Artists are mad because this tool empowers other people, who they view as less talented, to make art too. The camera and 1-hour film developing didn’t destroy oil paintings, it just enabled more people to have control over what was on their walls.
Re: On being listed as an artist whose work was used to train Midjourney
#684Earlier quoted context omitted.
So, zero. You yourself: zero. You completely ignored the premise of the question.
You should read the response more carefully. Generative models are just tools. If I use one to write a story it’s no less a story that I wrote than if I’d chiseled it into a Persian mountainside.
Re: On being listed as an artist whose work was used to train Midjourney
#685Earlier quoted context omitted.
I firmly believe that training models qualifies as fair use. I think it falls under research, and is used to push the scientific community forward. I also firmly believe that commercializing models built on top of copyrighted works (which all works start off as) does not qualify as fair use (or at least shouldn't) and that commercializing models build on copyrighted material is nothing more than license laundering. C…
> The reason that AI models are generating content similar to other people's work is because those models were explicitly trained to do that. Ah, just like humans who train against the output of other humans. AI models are not fundamentally different in kind in this regard, only scope, and even that isn't perfectly obvious to me a priori.
Have you ever tried to "train a human"?
They don't work that way, not unless your "training" involves so weird torture stuff you probably shouldn't be boasting about.
Maybe try ask some teachers (of both adults and children) how it works with people...
Re: On being listed as an artist whose work was used to train Midjourney
#686Earlier quoted context omitted.
Sure, but then you need an additional line of rationale and logic to cover those other cases.
I suspect that single case will catch all of them
it is a fringe case that rarely occurs, and only with a lot of user prompting.
Re: On being listed as an artist whose work was used to train Midjourney
#687Earlier quoted context omitted.
This is a load of bullshit and I sincerely hope you know that as well as I do. As a thought experiment, let's say I pirate enough ebooks to stock a virtual library roughly equivalent in scope to a large metropolitan library system, then put up a website where you can download these books for free. I make money on the ads I run on this website, etc. This is theft, but as "compensation" I put some percentage of my reve…
IMO that's a terrible thought experiment given the situation. LLMs do not store enough content or with enough accuracy to even close to a virtual library. Unlike, say, Google and the Wayback Machine, the former of which stores enough to show snippets from the pages it's presenting to you as search results (and they got sued for that in certain categories of result), and the latter is straight up an archive of all the…
I already addressed this: "the only difference between it and what OpenAI is doing is that OpenAI's product relies on a technical means of laundering intellectual property that seems tailor-made to dodge a body of existing copyright law designed by people who could not possibly have conceived of what modern genAI is capable of."
You are certainly welcome to disagree with what I've said, but you can't simply pretend I didn't say it.
> Furthermore, the "percentage" in question for OpenAI is "once we've paid off our investors, all of it goes to benefitting humanity one way or another" — the parent company is a not-for-profit.
A quick Googling suggests that OpenAI employees are not working for free—far from it, in fact. In this frame I don't particularly care whether the organization itself is nominally "non profit", because profit motives are obviously present all the same.
> Here's a different question for you: If a generative AI is trained only on out-of-copyright novels and open-licensed modern works, and then still deprives everyone of all publishing opportunities forever as in this thought experiment it's better and cheaper than any human novelist, is that any more, or any less fair on literally any person on the planet? The outcomes are the same.
They are certainly welcome to try! Given how profoundly incapable extant genAI systems are of generating novel (no pun intended) output, including but not limited to developing artistic styles of their own, I think it would be quite funny to see these companies try to outcompete human artists with AI generated slop 70+ years behind the curve of art and culture. As for modern "public domain"-ish content, if genAI companies actually decided to respect intellectual property rights, I expect those licenses would quickly be amended to prohibit use in AI training.
AI systems will probably get there eventually, though it's very difficult to predict when. However, that speculation does not justify theft today.
> I'm sure someone's already thought of making such a model, it's just a question of if they raised enough money to train such a model.
People are absolutely throwing money at genAI right now, so if nobody has thrown enough money at this particular idea to give it a fair shake then the obvious conclusion is that people who know genAI think it's a relatively bad one. I'm inclined to agree with them.
> You may have noticed from the version number that they're on versions 3 and 4. When version 2 came out in 2019, they said [...]
Why is this relevant? I'm not talking about AI safety or "X risk" or whatever—I'm talking about straightforward intellectual property theft, which OpenAI and their contemporaries are obviously very comfortable with. The models they sell to anybody willing to pay today could literally not exist without their training datasets.
Re: On being listed as an artist whose work was used to train Midjourney
#688Earlier quoted context omitted.
> This seems like a strange criticism to me On the contrary, I think it is very natural. The environment changed, and thus the deal. There's no clear way to negotiate the terms of the deal. It may be easy to say to just drop off the platform, but we've seen how difficult it can be to destroy a social media platform. Sure, Myspace and Google+ failed, but they didn't have the network base that Facebook and Twitter do w…
The author reminisces of a time that was favorable to them, and implies that back then somehow it was free to distribute and this was taken away. Which is interesting, because there were ads back then too. A lot. Banners. Toolbars. Google made money a shitton of money back then too. And there are amazing non-FAANG spaces today like the Fediverse where the author can distribute their work, and the number of users ther…
This sounds like another way to say that the environment changed and thus the deal did.
There's still a ton of ads. And it isn't like Meta is making less money. They're just below their all time high[0] and it's not like there was one of the largest global financial crises between these two periods or something. Meta is in the top 10 for market caps and just shy of getting into the trillion dollar club. I'm not sure what argument you're trying to make because I'm not seeing how Meta (or any FAANG) is struggling.
> And there are amazing non-FAANG spaces today like the Fediverse where the author can distribute their work, and the number of users there are easily comparable to the 'old internet' user counts.
Be realistic. We both know that 1) many of the artists are trying to distribute there, 2) that this doesn't prevent their work from being used in ML models since you can still download the material, 3) the audience is substantially smaller and Mastodon is not even close to replacing Twitter.
> platforms and powerful interests shape userbases
Also remember that platforms with large userbases shape users. There's a reason why the classic Silicon Valley Strategy is to run in the red and attempt to establish a (near) monopoly. Because once you have it, you have a lot of power and can make up for all the cash you burned. Or you now have a product to sell: users.
Re: On being listed as an artist whose work was used to train Midjourney
#689Earlier quoted context omitted.
I suspect that single case will catch all of them
Thats pretty silly. you can just put a gatekeeper to prevent it from spitting out anything too similar, or prevent a user from forcing it to. it is not an intractable or pervasive problem. it is a fringe case that rarely occurs, and only with a lot of user prompting.
legal discovery could almost certainly compel the LLM host to provide access to the output of the weights themselves without the "gatekeeper" present
Re: On being listed as an artist whose work was used to train Midjourney
#690Earlier quoted context omitted.
I firmly believe that training models qualifies as fair use. I think it falls under research, and is used to push the scientific community forward. I also firmly believe that commercializing models built on top of copyrighted works (which all works start off as) does not qualify as fair use (or at least shouldn't) and that commercializing models build on copyrighted material is nothing more than license laundering. C…
> I firmly believe that training models qualifies as fair use There's a hell lot of money to be made from this belief so of course the HN crowd will hold it. Some of us here who have been around the copyright hustle for a little longer laugh at this bitterly and pray that the courts and/or Doctorow's activism saves us. But there's so much money to be made from automatized plagiarism and the forces against are so weak…
That is a pretty unfair response when you're skipping the part about how commercializing should not be fair use.