"By continuing, you agree that using any content from this site in training Generative AI grants the site-owner a perpetual, irrevocable, and royalty-free license to use and re-license any and all output created by that Generative AI system, including but not limited to derivative works based on that output." Just just a GPL-esque idea I've been musing lately [0], I'd appreciate any feedback from actual IP lawyers. T…
OpenAI fails to deliver opt-out system for photographers
101–110 of 166 posts
Re: OpenAI fails to deliver opt-out system for photographers
#102Aren't lawsuits the proper way to address this? Seems like there's an argument that model weights are a derivative work of the training data, at least if the model is capable of producing output that would be ruled to be such a derivative work given minimal prompting. Although it may not work with photography since the model might just almost exclusively learn how the object of the photo looks in general and how phot…
It would seem more coherent to argue that a model output could be a derivative work, though it would need to include a significant portion of some given source. But even then, since the copyright office's position is that they're not copyrightable, I'm not sure they could qualify.
Re: OpenAI fails to deliver opt-out system for photographers
#103Earlier quoted context omitted.
Insofar as data for diffusion / image / video models are concerned, the rise of synthetic data and data efficiency will mean that none of this really matters anyway. We were just in the bootstrapping phase. You can bolt on new functional modules and train them with very limited data you acquire from Unreal Engine or in the field.
I don’t entirely agree. For example, it’s a very popular scheme on Etsy right now to use LLMs to generate posters in the style of popular artists. Any artist should be able to say hey I don’t want my works to be part of your training set to power derivative generations. And I think it should even apply retroactively so that they have to retrain their models that are already generating works from training data consume…
Re: OpenAI fails to deliver opt-out system for photographers
#104Good. Everyone gets big mad when someone with money acts like Aaron Swartz did. The only bad thing about OpenAI is that they're not actually open sourcing or open accessing their stuff. Mistral or Llama "training on pirated material" is literally a feature, not a bug and the tears from all the artists and others who get mad are delicious. These same artists would profess literal radical marxism but become capitalist…
Re: OpenAI fails to deliver opt-out system for photographers
#105I think it’s safe to assume anything Sam A says is an outright lie by now.
It's depressing that this understanding hasn't been the status quo for years now. It's not like this is his first gig, it's been publicly verifiable what kind of person he is for ages, long before GPT became famous. You don't need to be part of some insider Silicon Valley cabal to find out.
Re: OpenAI fails to deliver opt-out system for photographers
#106No way OpenAI will ever “good citizen” this. Tools to opt out of training sets will only come if they are legally compelled. Governments will have to make respecting some sort of training preference header on public content mandatory I think. The fact that photographers have to independently submit each piece of work they wanted excluded along with detailed descriptions just shows how much they DONT want anyone exclu…
Or was it that the record companies got to sue individuals for astronomic amounts of made up damages for every song potentially shared?
Which one was it?
Re: OpenAI fails to deliver opt-out system for photographers
#107Earlier quoted context omitted.
Hard to say what motivates them, from the outside looking in. There have been signs of cultlike behavior before, such as the way the rank and file instantly lined up behind Altman when he was fired. You don't see that at Boeing or Microsoft. Obviously it's a highly-commercial endeavor, which is why they are trying so hard to back away from the whole non-profit concept. But that's largely orthogonal to the question of…
> Especially given that only HN'ers are 100% certain that training a model is infringement. In the real world, this is not a settled question. Why worry about obeying laws that don't even exist yet? This is exactly why people are against it. Your argument is that there is no definitive law. Therefore the creators of the data you scrape to train, and their wishes are irrelevant. If the motivation was to help humanity,…
Correct, that is the position of the law. Here in America, we don't take the position, held in many other countries, that everything not explicitly permitted is forbidden. This is a good thing.
If the motivation was to help humanity, they’d think twice about stepping on the toes of the humanity they want to save
Whether it is permissible to train models with copyrighted content is up to the courts and Congress, not us. Until then, no one's toes are being stepped on. Everybody whose work was used to train the models still holds the same rights to that work that they held before.
Re: OpenAI fails to deliver opt-out system for photographers
#108Earlier quoted context omitted.
Should any artist be able to tell another artist: hey don't copy my work when you're learning, I don't want competition? It seems like they are deeply upset someone has figured out a way for a machine to do what artists have been doing since time immemorial.
Llms are not humans and shouldn’t be anthropomorphized as a strategy to get around copyright infringement.
Re: OpenAI fails to deliver opt-out system for photographers
#109Earlier quoted context omitted.
> An AI that has enough sense of self-awareness to not hallucinate It's not entirely clear that this is meaningful. Humans engage in confabulation, too.
Humans engage in confabulation but they’re mostly aware of it. In some mental disorders they may not be aware; though statistically that is not too significant and no, we normally don’t confabulate as much as the current crop of AI aka LLMs. As a tool LLMs are fantastic and am glad to look at them as solely as powerful tools. AGI is not here yet and maybe that’s a good thing. Who would want some kind of artificial in…
> Confabulators are usually unaware they are providing false information. They often display genuine surprise or confusion when evidence of facts contradicts their statements.
This is similar to LLMs actually. But it also seems like various "System 2" things like chain of thought could compensate for this issue in the LLM (and that possibly that is similar to how the brain works).
Re: OpenAI fails to deliver opt-out system for photographers
#110Earlier quoted context omitted.
There is an argument to be made that ChatGPT mildly rewording/misquoting info directly from my blog is copying.
I think to make that argument you would need evidence that someone prompted ChatGPT to reword/misquote info directly from your blog, at which point the argument would be that that person is rewording/misquoting info directly from your blog, not ChatGPT.
OpenAI would be the company actually committing the infringement and providing the copy in order to satisfy the request.
If the law suddenly worked the other way around, companies would no longer be able to prosecute people for hosting pirated content online, because the responsibility would lie with the users choosing to initiate the download.