Live data from Hacker News

OpenAI fails to deliver opt-out system for photographers

petapixel.com

101–110 of 166 posts

Re: OpenAI fails to deliver opt-out system for photographers

#101
post #19

"By continuing, you agree that using any content from this site in training Generative AI grants the site-owner a perpetual, irrevocable, and royalty-free license to use and re-license any and all output created by that Generative AI system, including but not limited to derivative works based on that output." Just just a GPL-esque idea I've been musing lately [0], I'd appreciate any feedback from actual IP lawyers. T…

That's cute, I'm going to put that at the start of any creative work I make so that anyone who sees it owes me a license to everything they ever made afterward because a nugget of their life experience legally belongs to me now and all their creative works are now tainted by that.

Re: OpenAI fails to deliver opt-out system for photographers

#102
post #25

Aren't lawsuits the proper way to address this? Seems like there's an argument that model weights are a derivative work of the training data, at least if the model is capable of producing output that would be ruled to be such a derivative work given minimal prompting. Although it may not work with photography since the model might just almost exclusively learn how the object of the photo looks in general and how phot…

I think that argument falls down though, because a derivative work is an expressive work in its own right, and model weights aren't.

It would seem more coherent to argue that a model output could be a derivative work, though it would need to include a significant portion of some given source. But even then, since the copyright office's position is that they're not copyrightable, I'm not sure they could qualify.

Re: OpenAI fails to deliver opt-out system for photographers

#103
post #7

Earlier quoted context omitted.

Insofar as data for diffusion / image / video models are concerned, the rise of synthetic data and data efficiency will mean that none of this really matters anyway. We were just in the bootstrapping phase. You can bolt on new functional modules and train them with very limited data you acquire from Unreal Engine or in the field.

I don’t entirely agree. For example, it’s a very popular scheme on Etsy right now to use LLMs to generate posters in the style of popular artists. Any artist should be able to say hey I don’t want my works to be part of your training set to power derivative generations. And I think it should even apply retroactively so that they have to retrain their models that are already generating works from training data consume…

Style isnt protected?

Re: OpenAI fails to deliver opt-out system for photographers

#104

Good. Everyone gets big mad when someone with money acts like Aaron Swartz did. The only bad thing about OpenAI is that they're not actually open sourcing or open accessing their stuff. Mistral or Llama "training on pirated material" is literally a feature, not a bug and the tears from all the artists and others who get mad are delicious. These same artists would profess literal radical marxism but become capitalist…

It is comical to me how fast the anti-RIAA internet turned into a bunch of copyright maximalists who expect organizations like the RIAA to protect them in some way. In actuality, if someone manages to weaponize copyright against AI, it will only successfully be used by massive rights holders to extract payouts from AI companies and none of the money will be given to any of the creatives, and creatives will naturally still not be very happy about it. Spotify 1.0 is right holders streaming your content and paying you fractions of pennies for it, "Spotify 2.0" will be them licensing your content to AI companies and paying you a fraction of a fraction of a penny just once.

Re: OpenAI fails to deliver opt-out system for photographers

#105
post #94

I think it’s safe to assume anything Sam A says is an outright lie by now.

It's depressing that this understanding hasn't been the status quo for years now. It's not like this is his first gig, it's been publicly verifiable what kind of person he is for ages, long before GPT became famous. You don't need to be part of some insider Silicon Valley cabal to find out.

Can you back that up with anything? I’ve gotten this as a vague sense, but it seems hard to find much actual background about how he manages to continuously fail upward.

Re: OpenAI fails to deliver opt-out system for photographers

#106

No way OpenAI will ever “good citizen” this. Tools to opt out of training sets will only come if they are legally compelled. Governments will have to make respecting some sort of training preference header on public content mandatory I think. The fact that photographers have to independently submit each piece of work they wanted excluded along with detailed descriptions just shows how much they DONT want anyone exclu…

Reminds me of the time when p2p music sharing became popular and the record companies had to submit every song they did not want to get shared along with an explanation to every person who had Napster installed.

Or was it that the record companies got to sue individuals for astronomic amounts of made up damages for every song potentially shared?

Which one was it?

Re: OpenAI fails to deliver opt-out system for photographers

#107

Earlier quoted context omitted.

Hard to say what motivates them, from the outside looking in. There have been signs of cultlike behavior before, such as the way the rank and file instantly lined up behind Altman when he was fired. You don't see that at Boeing or Microsoft. Obviously it's a highly-commercial endeavor, which is why they are trying so hard to back away from the whole non-profit concept. But that's largely orthogonal to the question of…

> Especially given that only HN'ers are 100% certain that training a model is infringement. In the real world, this is not a settled question. Why worry about obeying laws that don't even exist yet? This is exactly why people are against it. Your argument is that there is no definitive law. Therefore the creators of the data you scrape to train, and their wishes are irrelevant. If the motivation was to help humanity,…

Your argument is that there is no definitive law. Therefore the creators of the data you scrape to train, and their wishes are irrelevant.

Correct, that is the position of the law. Here in America, we don't take the position, held in many other countries, that everything not explicitly permitted is forbidden. This is a good thing.

If the motivation was to help humanity, they’d think twice about stepping on the toes of the humanity they want to save

Whether it is permissible to train models with copyrighted content is up to the courts and Congress, not us. Until then, no one's toes are being stepped on. Everybody whose work was used to train the models still holds the same rights to that work that they held before.

Re: OpenAI fails to deliver opt-out system for photographers

#108
post #40

Earlier quoted context omitted.

Should any artist be able to tell another artist: hey don't copy my work when you're learning, I don't want competition? It seems like they are deeply upset someone has figured out a way for a machine to do what artists have been doing since time immemorial.

Llms are not humans and shouldn’t be anthropomorphized as a strategy to get around copyright infringement.

And if they do get anthropomorphized... then the people in charge of that company need to be charged with the heinous crime of enslaving children.

Re: OpenAI fails to deliver opt-out system for photographers

#109

Earlier quoted context omitted.

> An AI that has enough sense of self-awareness to not hallucinate It's not entirely clear that this is meaningful. Humans engage in confabulation, too.

Humans engage in confabulation but they’re mostly aware of it. In some mental disorders they may not be aware; though statistically that is not too significant and no, we normally don’t confabulate as much as the current crop of AI aka LLMs. As a tool LLMs are fantastic and am glad to look at them as solely as powerful tools. AGI is not here yet and maybe that’s a good thing. Who would want some kind of artificial in…

https://www.medicalnewstoday.com/articles/confabulation#vs-l...

> Confabulators are usually unaware they are providing false information. They often display genuine surprise or confusion when evidence of facts contradicts their statements.

This is similar to LLMs actually. But it also seems like various "System 2" things like chain of thought could compensate for this issue in the LLM (and that possibly that is similar to how the brain works).

Re: OpenAI fails to deliver opt-out system for photographers

#110
post #74

Earlier quoted context omitted.

There is an argument to be made that ChatGPT mildly rewording/misquoting info directly from my blog is copying.

I think to make that argument you would need evidence that someone prompted ChatGPT to reword/misquote info directly from your blog, at which point the argument would be that that person is rewording/misquoting info directly from your blog, not ChatGPT.

I don't think so: The user is merely making a request for copyrighted material, which is not itself infringing, even if their request was extremely specific and their intent was obvious.

OpenAI would be the company actually committing the infringement and providing the copy in order to satisfy the request.

If the law suddenly worked the other way around, companies would no longer be able to prosecute people for hosting pirated content online, because the responsibility would lie with the users choosing to initiate the download.

Post reply on HN