Earlier quoted context omitted.
> you put that information on the internet, you have no expectation to privacy That was the expectation until Zuckerberg came along. People were putting tons of stuff on BBSs, intranets, and other stuff 20+ years beforehand because people had the idea of "consent" back then existed: using someones information for a purpose they didn't intend was (and still is) wrong.
This seems like a fantasy. Anything could end up in 4chan with a photoshoped penis on it since a long time before Zuck came along.
Can I remove my personal data from GenAI training datasets?
111–120 of 123 posts
Re: Can I remove my personal data from GenAI training datasets?
#112Earlier quoted context omitted.
That could be besides the point as for the training the machines will create a copy in their memory - or is that no longer the case?
I don't believe transient copies are the real issue here. Having a temporary copy stored locally in memory is more of a technical nuance than anything. That copy remains isolated and is not being distributed or shared. In contrast, using something like BitTorrent actively spreads copies across the internet. There are also cases where loading copyrighted content is unavoidable in order to even view the license terms.…
Re: Can I remove my personal data from GenAI training datasets?
#113Earlier quoted context omitted.
> The fundamental problem here is about consent. What would a reasonable person expect? 10 years ago certainly no one expected this. I just don't think that's the right angle. I think if you're in the public sphere, the way that you anticipated your creations to be used just shouldn't matter at all. If you've given something to the public, the public should not be limited to whatever you thought they'd use it for. An…
Why does everything have to be binary these days? We aren't machines. The world is messy and so are we humans. Removing nuance from conversations and ideas does not help anything. The scenario you laid out, continued, leaves no person with privacy anywhere. I specifically don't want to explain because I specifically want more people thinking and thinking hard. To specifically attempt to break from their way of thinki…
If we have to lose privacy and control in order to keep our participatory culture, I'm sorry, I'll take culture. If it's a multidimensional topic, I'm for tightening privacy in private spaces and loosening privacy in public spaces. If there's only one axis, I'll push the lever towards loose.
Re: Can I remove my personal data from GenAI training datasets?
#114Earlier quoted context omitted.
Why does everything have to be binary these days? We aren't machines. The world is messy and so are we humans. Removing nuance from conversations and ideas does not help anything. The scenario you laid out, continued, leaves no person with privacy anywhere. I specifically don't want to explain because I specifically want more people thinking and thinking hard. To specifically attempt to break from their way of thinki…
Personally speaking: all the art I consume is remixes. Almost all the books I read are fanfics. If the repurposing of cultural content without the consent of the author becomes illegal, my cultural universe will vanish overnight. As such, I think I'm just on the other side of that fight. If we have to lose privacy and control in order to keep our participatory culture, I'm sorry, I'll take culture. If it's a multidim…
You'd be surprised because these industries have figured these problems out. What is allow without license, what needs royalties, what needs permission, and what can't be done. This was a huge conversation in the 90's and you can find many discussions about hip hop (around sampling) and copyright law.
Your culture is not in danger, even if you don't know its history. I'd suggest learning it though, because that's the best way to ensure your culture stays out of danger (a culture I am also part of fwiw).
Re: Can I remove my personal data from GenAI training datasets?
#115Earlier quoted context omitted.
Did you read what I said? Anything posted on the internet is public. Just because you think it's private does not make it so. To that end, you should consider any file on your computer to be public as well if it's connected to the internet, as there are countless ways that data could be exfiltrated.
That may well be true, but it's not what the article is about. The article is about individuals removing personal data from GenAI products. GenAI companies in many places have a legal responsibility to facilitate this. Your concern about all internet-connected data being "public" is only a concern if you're dealing with malicious actors, which is not what's at hand here.
There's no difference between an AI being trained on public data, or a human being trained on public data. Likewise there should be no expectation to "unsee" something someone willingly posted in public.
Re: Can I remove my personal data from GenAI training datasets?
#116Earlier quoted context omitted.
Personally speaking: all the art I consume is remixes. Almost all the books I read are fanfics. If the repurposing of cultural content without the consent of the author becomes illegal, my cultural universe will vanish overnight. As such, I think I'm just on the other side of that fight. If we have to lose privacy and control in order to keep our participatory culture, I'm sorry, I'll take culture. If it's a multidim…
> Personally speaking: all the art I consume is remixes. Almost all the books I read are fanfics. If the repurposing of cultural content without the consent of the author becomes illegal, my cultural universe will vanish overnight. You'd be surprised because these industries have figured these problems out. What is allow without license, what needs royalties, what needs permission, and what can't be done. This was a…
Re: Can I remove my personal data from GenAI training datasets?
#117Earlier quoted context omitted.
>For most companies “deleting the model” is equivalent to dissolving the company I'm okay with that. It's sad that you're not. If you're willing to start a business on such shady foundations, there's a really good chance your business will continue to make shady decisions in the future. It's better to find and remove the cancer early
I just disagree that it's shady. If you put your personal info into public circulation I think it's absolutely morally fine to have a model train on it.
Re: Can I remove my personal data from GenAI training datasets?
#118Earlier quoted context omitted.
> Personally speaking: all the art I consume is remixes. Almost all the books I read are fanfics. If the repurposing of cultural content without the consent of the author becomes illegal, my cultural universe will vanish overnight. You'd be surprised because these industries have figured these problems out. What is allow without license, what needs royalties, what needs permission, and what can't be done. This was a…
To the best of my knowledge, all claims that fanfiction is in the clear, that it's okay to do derivative works if you aren't charging money, and so on, are all just fandom folklore. If any author wanted to legally shut down fanfic, to my knowledge they could. And I'm against that - hell, in my opinion fanfiction writers should be able to sell their work even without authorial consent. Look at the Touhou scene, look h…
Yeah, that's to the best of my knowledge true but IANAL. I'm with you about the fanfic scenario too. I'm also very open about sampling. But I think these issues are different than the data that we're talking about in this thread. If data can be recovered (an in some ways it is, others it isn't) then that's not really derivative. Derivative also needs some distance and not be too close. And importantly, these data are being used to create a product that is being sold. Where the processing of the data is the thing of value.
My point is that the environment has changed and there's a lot of gray area here. Turning this into a binary distinction is unhelpful. There's new nuances here and there were new nuances when sampling became popular in hip hop. We need to have open and honest discussions about these things, and I think a lot of discussions we have or observe are rather dismissive of these nuances (coming from both sides of the debate fwiw). I'm obviously very open to using data but we must also be aware of our data privacy, how it can be used, our social contracts, and what a reasonable level of a priori consent is. If we overly simplify these conversations then they aren't actually discussions. My points are to this, that there is gray and that there are very clear cases where you do not have unlimited access and usage to works that are publicly available. Private ownership is the root of capitalism afterall, and so it should be rather unsurprising that we have many laws and social contracts over ownership and the extent of what one may do with things they did not create. There certainly is a lot of anger and frustration in these conversations and I don't expect an artist making their living off of their art to understand all these nuances nor am I surprised that they are upset and possibly afraid. This is new territory and pretending it isn't is just as obtuse as calling generators fuzzy copy-paste machines.
But I want to be clear that we can have both goals. We can protect data rights, privacy, fairness AND have this sampling and creativity culture, for lack of better words. We just need to be careful, nuanced, and thoughtful to determine how to do this though. We won't be perfect and won't make everyone happy, but we can maximize social agreement conditioned under fairness and privacy. I just want to ensure that we are not approaching this conversation as that there are clear answers in what can be done with data and what can't be. Hell, we don't even have that answer for music, sampling, or fan fiction. We have answers as to what laws say, but even as you point out, that's ambiguous in many cases without even considering that the environment is not only changing, but changing rapidly. I think we all understand that there is a difference between using the Akira slide compared to the "Ice Ice Baby v Under Pressure" scenario. No one has the answers, and that's why we need to talk. And unfortunately "edge cases" are the norm in topics like these.
Note: I am an ML researcher. I use publicly available data to train models that have images of people, their art, their animals, their property, and such that I'm sure many do not know exists in these datasets. Similarly I do not even know all the data within some of these datasets. But I can still recognize that there is a gray area that exists here and personally I see it as my ethical duty to ensure we have these discussions in an open and honest way to determine the limits of what I should and shouldn't be able to do. It isn't up to me, it is up to our society to create a social contract.
Re: Can I remove my personal data from GenAI training datasets?
#119Earlier quoted context omitted.
There's a difference between being in public and being available for public use. Books in a book shop are "in public", but only those out of copyright are available for public use. Similarly, blog posts are in public, but not being given away for anyone to resell or include in their product. For an image/photography example, I may have my photo taken at an event and sign a release or agree to the event's T&Cs that sa…
> There's a difference between being in public and being available for public use. Books in a book shop are "in public", but only those out of copyright are available for public use Copying books can be fair use if it is sufficiently transformative, copyright doesn't block all use of copyrighted material. A major case on this was around Google building a search index for books: https://en.wikipedia.org/wiki/Authors_G…
Re: Can I remove my personal data from GenAI training datasets?
#120Earlier quoted context omitted.
Machines need to be prompted with the exact prefix to be able to retrieve any copyrighted fragment, and that doesn't work most of the times. So the intent for copyright circumvention is in the prompt. Temperature settings also matter.
That could be besides the point as for the training the machines will create a copy in their memory - or is that no longer the case?