Live data from Hacker News

Machine Unlearning in 2024

ai.stanford.edu

81–90 of 97 posts

Re: Machine Unlearning in 2024

#81
post #3

> However, RTBF wasn’t really proposed with machine learning in mind. In 2014, policymakers wouldn’t have predicted that deep learning will be a giant hodgepodge of data & compute Eh? Weren't deep learning and big data already things in 2014? Pretty sure everyone understood ML models would have a tough time and they still wanted RTBF.

I'm pretty sure that the policymakers did NOT understand ML models in 2014 - and still do NOT understand it today.

I also don't think that they care. They don't care that ML is a hodgepodge of data & compute, and they don't care how hard it is to remove data from a model.

They didn't care about the ease or difficulty of removing data from more traditional types of knowledge storage either - like search indexes, database backups and whatnot.

RTBF was not proposed with any specific technology in mind. What they had in mind, was to try and give individuals a tool, to keep their private information private. Like, if you have a private, unlisted phone number, and that number somehow ends up on the call-list of some pollster firm, you can force that firm to delete your number so that they can't call you anymore.

The idea is, that if your private phone number (or similar data) ends up being shared or sold without your consent - you can try to undo the damage.

In practice it might still be easier to get a new number, than to have your leaked one erased... but not all private data is exchangeable like that.

Re: Machine Unlearning in 2024

#82

I don't know — the post, reading the comments here, I am a little worried for the "sanity" of our AI that have been trained, untrained, retrained like a pawn in some kind of Cold War spy novel.

It's fine, the LLM AIs we have now are just fancy versions of autocorrect. They, and other LMs, guess at statistically probable words/datapoints, and because they don't understand context, you might need to put your thumb on the scales to make the output actually be useful. They're at best very janky tools as soon as you're working with things that require context that isn't easily contained in some kind of confined area of work.

Currently we are seeing the phenomenon 'habsburg AI' where AI's consume their own outputs as training data, which rapidly deteriorates their ability to actually be useful for much of anything.

The thing is that there literally isn't enough human-made data to keep feeding them (they already ate the entire internet), so if you both want to continue ramping their intake of data and you also don't want them to get rapidly weird and completely useless, you pretty much have to get in there with elbow grease. Removing or deprioritizing data that's tripping up the model is one of the few ways you can do human-assisted refinement of these things.

The sooner we all face the music that these things aren't magical truth machines, have a long way to go and there is no guaranteed rate of growth, the sooner this hype cycle can end.

Re: Machine Unlearning in 2024

#83

Earlier quoted context omitted.

A business cannot read a book, and your machine learning model is not given human rights.

> A business cannot read a book Assume the human read the book as part of their job . Is that using copyrighted material for commercial purposes? If that doesn't count then I'm not sure why you brought up "commercial purposes" at all. > This rules with harsh penalties for consumers/small companies but not for bigtech double standard is bullshit, though. Consumers and small companies get away with small copyright viol…

> Assume the human

Humans have rights. They get to do things that businesses, and machine learning models, or general automation, don't.

Just like you can sit in a library and tell people the contents of books when they ask, but if you go ahead and upload everything you get bullied into suicide by the US government[1]

> Consumers and small companies get away with small copyright violations all the time

Yeah, because people don't notice so they don't care. Everyone knows what these bigtech criminals are doing.

[1] https://en.wikipedia.org/wiki/Aaron_Swartz

Re: Machine Unlearning in 2024

#84
post #80
post #77

Earlier quoted context omitted.

All this is technically correct, but it also means this technology is absolutely not ready to be used for anything remotely involving humans or end user data.

Why? We use random data in lots of applications, and there's always the theoretical probability that it could 'spell something naughty'.

It's about models' ability to unlearn information or to configure their training environment so that something is never learned in the first place... is not exactly the same as "oups, we logged your IP in a log by accident".

A company is liable even if they have accidentally retained / failed to delete personal information. That's why we have a lot of standards and compliance regulation to ensure a bare minimum of practices and checks are performed. There is also the cyber resilience act coming up.

If your tool is used by/for humans, you need beyond 100% certitude exactly what happens with their data and how it can be deleted and updated.

Re: Machine Unlearning in 2024

#85

We need to consider the practicality of unlearning methods in real-world applications and the legal acceptance of the same. Given current technology and what advancements are needed to make Unlearning more possible, probably there should be a time-to-unlearn kind of an acceptable agreement that allows organizations to retrain or tune the response that does not involve any response from the to-be-unlearned copyright c…

> We need to consider the practicality of unlearning methods in real-world applications and the legal acceptance of the same. > probably there should be a time-to-unlearn kind of an acceptable agreement

A very important distinction is between data storage and data use/dissemination. Your comment hints at "use current model until retrained is available and validated", which is an extremely dangerous idea.

Remember old times of music albums distributed over physical media. Suppose a publisher creates a mix, stocks shelves with album and it becomes known that one of the tracks is not properly licensed. It would be expected that it takes some time to execute distribution shutdown: distribute order, clean up shelves, etc. However, time for another production run with a modified tracklist would be entirely the problem of the publisher in question.

The window for time-to-unlearn should only depend on practicality of stopping information dissemination, not getting updated source ready. Otherwise companies will simply wait for model to be retrained on a single 1080 and call it a day, which would effectively nullify the law.

Re: Machine Unlearning in 2024

#87

Earlier quoted context omitted.

> A business cannot read a book Assume the human read the book as part of their job . Is that using copyrighted material for commercial purposes? If that doesn't count then I'm not sure why you brought up "commercial purposes" at all. > This rules with harsh penalties for consumers/small companies but not for bigtech double standard is bullshit, though. Consumers and small companies get away with small copyright viol…

> Assume the human Humans have rights. They get to do things that businesses, and machine learning models, or general automation, don't. Just like you can sit in a library and tell people the contents of books when they ask, but if you go ahead and upload everything you get bullied into suicide by the US government[1] > Consumers and small companies get away with small copyright violations all the time Yeah, because…

> Humans have rights. They get to do things that businesses, and machine learning models, or general automation, don't.

So is that a yes to my question?

If humans are allowed to do it for commercial purposes, and it's entirely about human versus machine, then why did you say "Using copyrighted content for commercial purposes should be a violation" in the first place?

> Just like you can sit in a library and tell people the contents of books when they ask,

You know there a huge difference between describing a book and uploading the entire contents verbatim, right?

If "tell the contents" means reading the book out loud, that becomes illegal as soon as enough people are listening to make it a public performance.

> but if you go ahead and upload everything you get bullied into suicide by the US government[1]

They did that to a human... So I've totally lost track of what your point is now.

Re: Machine Unlearning in 2024

#88

Earlier quoted context omitted.

> Assume the human Humans have rights. They get to do things that businesses, and machine learning models, or general automation, don't. Just like you can sit in a library and tell people the contents of books when they ask, but if you go ahead and upload everything you get bullied into suicide by the US government[1] > Consumers and small companies get away with small copyright violations all the time Yeah, because…

> Humans have rights. They get to do things that businesses, and machine learning models, or general automation, don't. So is that a yes to my question? If humans are allowed to do it for commercial purposes, and it's entirely about human versus machine, then why did you say "Using copyrighted content for commercial purposes should be a violation" in the first place? > Just like you can sit in a library and tell peop…

> and it's entirely about human versus machine

It's not. Those were what's called examples. There is of course more to it. Stop trying to pigeonhole a complex discussion onto a few talking points. There are many reasons why what OpenAI did is bad, and I gave you a few examples.

Re: Machine Unlearning in 2024

#89

Earlier quoted context omitted.

> Humans have rights. They get to do things that businesses, and machine learning models, or general automation, don't. So is that a yes to my question? If humans are allowed to do it for commercial purposes, and it's entirely about human versus machine, then why did you say "Using copyrighted content for commercial purposes should be a violation" in the first place? > Just like you can sit in a library and tell peop…

> and it's entirely about human versus machine It's not. Those were what's called examples. There is of course more to it. Stop trying to pigeonhole a complex discussion onto a few talking points. There are many reasons why what OpenAI did is bad, and I gave you a few examples.

I'm not trying to be reductive or nitpick your example, I was trying to understand your original statement and I still don't understand it.

There's a reason I keep asking a very generic "why did you bring it up", it's because I'm not trying to pigeonhole.

But if it's not worth explaining at this point and the conversation should be over, that's okay.

Re: Machine Unlearning in 2024

#90
post #84
post #80

Earlier quoted context omitted.

Why? We use random data in lots of applications, and there's always the theoretical probability that it could 'spell something naughty'.

It's about models' ability to unlearn information or to configure their training environment so that something is never learned in the first place... is not exactly the same as "oups, we logged your IP in a log by accident". A company is liable even if they have accidentally retained / failed to delete personal information. That's why we have a lot of standards and compliance regulation to ensure a bare minimum of pr…

You can never even got to 100% certainty, yet alone 'beyond' that.

Google can't even get 100% certainty that they eg deleted a photo you uploaded. No AI involved. They can get an impressive number of 9s in their 99.9..%, but never 100%.

So this complaint when taken to the absolute like you want to take it, says nothing about Machine Learning at all. It's far too general.

Post reply on HN