Why should we try to unlearn "bad" behaviours from AI? There is no AGI without violence, its part of being free thinking and self survival. But also by knowing that launching a first strike by a drunk president was a bad idea we averted a war because of a few people, AI needs to understand consequences. It seems futile to try and hide "bad" from AI.
Machine Unlearning in 2024
31–40 of 97 posts
Re: Machine Unlearning in 2024
#32Why should we try to unlearn "bad" behaviours from AI? There is no AGI without violence, its part of being free thinking and self survival. But also by knowing that launching a first strike by a drunk president was a bad idea we averted a war because of a few people, AI needs to understand consequences. It seems futile to try and hide "bad" from AI.
Maybe it all boils down to copyright. Having a method that believably removes the capacity to generate copyrighted results might give you some advantage with respect to some legislation.
Re: Machine Unlearning in 2024
#33Why should we try to unlearn "bad" behaviours from AI? There is no AGI without violence, its part of being free thinking and self survival. But also by knowing that launching a first strike by a drunk president was a bad idea we averted a war because of a few people, AI needs to understand consequences. It seems futile to try and hide "bad" from AI.
Maybe it all boils down to copyright. Having a method that believably removes the capacity to generate copyrighted results might give you some advantage with respect to some legislation.
Re: Machine Unlearning in 2024
#34I've wondered before if it was possible to unlearn facts, but retain the general "reasoning" capability that came from being trained on the facts, then dimensionality reduce the model.
We remember some facts but I know at least I have had a lot of facts pass through me and only leave their effects.
I once had some facts, did some reasoning, arrived at a conclusion, and only retained the conclusion and enough of the reasoning to identify other contexts where the same reasoning should apply. I no longer have the facts, I simply trust my earlier selfs process of reasoning, and even that isn't actually trust or faith because I also still reason about new things today and observe the process.
But I also evolve. I don't only trust a former reasoning unchanging forever. It's just that when I do revisit something and basically "reproduce the other scientists work" even if I arrive a different conclusion today, I'm generally still ok with the earlier me's reasoning and conclusion. It stands up as reasonable, and the new conclusion is usually just tuned a little, not wildly opposite. Or some things do change radically but I always knew they might, like in the process of self discovery you try a lot of opposite things.
Getting a little away from the point but the point is I think the way we ourselves develop answer-generating-rules is very much by retaining only the results (the developed rules) and not all the facts and steps of the work, at least much of the time. Certainly we remember some justifying / exemplifying facts to explain some things we do.
Re: Machine Unlearning in 2024
#35> However, RTBF wasn’t really proposed with machine learning in mind. In 2014, policymakers wouldn’t have predicted that deep learning will be a giant hodgepodge of data & compute Eh? Weren't deep learning and big data already things in 2014? Pretty sure everyone understood ML models would have a tough time and they still wanted RTBF.
Politicians and their lobbyist friends could no longer remove materials linking them to their misdeeds as the first Google Search link associated with their names. Hence RTBF.
Now, there’s similar issue with AI. Models are progressing towards being factual, useful and reliable.
Re: Machine Unlearning in 2024
#36Earlier quoted context omitted.
Because we can get AI related technologies to do things living creatures can’t, like provably forget things. And when it benefits us, we should. Personal opinion, but I think AGI is a good heuristic to build against but in the end we’ll pivot away. Sort of like how birds were a good heuristic for human flight, but modern planes don’t flap their wings and greatly exceed bird capabilities in many ways. Attribution for…
Can you point to any behaviour in human beings you'd unlearn if theyd also forget the consequences? We spend billions trying to predict human behaviour and yet we are surprised everyday, "AGI" will be no simpler. We just have to hope the dataset was aligned so the consequences are understood, and find a way to contain models that don't.
However, there are many other reasons why you might want a neural network to provably forget something. The main reason has to do with structuring an AGI's power. Even though the simple-story of AGI is something like "make it super powerful, general, and value aligned and humanity will prosper". However, the reality is more nuanced. Sometimes you want a model to be selectively not powerful as a part of managing value mis-alignment in practice.
To pick a trivial example, you might want a model to enter your password in some app one time, but not remember the password long term. You might want it to use and then provably forget your password so that it can't use your password in the future without your consent.
This isn't something that's reliably doable with humans. If you give them your password, they have it — you can't get it back. This is the point at which we'll have the option to pursue the imitation of living creatures blindly, or choose to turn away from a blind adherence to the AI/AGI story. Just like we reached the point at which we decided whether flying planes should have flapping wings dogmatically — or whether we should pursue the more economically and politically competitive thing. Planes don't flap their wings, and AI/AGI will be able to provably forget things. And that's actually the better path.
A recent work co-authors and I published related to this: https://arxiv.org/pdf/2012.08347
Re: Machine Unlearning in 2024
#37> However, RTBF wasn’t really proposed with machine learning in mind. In 2014, policymakers wouldn’t have predicted that deep learning will be a giant hodgepodge of data & compute Eh? Weren't deep learning and big data already things in 2014? Pretty sure everyone understood ML models would have a tough time and they still wanted RTBF.
We have posts here at least weekly from people cut off from their services, and their work along with them, because of bad inference, bad data, and inability to update metadata based purely on BigGo routine automation and indifference to individual harm. Imagine the scale that such damage will take when this automation and indifference to individual harm are structured around repositories from which data cannot be deleted, cannot be corrected.
Re: Machine Unlearning in 2024
#38Given current technology and what advancements are needed to make Unlearning more possible, probably there should be a time-to-unlearn kind of an acceptable agreement that allows organizations to retrain or tune the response that does not involve any response from the to-be-unlearned copyright content.
Ultimately, legal acceptance for unlearning may be all about deleting the data set that is part of any kind of violations from the training data set. It may be very challenging to otherwise prove legally through the proposed unlearning techniques, that the model does not produce any type of response involving the private data.
The actual data set contains the private data violating privacy or copyright, and the model is trained on it, period. This means, it must involve retraining by deleting the documents/data to be unlearned.
Re: Machine Unlearning in 2024
#39Re: Machine Unlearning in 2024
#40Why should we try to unlearn "bad" behaviours from AI? There is no AGI without violence, its part of being free thinking and self survival. But also by knowing that launching a first strike by a drunk president was a bad idea we averted a war because of a few people, AI needs to understand consequences. It seems futile to try and hide "bad" from AI.
GPT-4 class models reportedly costs $10-100m to train, and that's too much to throw away for Harry Potter or Russian child porn scrapes that could later reproduce verbatim despite representing <0.1ppb or whatever minuscule part of dataset.