This paper is the most reasonable one I have read on AI safety. We need to fundamentally change the training pipelines by figuring out better ways to ‘reward’ behavior. Yoshua didn’t explicitly mention training data, but we probably need to only use synthetic data that contains no text that could motivate bad behavior via imitation. I feel like a heretic for saying this, but I will say it anyway: AI agents are great…
I’m sure it’s impossible to completely weed it out, but are the labs doing any of this kind of data sanitation?