May be useful, but it seems to me that the reward function still is relatively easy to specify? Much of the difficulty in AI safety is due to specify what humans really want. Perhaps the AI can observe a human playing the game and learn a reward function?
A subarea of AI research focused on learning the reward function is Inverse Reinforcement Learning. Here’s an article on it:
Learning from humans: what is inverse reinforcement learning? https://thegradient.pub/learning-from-humans-what-is-inverse...