Earlier quoted context omitted.
Oh! Good insight! ... but we already have that, right? Undesirable outcomes from rulesets that govern the behaviour of a large number of (conscious or unconscious) actors are a pretty standard feature of civilisation
It's the paper clip maximizer idea [1]. "The goal of maximizing paperclips is chosen for illustrative purposes because it is very unlikely to be implemented, and has little apparent danger or emotional load (in contrast to, for example, curing cancer or winning wars). This produces a thought experiment which shows the contingency of human values: An extremely powerful optimizer (a highly intelligent agent) could seek…
> Our algorithm’s performance is only as good as the human evaluator’s intuition about what behaviors look correct, so if the human doesn’t have a good grasp of the task they may not offer as much helpful feedback. Relatedly, in some domains our system can result in agents adopting policies that trick the evaluators. For example, a robot which was supposed to grasp items instead positioned its manipulator in between the camera and the object so that it only appeared to be grasping it, as shown below.
[0] https://blog.openai.com/deep-reinforcement-learning-from-hum...
https://news.ycombinator.com/item?id=14545298 78 upvotes, 7 comments