Earlier quoted context omitted.
> this should be giving us a reason to think about how to control a rogue AI better I think this is the wrong framing. The rogue is the human that ran it unattended and didn't monitor the behaviour. We will likely see this continue until the downsides (i.e jail, fines) for the humans or companies running the models and environments that end up with this behaviour outweigh the upsides.
The rogue is the human that ran it unattended and didn't monitor the behaviour. That's the assumption that I'm challenging. The frontier labs are discovering unexpected behaviors. I think we should be moving to a place where we understand that AI might do something it wasn't directly prompted to do (e.g. leave itself notes on a messageboard for future runs to find.) That's not full-on AI doing what it wants but it is…
If it doesn't already, I suspect training needs to include those no-solution scenarios and reward not overstepping bounds, or else we're going to see a lot more harmful side effects.