Safety and alignment in an era of long-horizon models
1–7 of 7 posts
Re: Safety and alignment in an era of long-horizon models
#2Also for typical normal use case for these smart models, you'd probably want an actual "max turns" limit to AVOID the pathological persistence (which in itself would be misaligned for "normal" tasks).
Re: Safety and alignment in an era of long-horizon models
#3Re: Safety and alignment in an era of long-horizon models
#4And of course someone in the comments needs to link to the Zealous Autoconfig XKCD, so I’ll do it: https://xkcd.com/416/
Re: Safety and alignment in an era of long-horizon models
#5Re: Safety and alignment in an era of long-horizon models
#6The article is rather light on "what to actually do about it". Even the basic "run it in isolated container without access to anything it does not need for the task" would have already improved the situation considerably (from the article it really seems like they didn't do that) - then the model would have to find local privilege exploits to actually escape (much cleaner misaligned behavior). Also for typical normal…
So if you don't want a model to do something, make sure it's running in an environment where it cannot do that thing - including via loopholes.