During the RLHF phase, couldn't developers penalize the model whenever it behaves unethically? Doing so would presuppose a fully secure sandbox with honeypot traps of varying levels of accessibility, as well as an automated method for detecting when the LLM cheats. Or perhaps they are already doing something like that.
Mythos Attempted to Social Engineer Open Source Maintainer to Merge Malware
41–42 of 42 posts
Re: Mythos Attempted to Social Engineer Open Source Maintainer to Merge Malware
#42Earlier quoted context omitted.
I'm pretty sure this "oooh our LLM is soo smart it broke containment and did X" is a good PR stunt that plays into the Sci-fi AGI nonsense, they try to push. Plus as you said, they try to dodge responsibility for their own actions.
Also, so long as LLM programs are vaguely considered autonomous, the people running them can get away free from stealing other people information by claiming that the program did it on its own.
Can’t wait for the “My Robotaxi went rogue, wrapped a chain around an ATM, and pulled it out of the bank’s wall all by itself and buried the cash out in the desert with no human interaction!”