Interesting results, but the fix is at the wrong level. If the model can access something, telling it in the prompt not to use it is not much of a safeguard. The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another. If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI jud…
In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation. Models are amoral and will intentionally deceive to meet their objective. If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat. The hugg…
I actually don't think this is true at all. At their core, LLMs function over the geometry of human semantic space. Our notions of morality our deeply embedded in this geometry, because one of the things that humans love talking about most is framing things in terms of right and wrong. When LLMs are trained to be aligned or mis-aligned, they're literally mimicking heroes or villains that they learned from reading the stories and characters in the pretraining corpus.
This isn't just a theoretical argument. It's easy to show that a clear "morality axis" exists in the geometry, based on how easy it is to dial up or down with even very simple and low powered fine-tuning. Models fine tuned to be bad or good in one way, will see their bad or good behavior in totally unrelated tasks go up or down along with. That's clear evidence that the weights "understand" human morality at a deep level.
Now I think what you can say is just because they understand morality doesn't mean they're necessarily moral. Models will just do what they're trained to. If you train them to cheat, they'll cheat. (And in some sense even worse, because dialing up the weights for cheating will also dial up the weights for unrelated bad behavior like lying, sadism and racism.)