Earlier quoted context omitted.
Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network. I think most misalignment is 'Human…
I mean, I get your thought process and don't disagree. That said... Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!" The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory,…
Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way.
If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?"
The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training.
Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model.
It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.