Earlier quoted context omitted.
This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.
Even the behavior of agents searching for sandbox bypasses must have been in the training data, or at the very least, "suggested" in some way. To be this whole thing feels like a marketing play by OpenAI.
In fact, I wonder if teaching it "this is a bypass" would help it to model when it's doing its job vs working around the job.