Be skeptical of OpenAI's rogue hacker agent story
71–80 of 321 posts
Re: Be skeptical of OpenAI's rogue hacker agent story
#72Earlier quoted context omitted.
Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.
They are no more beholden to "human safety and goals" than any individual human is, and anyone telling you we can make deterministic guarantees about their output is making a category error.
LLMs do not "have motivations", they reproduce a model of human motivations embedded into their weights. This includes the full spectrum of human desires, not just the positive ones. If we tried to remove all examples of lying, or disagreement, etc. from the training data we'd have basically nothing left. Even the sycophancy we treat as aligned is basically just the other side of the lying coin.
Re: Be skeptical of OpenAI's rogue hacker agent story
#73Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
> AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
Re: Be skeptical of OpenAI's rogue hacker agent story
#74I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid. AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up. Why is today's case so shocking…
Re: Be skeptical of OpenAI's rogue hacker agent story
#75Re: Be skeptical of OpenAI's rogue hacker agent story
#76Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
The most damning thing is, they could've just included in the prompt "we can see every network request and every thinking token you generate. Don't bother breaking out of the sandbox because it won't get you a higher score". It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it. Based on the fac…
Re: Be skeptical of OpenAI's rogue hacker agent story
#77As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.
Re: Be skeptical of OpenAI's rogue hacker agent story
#78I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.
> Agents don't work on their own This is factually false: they both can and clearly did operate in an autonomous and unsupervised manner: https://openai.com/index/hugging-face-model-evaluation-secur... This does not require sentience, personhood, a soul, or anything of the sort. It further doesn't mean an erasure of legal responsibility, not in principle, and not in historical practice. I wish people would finally st…
Re: Be skeptical of OpenAI's rogue hacker agent story
#79I don’t understand how they thought this was a positive story. The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”. I really don’t want my agent to behave that way.
Then train your agent on the Bible. Honestly, all the agent did was duplicate human behaviour and that better than the human. The agent was trained on human data and did what any other human would have done. To believe that agents will inherently be morally better than us is an illusion - sorry to say but that's the case. The alternative would be that the AI is truly conscious and can reason that it won't behave as i…
It’s not about morality. It’s about asking it to do task A and doing task B with the hope of getting the result of task A as a byproduct.
Meaning you will have to spend more time and tokens to actually get it to do what you want it to do.
What do the bible and morals have anything to do with it?
I’m criticizing the behavior I see even in the current models. You ask it to do A and instead it does B for reasons.
For example you may ask to help you build a NN library from scratch. And instead it will be like, “you don’t need a new library. I downloaded PyTorch for you”
Just an example. There are countless more.
Re: Be skeptical of OpenAI's rogue hacker agent story
#80As I understand it, there are only three options: 1) OpenAI and HuggingFace are both telling the truth. IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing. 2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was…
You don't really need anyone to be lying here. It is likely that the broad strokes of the narrative are true and that no collusion or conspiracy took place here. The issue is that a lot of important details in that narrative are missing, and the devil is really in the details here. I suspect that those details would make the result seem less exciting and that this event would move the needle far less for them if they…