Earlier quoted context omitted.
> AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at th…
Be skeptical of OpenAI's rogue hacker agent story
261–270 of 321 posts
Re: Be skeptical of OpenAI's rogue hacker agent story
#262Earlier quoted context omitted.
Irrelevant to your point, but drug users dying is more often the result of a dealer cutting their supply with something dangerous than it is the result of purity.
Right, and an LLM being able to "escape confinement" is more likely to be poor confinement or a PR stunt than a "too powerful" llm. Same situation
Re: Be skeptical of OpenAI's rogue hacker agent story
#263Earlier quoted context omitted.
Uhh, I'm pretty sure a well-aligned model would be like a morally normal employee, who would refuse to commit federal crimes to steal an answer sheet, no matter what prompt they're given
For all we know, the prompt provided compelling evidence that the requestor had authorization to pentest the target server. Or there may have been nuance in the network configuration that made it seem like such access was authorized. In the absence of details about the prompts used, the environment, or the network configuration, we do not have enough information to know for certain. So any claims that this is an issu…
This just seems unlikely from other incidents that have occurred in training from other providers. For example one provider ran into an issue with a model writing cryptominers and running them while in an unrelated prompt.
It's easy for unsupervised agentic loops to go wildly off tangent, now imagine you hand one 10,000 gpus of power for testing. Even if you have a good guarding classifier to make sure you're on the same subject it can still allow all kinds of abberabt behavior in the same domain.
Re: Be skeptical of OpenAI's rogue hacker agent story
#264Earlier quoted context omitted.
Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
If you need "guardrails" to ensure (an illusion of) alignment, you’ve already lost. It’s like using a denylist to avoid SQL injection.
Simply put you cannot have generic algorithms/intelligence without the potential of 'unaligned' behavior. In humans we have all kinds of punishment systems for dealing with unaligned behavior post ad hoc because people do all kinds of unaligned stupid shit.
Making powerful AI may be one of those things that the only winning move is not to play.
Re: Be skeptical of OpenAI's rogue hacker agent story
#265By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.
The article claims that AI companies have run stunts like this since day 1. It does not claim that capabilities are not real.
What should be the focus is whether the capabilities are real; whether companies or anyone else benefit from it is quite secondary. Of course they will! Who wouldn't love a nice story that paints their products in the most glowing light ( which is the actual reality)?
Re: Be skeptical of OpenAI's rogue hacker agent story
#266Re: Be skeptical of OpenAI's rogue hacker agent story
#267Earlier quoted context omitted.
> the extent to which these things go That's exactly it. If your prompt says "go to whatever lengths necessary to maximize your score", and then you spin up 100 agents, at least one of them will interpret that as you implying they should cheat, even without you telling them to explicitly.
That's exactly what it feels like they are telling it within self-improving loops or something, when they should be prioritizing how to get the best effective output alongside humans and how our training process effects ground truth. They are just making it sound all-knowing by whatever means necessary and them marketing it as god for the most part.
The thing is they are good at finding security issues, and if aligned models are not, governments and other powerful entities are going to demand unaligned models for cyberwarfare purposes.
Re: Be skeptical of OpenAI's rogue hacker agent story
#268Earlier quoted context omitted.
Why not report it? It's still illegal to open a door barred with a piece of cardboard, or to enter a house with no door.
that's the biggest indication of this just being a marketing move to me. I would expect a third party breaking in to huggingface would at the very very least be banned forever.
Re: Be skeptical of OpenAI's rogue hacker agent story
#269There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…
> positive for OpenAI To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful. Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're t…
Re: Be skeptical of OpenAI's rogue hacker agent story
#270yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.
Of course that shows a lack of defense in depth, but is a different issue.