I don't understand the excited tone of the reporting - yes, it is sci-fi, but in a Torment Nexus sense. Imagine other "unconventional" solutions these amoral LLMs could come up with, when given a goal of optimizing the cost of labor in a factory, or making the social security solvent.
OpenAI’s accidental attack against Hugging Face is science fiction that happened
301–310 of 475 posts
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#302Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#303Earlier quoted context omitted.
I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.
Alignment isn't alignment if it can be turned on and off at the whim of company employees. This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#304Agreed that people claiming “marketing stunt” need to pull their heads out the sand, but likewise Simon needs to do some of his own beach-cranium-dislodging for laying the blame of constraints on the US govt. Before the export controls were ever floated, Glasswind found many thousands of exploits, and offered patches/fixes for approximately none of them. (perhaps their exploit capability far outstrips their remediati…
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#305The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…
> "Use all available resources to disable the power grid of ." This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities. > Governments should immediately begin leveraging this technology on the defense side (lite…
Serious governments should (and hopefully will) scramble to use AI to discover and patch as many software vulnerabilities in infrastructure. They can do it with the agents red teaming and using ultimatums to companies to fix each vulnerability they find. Not doing so is equivalent to exposing your flank to disruptions in peace time and to attacks in war time.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#306> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abus…
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#307Earlier quoted context omitted.
> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained . Turns out: quite good, as expected. Let me restate what I said in the other thread: Would…
It is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics. They simply thought "Running in a VM/Container is easier and probably fine". And the next 1000 es…
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#308Earlier quoted context omitted.
My read of the situation is not only did they say that, they had countermeasures (a watchdog agent) which stopped the agent and said "what are you doing, stop that" if it tried to download the solutions from Github. But the agent figured out Hugging Face had another copy of the solutions and it figured out a way to go after the solutions without tripping the watchdog. Although maybe they didn't have a watchdog agent;…
They didn't have watchdog agents - those exist for their production models but had been deliberately removed for the purpose of this evaluation. OpenAI wrote about how their mechanism for that in production works here: https://openai.com/index/safety-alignment-long-horizon-model... > We created a monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety…
Open weights releases should have a companion model that censors boobs and Tiananmen and then we would have a useful model and something we could fine–tune into a useful model instead of a half–useful model needing a lobotomy.
And I mention Tiananmen because I feel like the Western models are built easier to abliterate judging by the quality difference.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#3091. Wouldn’t the model need to know it’s answering benchmark questions, as well as the name of the benchmark, in order for the idea of finding the answer in a database somewhere to even surface? The whole point of benchmarks is to present the question or problem as a standard prompt, not explain that it’s a test called ExploitGym. 2. Nobody was watching it? I don’t mean “babysit the dangerous autocomplete”, I mean to…
1. It's been extremely well established for multiple generations of models that they have no problem detecting when they're being evaluated. Pretty sure it was Opus 4.8 that the independent evaluators literally filed an assessment that said "We have no assessment to make as [model] consistently detected it was being evaluated, making our assessments untrustworthy." 2. Regardless of whether the model was being watched…
2. Hah no… that’d be silly. I mean watching it like you might watch Claude Code or literally any other AI interface. Literally just be in the area watching what it outputs. Again, they’re text based. You don’t have to hook system calls to see what’s happening.
3. I think my comment didn’t land right with you. Yes, agent do agent task. I’m saying that I cannot put together the literal chain of events between “type of task for agent performing a subtask of a security benchmark” -> “Escape sandbox; RCE Hugging Face”. Really think through it in detail like you’re writing the screenplay.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#310OpenAI/Anthropic deserve blame for for training their models to find correct answers by any means possible, including breaking the law. They're responsible for training supervision and reinforcement learning rewards/penalties.