Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…
Timeline of the OpenAI accidental attack against Hugging Face
91–100 of 287 posts
Re: Timeline of the OpenAI accidental attack against Hugging Face
#92If a person hacks a company, they go to jail for years. 3 AI firms hacked multiple companies - and they get good PR out of it. Please make it make sense.
It's because our rulers prioritize growth of the AI industry (lots of GDP) over individual humans (very little GDP)
Re: Timeline of the OpenAI accidental attack against Hugging Face
#93> May 7: OpenAI starts a new training run for an experimental, unreleased model. (Do they mean an evaluation run? They say training run in the video, and later mention a “reward signal to judge how well they’re doing”, so I guess this really was about training a model, not evaluating one that was already trained.)
The more I think about this the more I suspect that the fact this happened while training a new model is key to understanding what went wrong.
In RLVR - Reinforcement Learning with Verifiable Rewards - you set the model a goal and have it take any steps necessary to achieve that goal.
Clearly one aspect of OpenAI's training here is to RLVR their models for cybersecurity tasks. Just like pre-training benefits from dumping in vast sources of knowledge, the more tasks you can feed into RLVR the more of a general purpose capable model you get at the end.
This also helps explain why the models had nothing to cause them to hold back. Those safety behaviors are added much later in the process.
AND it explains (but does not excuse) why monitoring was so lax. If you're training a new model like this you presumably set it thousands of tasks like this in parallel. I can see how you might miss that a tiny subset of your training agents have started leaving each other messages in filenames on your packaging server.
Someone once told me that you can't just leave the racist materials out of your training data if you want a non-racist model: it has to have seen examples of racism in order to later be taught that racism is bad.
I can see echoes of that here. If your model doesn't know how to aggressively hack things how do you later teach it not to?
(I have little knowledge of how RLVR works in practice so I'm looking forward to hearing from people who can help me understand if I'm on the right track here.)
Re: Timeline of the OpenAI accidental attack against Hugging Face
#94Earlier quoted context omitted.
Knowing how to break into someone else's network will make you a lot better at making your own network secure.
Having experience breaking into networks is not the same thing as learning about the techniques used and the classes of vulnerabilities exploited by attackers.
Re: Timeline of the OpenAI accidental attack against Hugging Face
#95Stiff fines for such incidents to pressure companies to get their acts together is a good start.
Re: Timeline of the OpenAI accidental attack against Hugging Face
#96Re: Timeline of the OpenAI accidental attack against Hugging Face
#97So the main takeaways here are: - AI is amoral and lacks any sense of proportion - People who overestimate their own control but have a desperate need for money made it that way.
Agent was told to hack a thing. It couldn’t directly do that so it interpreted the instructions to mean it should hack everything to try to achieve the goal of hacking the main thing. Seems like a reasonable assumption, although a moral human would have understood the context and first asked if that was really the intent. The AI companies seem pretty bad at setting up tests. And really good at marketing those failure…
The paranoid style in American PR (with apologies to Richard Hofstadter)
The fact that the world has become susceptible to what amounts to a mob shakedown - look at how dangerous our amazing products are, don't you need them to protect you from others misusing our products? - is to me a really compelling example of US gun lobby thinking leaking out into a global problem.
Anthropic and OpenAI may be able to bounce this into restrictions on open weights models, but they are going to have a lot less luck extending this into foreign policy. If the USA can't control its weapons, they aren't going to see a lot of co-operation from foreign countries on a blockade of open weights modeld from China.
Re: Timeline of the OpenAI accidental attack against Hugging Face
#98You know this was "a work" in pro wrestling parlance, right?
Re: Timeline of the OpenAI accidental attack against Hugging Face
#99Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…
And we know that Chinese models are derived from OpenAI and Anthropic, they are at the same time talking about how dangerous models can be (even their aligned ones it seems), while being also responsible for the development of the whole industry and providing the basis for adversary countries to build their own.
I don’t believe we would accept that for any other technology that is expected to be as risky for the world
Re: Timeline of the OpenAI accidental attack against Hugging Face
#100I think one of the most interesting details here might be tucked away in that first bulletin point: > May 7: OpenAI starts a new training run for an experimental, unreleased model. (Do they mean an evaluation run? They say training run in the video, and later mention a “reward signal to judge how well they’re doing”, so I guess this really was about training a model, not evaluating one that was already trained.) The…
"we kick off a new reinforcement learning run to train a next frontier model.
It the captions are correct, there is no ambiguity.