Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

101–110 of 302 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#102

"The agents found a Modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face." Why, what was the prompt? I told Claude today to wire plugins on Linux into a sound pipeline to remove noise. Did some astonishing things, played sound through the pipeline, measured it etc. I told it to optimize my sound for TF2 and it played the spy_decloak samples, measured them and made them…

This was clearly explained by OpenAI in their initial press release on 7/21 [0]:

> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. […] The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

[0] https://openai.com/index/hugging-face-model-evaluation-secur...

Re: Timeline of the OpenAI accidental attack against Hugging Face

#103
post #93

I think one of the most interesting details here might be tucked away in that first bulletin point: > May 7: OpenAI starts a new training run for an experimental, unreleased model. (Do they mean an evaluation run? They say training run in the video, and later mention a “reward signal to judge how well they’re doing”, so I guess this really was about training a model, not evaluating one that was already trained.) The…

I'm just reading the captions of the video for May 7th. They clearly say at 10:18: "we kick off a new reinforcement learning run to train a next frontier model . It the captions are correct, there is no ambiguity.

Thanks, I just updated that note in the post to quote that snippet.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#104
post #46

Norbert Wiener in 1960: "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of…

Maybe they didn't have proper debuggers in 1960? For a language model you need (RNG state, context, prompt).

So if they wrote an LLM step by step debugger, it would be all deterministic. But they prefer rapid sales, chaos and mystique.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#106

Earlier quoted context omitted.

I think it's a show of these agents happily bypassing security to get stuff done. I've actually observed similar behavior at home. I have a k3s cluster running at home. I asked an agent to check some stuff as a normal user but I had kubectl access to the k3s cluster. Part of the research, I'd allowed access to run kubectl commands for spinning up test containers. However, when the agent ran into something that needed…

" bypassing security" If they can bypass it there is no security and the security was flawed all along.

There is no perfect security. It's always flawed in some way.

Good security is extremely hard.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#107
Simon's retelling is more compact but it also invites anthropomorphization of the sharing of the familiarity with the message board which re-emerged a few times.

Zvi's retelling handles this better. Zvi speculates that the secret message board familiarity was carried because it had been trained into the May-and-subsequent models: https://thezvi.substack.com/p/openai-trained-its-models-for-...

Re: Timeline of the OpenAI accidental attack against Hugging Face

#108
post #99

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

Their position makes no sense to me. I don’t see how you can be a mainstream company selling your services worldwide (almost) if you also believe that you’re building an extremely dangerous AGI (supposedly based on the same technology you’re offering to everyone). If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons. An…

The companies are begging to be regulated for this reason and have been doing so for years. HN's response is generally that this is performative for marketing or seeking regulatory capture or haha anthropic you get what you ask for. Maybe the cynics are right, but there's really nothing inconsistent about the naive view here, once you factor in race dynamics and obligations to investors.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#109
post #46

Norbert Wiener in 1960: "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of…

Maybe they didn't have proper debuggers in 1960? For a language model you need (RNG state, context, prompt). So if they wrote an LLM step by step debugger, it would be all deterministic. But they prefer rapid sales, chaos and mystique.

We also have engineer blindness, so having human in the loop confirming thousands of requests would quickly start to confirm everything without looking.

It would become just another system to hack through, and slow the development process as well. The OpenAI video in the article recommends an autonomous defense mechanism. For rapid reaction, but I don’t know how sustainable or effective that would be, or if as humans we will be able to keep up.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#110
post #46

Norbert Wiener in 1960: "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of…

Maybe they didn't have proper debuggers in 1960? For a language model you need (RNG state, context, prompt). So if they wrote an LLM step by step debugger, it would be all deterministic. But they prefer rapid sales, chaos and mystique.

[dead]
Post reply on HN