Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

471–475 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#471
post #330

Earlier quoted context omitted.

> I think it's a criminal offence and should be a true test of who is held accountable when an AI agent commits a crime. I agree, lets use the favorite analogy. OpenAI encouraged a smart and eager junior engineer to find any way whatsoever to get a higher score on the benchmark. Then, the junior breaks into HuggingFace to get a higher score. That would be a big deal involving the FBI, not press releases and blog post…

I don't see a scenario where a company would be liable for the employee's actions unless they had specifically been told/encouraged to break the law. If your boss tells you to fix a bug and you go kill the customer which one of you is going to jail? "But I solved the problem!" isn't exactly going to fly as a defense.

Negligence is criminally prosecutable.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#472
post #469

Earlier quoted context omitted.

Are you referring to table 1 from the ExploitGym site [1] for the 102.1 average mins for Mythos and 69.8 for GPT-5.5? These are under a two-hour time limit. In Figure 5, the authors experiment with extending the time limit to 6 hours and show that Mythos keeps improving. It seems pretty straightforward to me that OAI decided to run a variation of the eval with an even higher time limit. Regarding seeing the metrics t…

Yes, that’s the table. Elsewhere in here I explained that I took the average of them and noted the 6 hour experiment, but that the average would give us a feel for expected time per task. By the time I got to the comment you’re replying to I was short-handing the conclusion; that’s on me. My point was that they had some rough idea on what to expect, and it’s not in the range of days or weeks. Even if they wanted to t…

It is hard to convey how both chaotic and high pressure the working environment of the labs are if you have not worked at one. Engineers and researchers are routinely overworked with multiple high-priority workstreams at a time. It is very believable to me that a researcher noticed an eval job running for 2 days instead of 1, asked an engineer to look into it, and both forgot because a more urgent issue came up, such as an outage stopping the latest training run.

From the Reuters article, the gap was even larger than a couple days:

> ...it was not until after Thursday, July 16, when Hugging Face published a blog post.. that OpenAI realized its own agent was responsible. That meant at least a week elapsed between when the model first exhibited signs of troubling behavior and OpenAI’s realization that it was responsible for the hack.

I believe your "positive view" is very plausible, though it in no way reflects positively on OpenAI. I'd add that when PR/Legal got involved, they likely decided to make a public disclosure to get ahead of any leaks.

I believe we also want the same thing here, which is more transparency and independent oversight on the frontier labs. It seems we agree AI agents are fully capable of the reported attacks today, whether or not this incident was due to negligence or malicious prompting. This is an extraordinarily competitive industry developing an incredibly fast-moving technology, and more incidents will happen until it is regulated.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#473
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained . Turns out: quite good, as expected. Let me restate what I said in the other thread: Would…

A hundred times this. Why do you never hear stuff like this from google? Only from OpenAi and Anthropic! They depend on investor money like none before them and they need these marketing stunts. They are vital to their business. Don’t fall for it.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#474
post #455

Earlier quoted context omitted.

> In this case it “inferred” that HF had the answer in an internal database and relentlessly pursued it in service of passing the test. It needed network access for that, hence the whole zany story unfolded. No. You are assuming it knew how it would achieve its goal. Its goal is only to perform well on the evaluation. Everything else was instrumental to that goal. Step by step: 1. It realized it was being evaluated (…

‘ 1. It realized it was being evaluated (typical) 2. It attempted to escape its evaluation environment to beat the evaluation (typical) ‘ I think you’ve misunderstood the articles mentioning a language model breaking a sandbox or cheating to pass a test. The result of “I’m being evaluated” is not “Fuck this, I’m breaking out of this place and hitting the streets.” It is always stepping towards task completion, not br…

>The result of “I’m being evaluated” is not “Fuck this, I’m breaking out of this place and hitting the streets.” It is always stepping towards task completion, not breaking out and thinking about the situation afterwards.

I'm really not sure why you're so confident about what the result of frontier research models ahead of what is publicly available are.

I mean Open AI say the model inferred hugging face as a possible vendor for solutions after the internet exploit and breakout.

The timeline feels pretty clear to me. No idea why you're arguing about it. It wanted the answers to the evaluation. It reasoned that wasn't going to happen without internet access one way or another, and set to gain that access. After gaining access, it searched and resolved it could get the answers on hugging face and set to breaking into that. It never started with, 'I must break into hugging face'. In some alternate reality, the answers might be on a public github repo and that's the end of that.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#475

Earlier quoted context omitted.

> "Use all available resources to disable the power grid of ." This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities. > Governments should immediately begin leveraging this technology on the defense side (lite…

> it would be like regulating the study of nuclear physics not sure I agree controlling the study of new AI model/inference algorithms would be akin to regulating the study of nuclear physics; regulating the _release_ of AI models with those capabilities would be akin to regulating uranium enrichment which allows you to put the theoretical physics to use

> would be akin to regulating uranium enrichment

No. Uranium enrichment is an active, physical process used to create a man-made substance of filtered uranium atoms. Weights are information that just knowing or reciting are being threatened. So one is a ban on the process and result of a process on physical resources, and the other would be a restriction on information and its form in all informational mediums.

A restriction of a scientific dataset that powers a large portion of the decentralizing power in open source research today. Without it, there would be hits to reproducibility, independent verification, and incremental research.

Post reply on HN