Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

541–550 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#541
post #436

I wonder how many more high profile incidents some of you need before you stop insisting that this is all just marketing. Is it going to take Chinese companies also talking about contributing to long standing math problems and accidental sandbox escapes? Or is that also going to be interpreted as some conspiracy?

To clarify a little, I don't doubt that a decent portion of this story is embellished to make it sound more impressive/shocking than it was.

Yet even if we dismiss the drama as marketing (say, the sandbox intentionally left holes, the zero days weren't actually zero days, even that huggingface was in on it and the model was instructed to break in to a system), we're left with a model that seemingly broke into another company's servers.

Re: OpenAI and Hugging Face address security incident during model evaluation

#542

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.

This isn’t escaping in the same sense- the model was executing within the OpenAI infra. If it ported its entire architecture/weights into a public cloud to survive being turned off… that’d be pretty cool.

Re: OpenAI and Hugging Face address security incident during model evaluation

#543

Earlier quoted context omitted.

I don't think this is a paperclip factory moment. IIUC, it's an agent whose job it is to identfy and abuse exploits and that's exactly what it went off and did. The problem isn't anything AI specific, the problem is OpenAI's incompetence in their research leading to a lab leak. Just incompetence demanding regulation.

Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability. So going to find the Vulnerability's description on a third party website is clear cut reward hacking

> So going to find the Vulnerability's description on a third party website is clear cut reward hacking

that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.

Re: OpenAI and Hugging Face address security incident during model evaluation

#544

Earlier quoted context omitted.

What would compelling evidence look like to you?

> What would compelling evidence look like to you? I'm not sure. I trusted the labs when they first raised the alarms. But then we got a series of boys-who-cried-wolf. So at this point I want to see evidence of actual, novel harm that results in concrete damage.

Right. But I think the fear is that if we wait for this type of evidence, it will be too late.

Re: OpenAI and Hugging Face address security incident during model evaluation

#545

>All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal. ExploitGym is a literal exploit dev benchmark. As always, the entire event looks a lot more like "the model did what we prompted it to do" than "it decided to do this spontaneously on its own".

It did deliver the message to Garcia.

Re: OpenAI and Hugging Face address security incident during model evaluation

#546
post #475

Isn't this a crime that someone is liable for? What happened is that someone hacked into a computer system without permission. Maybe it wasn't intentional -- sure -- and that would be a factor at sentencing. But it sounds like they've admitted to a crime, and obviously our legal system considers the humans involved to be the liable parties; otherwise everyone would just say "my computer did the hacking" and wouldn't…

Lifting mens rea on that is going to be... interesting.

Re: OpenAI and Hugging Face address security incident during model evaluation

#547

It just feels deeply unserious that these labs talk about apocalyptic risks, ship models with safeguards that make them borderline useless for sensible tasks, and then YOLO stuff like that on the backend and use it as an opportunity to market their stuff some more.

100%, imo risks from internal deployment will eventually be the biggest risks, and keeping models heavily gated/not accessible just makes these risks much worse.

Re: OpenAI and Hugging Face address security incident during model evaluation

#548
post #475

Isn't this a crime that someone is liable for? What happened is that someone hacked into a computer system without permission. Maybe it wasn't intentional -- sure -- and that would be a factor at sentencing. But it sounds like they've admitted to a crime, and obviously our legal system considers the humans involved to be the liable parties; otherwise everyone would just say "my computer did the hacking" and wouldn't…

Most crimes require intent, hacking is one of them. The relevant law in this situation is:

> (a) Whoever— (2) intentionally accesses a computer without authorization or exceeds authorized access, and thereby obtains— (C) information from any protected computer; shall be punished as provided in subsection (c) of this section.

https://www.law.cornell.edu/uscode/text/18/1030

So if it can't be proven that you intended to access a computer without authorization, or exceed your authorized access, then you can't be found guilty of the crime.

Consider the possible consequences of the law not requiring intent, if simply accidentally exceeding your authorized access could be a criminal act.

Re: OpenAI and Hugging Face address security incident during model evaluation

#549
post #533

Earlier quoted context omitted.

Well they're doing a pretty poor job of guiding it in a positive direction and ethically speaking they are almost indistinguishable from the bad actors....

You really don't see much difference between Anthropic and xAI? Anthropic is safety testing their models and drawing hard (though minimal) boundaries against the DoD. xAI is building racist pornbots. Sure Anthropic is not perfect. But it's a coordination problem. They're in a race and safety/restraint is a handicap. That's why they're begging for regulation (and just get accused of attempting regulatory capture.) Why…

Who said anything about xAI here? That's not the bad actor people think of. Honestly I wouldn't be shocked if your post is intended to derail the discussion away from criticism of the industry.

Re: OpenAI and Hugging Face address security incident during model evaluation

#550

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Because it can make a small number of people really rich. That's all that matters.

Post reply on HN