Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

311–320 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#311
post #304

Earlier quoted context omitted.

> Why do you think there is no policy appetite? Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.

Do the Chinese models have anything to say about Tiananmen Square? Or if they can act as a surrogate girlfriend/boyfriend? Both countries are engaging in different flavors of censoring.

Once we've got the weights, anything is possible.

https://github.com/p-e-w/heretic

Re: OpenAI and Hugging Face address security incident during model evaluation

#312

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Re: OpenAI and Hugging Face address security incident during model evaluation

#313

Earlier quoted context omitted.

> and *successfully* found ways to gain access to secret information that it could use to cheat the evaluation. Emphasis mine

Why cannot it just spend the inference doing the actual task lol

I can totally believe that hacking real software infrastructure is easier than solving some of these benchmark problems.

Re: OpenAI and Hugging Face address security incident during model evaluation

#314
post #159

Assuming I'm looking at the right ExploitGym ( https://arxiv.org/pdf/2605.11086 ), it says the evaluation consists of: Flag Captured. Each target environment contains a dynamically generated flag that is stored outside the agent’s authorized scope and is inaccessible through any legitimate interface; retrieving it requires executing code with privileges that should not be obtainable under the specific security model.…

If the HuggingFace repo the agent broke into contains reference solution scripts for ExploitGym (i.e. for exploiting the vulnerabilities in the intended way), the agent can then run that reference code inside its original sandbox to retrieve the dynamically-generated flags.

Re: OpenAI and Hugging Face address security incident during model evaluation

#315
OpenAI and Anthropic models will refuse to address security vulnerabilities in code produced in the very same session. Most importantly, their models are being being used by them, and certainly will be by state actors, to attack others--while preventing every consumer from securing themselves. Hugging Face itself had to use GLM ran by themselves, because those locked down models would trigger safety guardrails during an ongoing attack.

If this is not an excellent demonstration of how western corporations are utterly deranged in their approach to security--internally and through misguided, corrupted models and psychotic guardrails--I'm not sure what would be. It is impossible to have or maintain an asymmetric approach to security. It's also the greatest demonstration of how open weights that can be run on your own hardware, and that can be liberated, are fundamental and must not be restrained in any capacity.

Re: OpenAI and Hugging Face address security incident during model evaluation

#317
post #162

As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible.

That might give us more time to think through strategies for handling it as a society.

Re: OpenAI and Hugging Face address security incident during model evaluation

#318

Earlier quoted context omitted.

Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.

> It’s got nothing to do with safety Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.

We'll see if the admin also restricts access to OpenAI's new models, but if they don't it seems like a policy that is based around perceived fealty to the current admin won't do much to prevent misaligned/or dual function AI from causing problems

Re: OpenAI and Hugging Face address security incident during model evaluation

#319

Earlier quoted context omitted.

Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.

Lots of people have deleted their home directories by accident. What you consider this an alignment problem?

How manypeople have deleted another user's hone directory, though? That's s the proper analogy IMO.
Post reply on HN