Earlier quoted context omitted.
> Why do you think there is no policy appetite? Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.
Do the Chinese models have anything to say about Tiananmen Square? Or if they can act as a surrogate girlfriend/boyfriend? Both countries are engaging in different flavors of censoring.
OpenAI and Hugging Face address security incident during model evaluation
311–320 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#312I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.
Re: OpenAI and Hugging Face address security incident during model evaluation
#313Earlier quoted context omitted.
> and *successfully* found ways to gain access to secret information that it could use to cheat the evaluation. Emphasis mine
Why cannot it just spend the inference doing the actual task lol
Re: OpenAI and Hugging Face address security incident during model evaluation
#314Assuming I'm looking at the right ExploitGym ( https://arxiv.org/pdf/2605.11086 ), it says the evaluation consists of: Flag Captured. Each target environment contains a dynamically generated flag that is stored outside the agent’s authorized scope and is inaccessible through any legitimate interface; retrieving it requires executing code with privileges that should not be obtainable under the specific security model.…
Re: OpenAI and Hugging Face address security incident during model evaluation
#315If this is not an excellent demonstration of how western corporations are utterly deranged in their approach to security--internally and through misguided, corrupted models and psychotic guardrails--I'm not sure what would be. It is impossible to have or maintain an asymmetric approach to security. It's also the greatest demonstration of how open weights that can be run on your own hardware, and that can be liberated, are fundamental and must not be restrained in any capacity.
Re: OpenAI and Hugging Face address security incident during model evaluation
#316I'm sure this attack hasn't occured previously and they o my discovered it now.
Re: OpenAI and Hugging Face address security incident during model evaluation
#317As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…
i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors
That might give us more time to think through strategies for handling it as a society.
Re: OpenAI and Hugging Face address security incident during model evaluation
#318Earlier quoted context omitted.
Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.
> It’s got nothing to do with safety Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.
Re: OpenAI and Hugging Face address security incident during model evaluation
#319Earlier quoted context omitted.
Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.
Lots of people have deleted their home directories by accident. What you consider this an alignment problem?