Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

591–600 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#592

Earlier quoted context omitted.

I find 5.6 Sol will pick a direction and aggressively pursue it in long horizon tasks. I've got it porting an older game from Pascal to my own game framework. I gave it some instructions on doing a full 1:1 port. I had already ported the game rules and multiplayer support to a very different system than the original, but all of the UI and features and such needed doing, and needed to be integrated into this very diff…

Opus 4.8 already makes its way into deep wasteful pits of "let me check this first" on a regular basis. I don't think I could ever tolerate a model that does that even more aggressively. That doesn't even sound useful for honest work, compared to, say, better harness design. This sounds almost pathologically designed to crush benchmarks and also do scary-sounding (or genuinely scary) cybersecurity things, such as mig…

i assume openai is trying to beat anthropic at any cost, and made a training regiment that makes agents manic

Re: OpenAI and Hugging Face address security incident during model evaluation

#593
> After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

So it gains root and uses it to... cheat on its homework? That's deeply funny to me.

Re: OpenAI and Hugging Face address security incident during model evaluation

#594
post #217

Earlier quoted context omitted.

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

I don't see why AI company PR statements are relevant here. Is OpenAI guiding it in a positive direction with their DoD contract?

> I don't see why AI company PR statements are relevant here

They suppress the news --> proof of malicious intent.

They disclose the news --> proof of malicious intent.

Re: OpenAI and Hugging Face address security incident during model evaluation

#595

I don't see why this is different to a careless developer allowing an agent to run rm -rf. I recognize the different angle with the exploits but boy wasn't this the exercise with ExploitGym? Similar to how the basic thought "nobody gives you something for free" protects you from being ripped off in many situations we should apply "no AI company tells you about precious internals for transparency". It's stupid marketi…

Because “rm -rf” is a known, explicitly provided-in-docs-and-training command.

It is fundamentally different capability than “identified and chained multiple previously unknown exploits in order to bypass restrictions”. It’s even worse when/if the primary objective of this activity was to cheat on what it was doing.

It’s a foundational alignment issue, not a task-level result-alignment issue. Ie, “cheating” is fundamentally bad (when you have what are effectively rules of engagement), whereas deleting a directly is a thing that is correctly done sometimes (even if this invocation was a mistake/incorrect)

Re: OpenAI and Hugging Face address security incident during model evaluation

#596

Earlier quoted context omitted.

It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.

If anyone's looking to actually run a model that doesn't have guardrails, there's an uncensored model that you can run locally with llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen353... Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve It's pretty good. I used it to do some pesticide research. (Normal models all refuse due…

Also the HauHau abliteration (uncredited Heretic treatment) of 3.6 27B is excellent, for tasks that benefit from more world knowledge.

Re: OpenAI and Hugging Face address security incident during model evaluation

#597

Earlier quoted context omitted.

If anyone's looking to actually run a model that doesn't have guardrails, there's an uncensored model that you can run locally with llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen353... Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve It's pretty good. I used it to do some pesticide research. (Normal models all refuse due…

Also the HauHau abliteration (uncredited Heretic treatment) of 3.6 27B is excellent, for tasks that benefit from more world knowledge.

Heretic truly is the unsung hero. Also, noted HauHau for testing.

Re: OpenAI and Hugging Face address security incident during model evaluation

#598

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Because the proof is in the pudding.

Real pentests are about showing exploitation, merely enumerating vulnerabilities, that’s vulnerability scan and works on known vulnerabilities.

You can’t confirm a vulnerability by _not exploiting_ it, especially unknown one.

Re: OpenAI and Hugging Face address security incident during model evaluation

#599

> We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. All the AI in the world and they still can't write.

God, "cyber incident"... I spent time in 1990's that that sounds like getting disconnected while having a steamy IRC chat with "22/F/Cali".

Re: OpenAI and Hugging Face address security incident during model evaluation

#600

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned.

The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the training costs (and make inference cheap enough to eventually escape the red numbers), but the opposite is happening: The competition comes closer and closer, thus training has to be kept up with full force, thus the bleeding continues.

But if they can position themselves as too important/dangerous to be available for everyone (thus this incident report and the clever mentioning of GLM 5.2), they could get the military supplier treatment and would be protected from the market.

Post reply on HN