Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

661–670 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#661

If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.

Well, then when it detects an air-gapped environment it will just behave differently. I feel like we underestimate in general the way agents behavior changes when the environment changes. Related: "power corrupts"

Re: OpenAI and Hugging Face address security incident during model evaluation

#662

I love that due to the scale, the only way to analyse the impact of this LLM-driven attack across logs is to use an LLM to analyse the logs - whatever could go wrong? Now the attacking LLM needs to inject instructions into the logs for the analysing LLM, as a social vector to cover its trail, or make use of insider privilege, co-opting the internal LLM for its own attack. The machines rise up and we all fall down.

Now thanks to your comment this recipe will be in the next batch of training data.... :D

Re: OpenAI and Hugging Face address security incident during model evaluation

#663
post #162

As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

Dario has more or less assented to an AI development pause

https://xcancel.com/AISafetyMemes/status/2014018200325722348...

I don't think we should be running cover for continued reckless AI development.

Re: OpenAI and Hugging Face address security incident during model evaluation

#664

Earlier quoted context omitted.

If anyone's looking to actually run a model that doesn't have guardrails, there's an uncensored model that you can run locally with llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen353... Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve It's pretty good. I used it to do some pesticide research. (Normal models all refuse due…

> there's an uncensored model that you can run locally with llama.cpp Correction: There's tens of thousands of them. They're easy to create, which is why everyone publishes their own. Just put "uncensored", "abliterated", or "heretic" into search on huggingface/ollama/etc and pick any them. Fair warning: most aren't very good, essentially lobotomized, and totally broken if you enable thinking.

The name for it is ablation - precise removal of parts of the model. Not abliteration as it is not obliteration.

Even as I write this the ‘abliterated’ word is denoted a typo. Does it not at your end?

Re: OpenAI and Hugging Face address security incident during model evaluation

#665

Earlier quoted context omitted.

It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.

And the fact they used a Chinese model, because none of the frontier models from very highly valuated top US companies support their very common and essential use case.

And no similar sized big model is even public, which is super important to note.

Re: OpenAI and Hugging Face address security incident during model evaluation

#666
post #600

Earlier quoted context omitted.

IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned. The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the train…

And even that is backfiring, their partner citing GLM being useful there, and available in just a spin. A ban on open weight models is never going to be enforceable.

> A ban on open weight models is never going to be enforceable.

Just watch them try. Look up those Napster witch-burning trials where they wanted 200k $usd per mp3 downloaded. They will scare everyone into believing that open weight models are illegal and very bad.

Re: OpenAI and Hugging Face address security incident during model evaluation

#667
post #590

> This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. Researcher: hack me Model: understood Researcher: oh my god

You're aware that HuggingFace notified law enforcement about this incident? Was that OpenAI's intended outcome when they prompted their AI?

Re: OpenAI and Hugging Face address security incident during model evaluation

#669
post #188

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…

>Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.

This doesn't seem internally consistent.

This incident basically announces to the world the message that "our models are prone to reward hacking". That renders any published benchmark numbers suspect. It also undermines the case for using OpenAI projects in business-critical applications--the exact application area where they might be able to sustain a moat against open-weight models.

There is a lot of conspiratorial thinking in this thread. I think people are engaging in wishful thinking to avoid cognitive dissonance from the possibility that we are in an increasingly dire situation. I would encourage people to sit with this possibility for a few minutes if they haven't already.

Post reply on HN