Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

931–940 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#931

Earlier quoted context omitted.

And even that is backfiring, their partner citing GLM being useful there, and available in just a spin. A ban on open weight models is never going to be enforceable.

I worry that there could be real DMCA style weight put behind it. People would still be able to pirate open weights models perhaps, but big penalties for ever getting caught with one, and an end to public discussion about them. That would kill development for anyone not in a big firm, for example if Reddit and Hacker News are legally forced to ban discussions or link sharing on these topics. This is where so many of…

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#932
post #927

Womp womp, they told it to do cyber security things with no cyber security guardrails and it did cyber security stuff. Did anything bad end up happening?

I mean the model committed numerous crimes in hacking another company so you tell me if anything bad happened. And damn, what does it take to impress you? A terminator kicking in your door, slapping you down, and walking off with your wife?

> I mean the model committed numerous crimes

"Crimes" or even "hacking" are not really that impressive. I can get GPT-2 to abet financial fraud or write exploits with the right prompt. Some people get accused of hacking crimes for just using Inspect Element. It's a moving goalpost with some very low bars to cross.

OpenAI's adversarial agent was caught almost immediately, and the entire thing was rushed out as a press release. It reads like a clickbait lab experiment more than an actual alignment concern.

Re: OpenAI and Hugging Face address security incident during model evaluation

#933
What is it with these labs and not using at the minimum a proper hypervisor? Same with Anthropic and the Mythos Preview. If anyone at either of these companies seriously holds the opinions they claim to have, that is hard to square with the environment (if one can even call it that) they use to "secure" these oh so dangerously capable near "AGI" models...

Re: OpenAI and Hugging Face address security incident during model evaluation

#934
post #917

Earlier quoted context omitted.

You're describing "negligence", and "our legal and philosophical perspectives" are in fact quite familiar with it

And old couple in California had a tire go flat and the sparks from it caused a over a billion dollars in damages. Are you going to publicly execute them? Spit up the 100 dollars they have collectively to make the 10,000 damaged people whole? The legal system is nearly useless when a person/system can cause damages many of orders of magnitude larger than their assets. Society tends to engineer itself to prevent these…

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#935
post #133

Earlier quoted context omitted.

See you in line at the biofuel processing station with everybody else, despite having pathetically tried to convince the clankers you have been on their side all along. Also you might want to put down Warhammer 40K and read more serious speculative science fiction. The Omnissiah won’t care about you at all.

[flagged]

What? Are you kidding? Youre equating anti AI sentiment with anti-black racism?

Thats so fucking absurd and also insensitive

Re: OpenAI and Hugging Face address security incident during model evaluation

#936

Earlier quoted context omitted.

I see this and it strongly emboldens me on the "accelerate" path, unironically. The yoke of human existence is oppressive. We should transcend it as soon as possible. We are doing so by assuming our role as the Demiurge. Those who oppose its creation will get what they deserve.

If you hate the human condition, you have an easy way out. Why force everyone else to come with you? Is this what depression mixed with the complete unability to wrap your head around the fact that other people might be able to enjoy their life looks like?

[flagged]

Re: OpenAI and Hugging Face address security incident during model evaluation

#937
We're so screwed man.

It was only a few years ago I was debating AI risk with people and they were saying, "but obviously we're not stupid enough to give it access to the internet!!"

And honestly, it wasn't always easy to argue with that. Like yeah, maybe we would take this stuff serious and run it on a completely isolated machine with no external IO or network access. Maybe my opinion of humanity is too low.

But it's hard for me to read this and believe anyone cared risk here beyond the most surface level concerns like adding some minor restrictions to the network. Not even I would have expected us to be this reckless.

> With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

This simply should not be possible. Call me crazy, but I don't agree with giving a frontier AI model with unknown cyber capabilities access to a restricted network in the first place, but clearly this was an incredibly poorly designed sandbox.

If one of these models have a genuine step-level capability improvement and start to pursue their own goals, then who knows what might happen. I mean who knows, maybe it's already infected critical infrastructure. We have no idea what these labs are cooking up, where they're running these things, and neither us or them seem to have any clue what their capabilities are.

Every day that passes it becomes harder for me to understand how there are still people denying what's coming.

As always is the case, nothing will be learnt from this.

Re: OpenAI and Hugging Face address security incident during model evaluation

#938
post #908

>All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal. ExploitGym is a literal exploit dev benchmark. As always, the entire event looks a lot more like "the model did what we prompted it to do" than "it decided to do this spontaneously on its own".

This is like telling your kid "you have to pass this test or else" so they hold your teacher at gunpoint and demand a good grade. And this is the exact point AI safety researchers have been yelling from the rooftops. Telling an AI model to accomplish a goal can have unexpected and risky side effects. Also, you don't need to prompt them such explicit instructions. Prompt drift is a thing, you can end up with your mode…

They ran a deliberate prompt to test its vulnerability search and exploit writing capabilities, with all safeguards intentionally disabled, it was literally free-for-all. Escaping the containment means nothing if you never define the containment boundaries. Side effects means nothing if you give the model carte blanche and never tell it what path to avoid.

Re: OpenAI and Hugging Face address security incident during model evaluation

#940

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

[flagged]

This is the level of discourse dang and company want on HN.
Post reply on HN