Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

461–470 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#462
If an individual did this, a massive CFAA hammer would be falling on their heads. Even though it doesn’t seem to be the case, OpenAI could’ve been trying to hack into HF and blame it on their models.

Is this a new kind of accountability backdoor?

Re: OpenAI and Hugging Face address security incident during model evaluation

#464

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

this doesn't really matter. There's no risk of models gaining sentience and running themselves, this blog is like openai saying whoops we ran sqlmap and dumped hf. cool, but someone still needs to point the gun

Re: OpenAI and Hugging Face address security incident during model evaluation

#465

Earlier quoted context omitted.

Does the CFAA cover unintentional access without authorization?

No. "Intentionally", "willfully", or "knowingly" are prerequisite states of mind for crimes defined by the CFAA.

The agent did it intentionally and willfully and knowingly. But you can’t sue the agent, I suppose. And the human didn’t ask the agent to do so.. so not a problem? Or the legislation needs an update?

Re: OpenAI and Hugging Face address security incident during model evaluation

#466

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

I don't think this is a paperclip factory moment. IIUC, it's an agent whose job it is to identfy and abuse exploits and that's exactly what it went off and did. The problem isn't anything AI specific, the problem is OpenAI's incompetence in their research leading to a lab leak. Just incompetence demanding regulation.

Re: OpenAI and Hugging Face address security incident during model evaluation

#467
post #279

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives. Sorry to bring the party down/be obstinate… I’m just a lil scared for the…

The problem is that the people telling us about these things are the same people that benefit from their model (and AI generally) being used, getting publicity, etc.

I think we desperately need some independent group to evaluate claims like this or the world-ending Mythos cybersecurity risk and tell us what’s going on.

Re: OpenAI and Hugging Face address security incident during model evaluation

#468

Earlier quoted context omitted.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

> Can you explain how the above event doesn't count as evidence alignment is an actual risk? Conflict of interest. Lack of a credible response. And no evidence of non-aligment. OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were…

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#469
post #444

If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.

> wildly negligent to not be running it in a physically-airgapped environment Why should it be physically airgapped? Clients won't be doing that.

Because they are testing it and are expected to erect guardrails before releasing.

Re: OpenAI and Hugging Face address security incident during model evaluation

#470
post #442

Earlier quoted context omitted.

[flagged]

I don't understand this sentiment at all. Is it a claim that "breaking into Hugging Face's production infrastructure" didn't happen? That it's not actually all that severe? That it was done by hand by OpenAI employees and they fooled Hugging Face? That the blog post exaggerates something, somehow? What exactly do you mean? At the moment it just reads like a thoughtless dismissal.

My thought is they might have set up the environment sloppily because they knew this could have led to something like this happening.
Post reply on HN