To invent another reason to ban powerful Chinese open weight models.
OpenAI and Hugging Face address security incident during model evaluation
461–470 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#462Is this a new kind of accountability backdoor?
Re: OpenAI and Hugging Face address security incident during model evaluation
#463If you’ve ever doubted the “paperclip maximizer” scenario, or doubted the Orthogonality Thesis, it’s time to put it to rest.
Re: OpenAI and Hugging Face address security incident during model evaluation
#464I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Re: OpenAI and Hugging Face address security incident during model evaluation
#465Earlier quoted context omitted.
Does the CFAA cover unintentional access without authorization?
No. "Intentionally", "willfully", or "knowingly" are prerequisite states of mind for crimes defined by the CFAA.
Re: OpenAI and Hugging Face address security incident during model evaluation
#466This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…
Re: OpenAI and Hugging Face address security incident during model evaluation
#467I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives. Sorry to bring the party down/be obstinate… I’m just a lil scared for the…
I think we desperately need some independent group to evaluate claims like this or the world-ending Mythos cybersecurity risk and tell us what’s going on.
Re: OpenAI and Hugging Face address security incident during model evaluation
#468Earlier quoted context omitted.
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
> Can you explain how the above event doesn't count as evidence alignment is an actual risk? Conflict of interest. Lack of a credible response. And no evidence of non-aligment. OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were…
Re: OpenAI and Hugging Face address security incident during model evaluation
#469If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.
> wildly negligent to not be running it in a physically-airgapped environment Why should it be physically airgapped? Clients won't be doing that.
Re: OpenAI and Hugging Face address security incident during model evaluation
#470Earlier quoted context omitted.
[flagged]
I don't understand this sentiment at all. Is it a claim that "breaking into Hugging Face's production infrastructure" didn't happen? That it's not actually all that severe? That it was done by hand by OpenAI employees and they fooled Hugging Face? That the blog post exaggerates something, somehow? What exactly do you mean? At the moment it just reads like a thoughtless dismissal.