Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

331–340 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#331

Earlier quoted context omitted.

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible. That might give us more time to think through strategies for handling it as a society.

MY cynical take: Until the compute needs get so enormous that only governments can fund it and there is a consensus internationally, its either company A in country X or company B in country Y. And since everyone thinks THEY are the good guys the competition will continue.

Re: OpenAI and Hugging Face address security incident during model evaluation

#332

Earlier quoted context omitted.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

> Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)

Re: OpenAI and Hugging Face address security incident during model evaluation

#333
I don't think this is fiction, but it's pretty clearly a marketing-release rather than a normal security disclosure.

OpenAI has strongly fallen behind after the incredible lore surrounding Mythos/Glasswing security capabilities, even though the frontier models should be relatively similar.

I think making sure eyes on this is absolutely a marketing move, regardless of the facts of the case. It feels a little silly.

Re: OpenAI and Hugging Face address security incident during model evaluation

#334

Earlier quoted context omitted.

> all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.

> datacentre moratoria What infrastructure will these open weight models be trained on?

Chinese infrastructure, presumably.

Re: OpenAI and Hugging Face address security incident during model evaluation

#335

Earlier quoted context omitted.

> all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.

> datacentre moratoria What infrastructure will these open weight models be trained on?

> What infrastructure will these open weight models be trained on?

One, the infrastructure is being built for inference. Not training. If all we were doing was training on datacentres, I think America probably has enough already for near-term commercial needs.

Re: OpenAI and Hugging Face address security incident during model evaluation

#336

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

I think the US labs are going with scare marketing as a regulatory moat.

Force US into putting laws in place that block out China firstly.

But secondly create regulations that have some cost to comply with such that the big 2-3 labs are grandfathered in by their scale.

Re: OpenAI and Hugging Face address security incident during model evaluation

#337

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#338

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

It really hinges on what you consider alignment and risk. For the widest definitions of alignment, we have never had an aligned model - One that will refuse to break the law or work against another persons interests.

Use to discover exploits, hack, or simply aid terrorist groups with mundane information are already risks manifest.

This is why many argue that alignment is impossible. You cant have LLMs that are both useful tools and safe as milk.

[Edit] It seems like you are operating under the assumption that alignment is synonymous with obedience. This is not a common convention and one of the problems that plague the discourse

Post reply on HN