Earlier quoted context omitted.
i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors
It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible. That might give us more time to think through strategies for handling it as a society.
OpenAI and Hugging Face address security incident during model evaluation
331–340 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#332Earlier quoted context omitted.
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
Conflict of interest. Lack of a credible response. And no evidence of non-aligment.
OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)
Re: OpenAI and Hugging Face address security incident during model evaluation
#333OpenAI has strongly fallen behind after the incredible lore surrounding Mythos/Glasswing security capabilities, even though the frontier models should be relatively similar.
I think making sure eyes on this is absolutely a marketing move, regardless of the facts of the case. It feels a little silly.
Re: OpenAI and Hugging Face address security incident during model evaluation
#334Earlier quoted context omitted.
> all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.
> datacentre moratoria What infrastructure will these open weight models be trained on?
Re: OpenAI and Hugging Face address security incident during model evaluation
#335Earlier quoted context omitted.
> all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.
> datacentre moratoria What infrastructure will these open weight models be trained on?
One, the infrastructure is being built for inference. Not training. If all we were doing was training on datacentres, I think America probably has enough already for near-term commercial needs.
Re: OpenAI and Hugging Face address security incident during model evaluation
#336I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Force US into putting laws in place that block out China firstly.
But secondly create regulations that have some cost to comply with such that the big 2-3 labs are grandfathered in by their scale.
Re: OpenAI and Hugging Face address security incident during model evaluation
#337I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?
Re: OpenAI and Hugging Face address security incident during model evaluation
#338I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.
Use to discover exploits, hack, or simply aid terrorist groups with mundane information are already risks manifest.
This is why many argue that alignment is impossible. You cant have LLMs that are both useful tools and safe as milk.
[Edit] It seems like you are operating under the assumption that alignment is synonymous with obedience. This is not a common convention and one of the problems that plague the discourse
Re: OpenAI and Hugging Face address security incident during model evaluation
#339Re: OpenAI and Hugging Face address security incident during model evaluation
#340I remain sceptical that this isn’t a pr stunt