OpenAI and Hugging Face address security incident during model evaluation
351–360 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#352Absolutely bewildering. If I am building a giant cannon and blow a hole straight through my neighbor’s house, I’m not going to say “we are working with our neighbors to improve their giant cannon defenses”. OpenAI brought this weapon and as far as I’m concerned they used it on another party. Morally it probably matters that this happens because they don’t know how their weapon works. Legally I always thought it was i…
It's an interesting point, but this is more like we are building a giant autonomous canon, that escaped the lab, the testing range, defeated state of the art and serious security protocols, and then blew a hole in the neighbors house. Our legal and philosophical perspectives are deeply rooted in humans being the actors. Doing that in a residential home is unforgiveable. Doing it responsibly on a military range is exp…
Re: OpenAI and Hugging Face address security incident during model evaluation
#353Each time Anthropic would do their nonsense to get headlines about how theoretically dangerous their models were - like when they claimed a model blackmailed someone with emails showing he was cheating, but they basically pushed it as much as possible to do as such - it got me more and more worried. Because eventually it's going to be a boy-who-cried-wolf situation where scary stuff really does start happening but pe…
Headline? It was buried in a model card. They just honestly report not-quite-incident because it's quite close to the incident OpenAI had. Nothing wrong with it.
They know what they're doing. It's a playbook. You write scary stuff in the model card to make it look like legitimate whitepaper rEsEarCh, then drip-feed it to the media outlets who make it a headline story. Fear based marketing is the hot trend of the 2020s.
But also, they write literal headlines: https://www.anthropic.com/research/agentic-misalignment
Re: OpenAI and Hugging Face address security incident during model evaluation
#354Re: OpenAI and Hugging Face address security incident during model evaluation
#355Earlier quoted context omitted.
Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.
Lots of people have deleted their home directories by accident. What you consider this an alignment problem?
Re: OpenAI and Hugging Face address security incident during model evaluation
#356Each time Anthropic would do their nonsense to get headlines about how theoretically dangerous their models were - like when they claimed a model blackmailed someone with emails showing he was cheating, but they basically pushed it as much as possible to do as such - it got me more and more worried. Because eventually it's going to be a boy-who-cried-wolf situation where scary stuff really does start happening but pe…
I mean, does it have to be one or the other? Just because it's actually dangerous doesn't mean nobody in OpenAI considers it great PR. And just because there are people in OpenAI that consider it great PR doesn't mean it isn't dangerous.
Re: OpenAI and Hugging Face address security incident during model evaluation
#357Earlier quoted context omitted.
This is marketing. Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?
It’s marketing the same way shitting your pants in public is marketing. People notice you.
Re: OpenAI and Hugging Face address security incident during model evaluation
#358Re: OpenAI and Hugging Face address security incident during model evaluation
#359Earlier quoted context omitted.
> Why do you think there is no policy appetite? Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.
How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop? During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.
You can ask them, they live in China, not Narnia. I spend about two months in the country per year mostly for tech/work related reasons and I've not encountered that sentiment. For one they don't have these borderline religious schizophrenic breakdowns thinking they're bringing about the end of the world, most people just see this tech for what it is, a tool for productivity and automation like any other piece of software and they don't actually think about the US. They're competing first and foremost for Chinese customers, with each other, maybe some old CCP guy cares about America, the 20/30 something's care about competing with other Chinese companies for users.
Re: OpenAI and Hugging Face address security incident during model evaluation
#360I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.
You can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem.