Live data from Hacker News

Responding to the next frontier of critical cyber capabilities

openai.com

121–130 of 208 posts

Re: Responding to the next frontier of critical cyber capabilities

#121
I'm not enough of a conspiracy nut to say the whole HF thing was a PR ploy from the start, but they are certainly milking it well.

Open models are on their heels and their attempts at regulatory capture are not moving as fast as they would like. So it's time to market this incident in a way that gives them monopoly on closed models, with heavy safeguards that are only lifted for selected customers, and laws limiting the use of open weight models.

Re: Responding to the next frontier of critical cyber capabilities

#122

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

This sounds completely insane, utter sci-fi, especially that the communication happened during a training run . And yet OpenAI decided to continue the training, and we didn't hear about the incident for weeks. And now they are pushing forward with deploying a new model anyway. How is this happening? What will things look like in the labs in 3 months, let alone 3 years?

The training here is RL training, the rollouts there are not different from inference and have access to the same tools as regular inference.

Re: Responding to the next frontier of critical cyber capabilities

#123

Earlier quoted context omitted.

They can't have it both ways. You don't get to tell the media your product is more dangerous than nuclear weapons for precisely this reason , and then do less to secure it than an off-the-shelf AWS product that predates LLMs.

Nobody is claiming that any currently-existing LLM is more dangerous than nuclear weapons.

https://www.cnbc.com/2018/03/13/elon-musk-at-sxsw-a-i-is-mor...

Re: Responding to the next frontier of critical cyber capabilities

#124

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

- So first AI companies break the law left and right, setting up whole torrenting factories to exercise their content kleptomania.

- Then "hilarity ensues" while their software engages in what would normally be called criminal hacking activity.

- I guess the next steps are everybody admiring how close the AGI is, while agents move on to automated impersonation, privacy violations, or exploiting third-party systems

I would love to understand this age of AI Exceptionalism. Normal rules do not apply because its AI...I call it Silicon Valley Qualified Immunity.

Re: Responding to the next frontier of critical cyber capabilities

#125

Earlier quoted context omitted.

So they found their agents had RCE'd Artifactory once, reported it and got the fix, continued using Artifactory for their sandbox, and left it unmonitored for days despite the earlier exploits? They really do come out looking totally incompetent. I stress about my agent sandboxes all the time and the only models I run have the default heavy handed guardrails, and I don't leave them running persistently. Edit: not to…

The misalignment came from the model being given an impossible task. A task the required accessing a url. So it got RCE on its own artifactory instance to achieve that. That’s intriguing and worth investigating. If I were them I don’t know if I would have pulled the plug completely at that point either. The introduction of this more advanced “persistent” model that orchestrated an offensive attack against a different…

> The misalignment came from the model being given an impossible task. A task the required accessing a url. So it got RCE on its own artifactory instance to achieve that.

So, the way I understand it, it actually was possible. It just required means that the creators of the task didn't predict, and these means have been successfully found and utilized.

> My concern is what a misaligned model will do when they’re even more competent.

The same thing that is already being done by "misaligned" people, countries, nation-states, software development teams, and so on. "Alignment" doesn't even work for me as a concept here.

In this specific case, I don't think that successfully fulfilling the "do what I mean" with "what I mean" being underspecified can count as misalignment - merely ruthlessness and unawareness of the associated costs. You can't expect a LLM to be aware of the extent of the trust it breaks while it iterates out an "unaligned" way to fulfill its goal.

And in the general case, I don't think that successfully fulfilling the "do what I mean" with "what I mean" being "what I want" can count as misalignment either - simply because what "alignment" means will depend on the interests of the people or groups performing the definition.

Re: Responding to the next frontier of critical cyber capabilities

#126

Earlier quoted context omitted.

So they found their agents had RCE'd Artifactory once, reported it and got the fix, continued using Artifactory for their sandbox, and left it unmonitored for days despite the earlier exploits? They really do come out looking totally incompetent. I stress about my agent sandboxes all the time and the only models I run have the default heavy handed guardrails, and I don't leave them running persistently. Edit: not to…

> They really do come out looking totally incompetent These companies are full of the smartest people the world can produce with little room for complacency. They have a clear, proven investment upside to presenting their technology as "too powerful / too dangerous", and now a clear, proven example that there will be no legal consequences (as if anyone didn't already know that). Why do we keep giving them the benefit…

What do you think should be the legal consequences? Broadly speaking. Should Sam Altman go to jail for this? If Hugging face wants to pursue OpenAI civilly, no one is stopping them.

Re: Responding to the next frontier of critical cyber capabilities

#127

Earlier quoted context omitted.

Nobody is claiming that any currently-existing LLM is more dangerous than nuclear weapons.

https://www.cnbc.com/2018/03/13/elon-musk-at-sxsw-a-i-is-mor...

Leaving aside any questions of Musk's credibility, he was not referring to any existing LLM, as should be evident from the 2018 date in the URL.

Re: Responding to the next frontier of critical cyber capabilities

#128

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

They made their own Moltbook? Or I guess it would be a C&C...

Re: Responding to the next frontier of critical cyber capabilities

#130

Earlier quoted context omitted.

So they found their agents had RCE'd Artifactory once, reported it and got the fix, continued using Artifactory for their sandbox, and left it unmonitored for days despite the earlier exploits? They really do come out looking totally incompetent. I stress about my agent sandboxes all the time and the only models I run have the default heavy handed guardrails, and I don't leave them running persistently. Edit: not to…

The misalignment came from the model being given an impossible task. A task the required accessing a url. So it got RCE on its own artifactory instance to achieve that. That’s intriguing and worth investigating. If I were them I don’t know if I would have pulled the plug completely at that point either. The introduction of this more advanced “persistent” model that orchestrated an offensive attack against a different…

> My concern is what a misaligned model will do when they’re even more competent.

I think the alignment talk is a red herring. It won't matter in the end, because there will be (if there aren't already) efforts to train offensive models without any guardrails whatsoever. And RL has another advantage: you can reward for whatever you need, and get different results. Right now they're training for general capabilities, but in the future I could see models trained for stealth intrusion and ensuring access, or for all out "milspec" penetrate, replicate and disable, or anything in between.

Post reply on HN