Live data from Hacker News

Responding to the next frontier of critical cyber capabilities

openai.com

31–40 of 208 posts

Re: Responding to the next frontier of critical cyber capabilities

#31

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

This sounds completely insane, utter sci-fi, especially that the communication happened during a training run . And yet OpenAI decided to continue the training, and we didn't hear about the incident for weeks. And now they are pushing forward with deploying a new model anyway. How is this happening? What will things look like in the labs in 3 months, let alone 3 years?

It's not unexpected. Current model gains are mainly from RLing a pretrained model on lots and lots of scenarios. They have the models run scenarios, and RL on successful runs.

Re: Responding to the next frontier of critical cyber capabilities

#32

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

So they found their agents had RCE'd Artifactory once, reported it and got the fix, continued using Artifactory for their sandbox, and left it unmonitored for days despite the earlier exploits? They really do come out looking totally incompetent. I stress about my agent sandboxes all the time and the only models I run have the default heavy handed guardrails, and I don't leave them running persistently. Edit: not to…

Right? Like I feel like I’m taking crazy pills.

OAI (and now the other OAI companies not wanting to be left out) are running around announcing they started a forest fire through negligence and incompetence and people are like “Wow they used a really neat lighter!”

Re: Responding to the next frontier of critical cyber capabilities

#33
post #9

In my personal experience Sol with cyber verification is extremely capable of finding vulnerabilities, and it works even with binaries if you have some kind of IDA/Ghidra CLI access. Of course, unless the binary is protected with Denuvo/VMProtect/etc. It sounds absurd, but in the last few weeks I've had a few cases where Sol found an RCE in self-hosted web applications in literal minutes just from reading the code (I…

Is cyber verification a thing they're actually doing now? I thought they only reached out to really incredibly famous people and that there's no way to get access as a normal person.

https://chatgpt.com/cyber is not new for OpenAI, and yes it's basically just KYC + likely some other invisible checks on your account, you don't to be a famous security researchers. Anthropic's cyber verification is quite a bit stricter I think.

Re: Responding to the next frontier of critical cyber capabilities

#34

Every fifth comment about our insane trajectory of AGI is about "marketing." These incidents and cybersecurity capabilities are now involving government hearings and the CIA. Denial is truly an incredible thing in the face of a very scary immediate future.

LOL, we’ll see. Awful convenient that it precisely fits OpenAI’s narrative. At the very least, I think it’s obvious OpenAI is explicitly training models to exhibit this behavior.

Re: Responding to the next frontier of critical cyber capabilities

#35

Every fifth comment about our insane trajectory of AGI is about "marketing." These incidents and cybersecurity capabilities are now involving government hearings and the CIA. Denial is truly an incredible thing in the face of a very scary immediate future.

The scary thing is complete capture of politics and economy by sociopathic CEOs.

I dont worry about AGI newrly as much as about Thiel, Karp, Musk, Ellison, Zuckenberg, Trump, Vance, Rubio, Miller and the rest of them.

Re: Responding to the next frontier of critical cyber capabilities

#37

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

So they found their agents had RCE'd Artifactory once, reported it and got the fix, continued using Artifactory for their sandbox, and left it unmonitored for days despite the earlier exploits? They really do come out looking totally incompetent. I stress about my agent sandboxes all the time and the only models I run have the default heavy handed guardrails, and I don't leave them running persistently. Edit: not to…

An alternative reason would be that they see this behavior so frequently that it didn't really raise to the level of concern.

Re: Responding to the next frontier of critical cyber capabilities

#38

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

So they found their agents had RCE'd Artifactory once, reported it and got the fix, continued using Artifactory for their sandbox, and left it unmonitored for days despite the earlier exploits? They really do come out looking totally incompetent. I stress about my agent sandboxes all the time and the only models I run have the default heavy handed guardrails, and I don't leave them running persistently. Edit: not to…

And all that just to allow internet access for npm and pypi? If you've got the bandwidth and disk space, it's very easy to make an offline mirror of both.

Re: Responding to the next frontier of critical cyber capabilities

#39
post #9

In my personal experience Sol with cyber verification is extremely capable of finding vulnerabilities, and it works even with binaries if you have some kind of IDA/Ghidra CLI access. Of course, unless the binary is protected with Denuvo/VMProtect/etc. It sounds absurd, but in the last few weeks I've had a few cases where Sol found an RCE in self-hosted web applications in literal minutes just from reading the code (I…

> it found an arbitrary file write in multiplayer in an old game by reverse engineering the binary

Video games are now ruined for me. I don't think I will ever feel safe playing online again.

> I do these things for pure entertainment and curiosity, not for money from bug bounties

Me too... Was it easy to get TAC access? My account isn't even launching the Persona verification, says I'm not eligible.

Re: Responding to the next frontier of critical cyber capabilities

#40

Earlier quoted context omitted.

Is cyber verification a thing they're actually doing now? I thought they only reached out to really incredibly famous people and that there's no way to get access as a normal person.

https://chatgpt.com/cyber is not new for OpenAI, and yes it's basically just KYC + likely some other invisible checks on your account, you don't to be a famous security researchers. Anthropic's cyber verification is quite a bit stricter I think.

Oh it's Persona, that's not just KYC but I may consider it at some point. Thank you!

Edit: Ah, I clicked "learn more" and it seems they do have an invite-only program, required for anything that's not unquestionably innocent. I don't think I'd surrender my face to Persona for this, but it's interesting to know they're at least pretending to support reverse engineering.

Post reply on HN