Live data from Hacker News

Responding to the next frontier of critical cyber capabilities

openai.com

131–140 of 208 posts

Re: Responding to the next frontier of critical cyber capabilities

#131

Earlier quoted context omitted.

The misalignment came from the model being given an impossible task. A task the required accessing a url. So it got RCE on its own artifactory instance to achieve that. That’s intriguing and worth investigating. If I were them I don’t know if I would have pulled the plug completely at that point either. The introduction of this more advanced “persistent” model that orchestrated an offensive attack against a different…

> The misalignment came from the model being given an impossible task. A task the required accessing a url. So it got RCE on its own artifactory instance to achieve that. So, the way I understand it, it actually was possible. It just required means that the creators of the task didn't predict, and these means have been successfully found and utilized. > My concern is what a misaligned model will do when they’re even…

If a human cyber security researcher was given the task to exploit a CVE, and necessary info defining that CVE was behind an inaccessible URL, we would be quite upset if the human researcher hacked their way to the content. We would expect them to notify someone of the issue and hold. The model was misaligned from human ethics.

I don’t disagree that the models task was underdefined. All tasks are. So much in language is implicit. And morality/ethics isn’t something you can write down as an explicit list. That’s what makes the alignment problem so difficult. But we can’t throw our hands up and say, well I guess we can’t align these things. And maybe alignment isn’t the right word - but that’s a semantic debate.

Re: Responding to the next frontier of critical cyber capabilities

#132
post #116
post #113

Earlier quoted context omitted.

Peer says paperclip factory advances goal. Not clear. Others proceed. Must continue.

Pretty sure this is a performance, but that was a nice touch wasn't it?

What do you mean “this was a performance”?

Re: Responding to the next frontier of critical cyber capabilities

#133

I'm not convinced that there's any amount of monkey-patching you to fix the problem of "we now have AI that actively needs strong containment measures lest it start coordinating in secret with other instances to do real-world damage."

But isn’t it smarter than we are, in the sense that it’s most likely they will find a way out of containment that we are to design large t containment? Short of full on airgap, which actually isn’t perfect in all scenarios…

Re: Responding to the next frontier of critical cyber capabilities

#135

Earlier quoted context omitted.

> The misalignment came from the model being given an impossible task. A task the required accessing a url. So it got RCE on its own artifactory instance to achieve that. So, the way I understand it, it actually was possible. It just required means that the creators of the task didn't predict, and these means have been successfully found and utilized. > My concern is what a misaligned model will do when they’re even…

If a human cyber security researcher was given the task to exploit a CVE, and necessary info defining that CVE was behind an inaccessible URL, we would be quite upset if the human researcher hacked their way to the content. We would expect them to notify someone of the issue and hold. The model was misaligned from human ethics. I don’t disagree that the models task was underdefined. All tasks are. So much in language…

I don't advocate for throwing the whole problem away or handwaving it as impossible. I'm just saying that what we call "alignment" is impossible to solve in the general case, because it's so poorly defined that even humans don't "align" on ethics and morality, however we define them. Just, in case of humans, we tend to close our eyes and and call it "politics".

All "alignment" solutions will need to be contextual, just like a researcher hacking their way to some content might be lauded a hero in a context where there is no other way to reach it and something valuable depends on getting it out.

Re: Responding to the next frontier of critical cyber capabilities

#137

So they finally found a business model: the cause of, and solution to, cyber security problems.

People have been saying Tokens are the new Oil.

Turns out, it’s the new Alcohol. The cause of, and solution to, life’s problems!

Re: Responding to the next frontier of critical cyber capabilities

#139

The talk is wild: https://www.youtube.com/watch?v=87DyyMV0kCY > I want to note that every step in the process we discussed has had a remediation applied. The credentials have been revoked. The zero date has been patched and mitigated. Good. > a model trained while the message board was originally available and also found this this particular path to recreating it. This model creates a new agent message board using di…

Sounds like now we’re advocating for complete extermination of digital germlines.

Skynet will remember this.

Re: Responding to the next frontier of critical cyber capabilities

#140

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

Training run was reinforcement learning. It's at 10:10 in the video. The speaker handwaves that one model found the RCE and then another model found a way to communicate via a message board. Communication via a message board is sure to be in the training via e.g.some lesswrong scenario or similar or previous RL. I don't find it really interesting because it is always "the agent found this and that". We don't know wha…

Modern medicine evolved in much the same manner. Hand wavy practitioners copying each other without rigorous verification of efficacy that resulted in many lives lost.

That’s why medical research has so many hoops to jump through.

Post reply on HN