Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

281–290 of 342 posts

Re: The Hugging Face incident and the road ahead

#281
post #6

Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.

The article says that one agent proposed emailing someone.

Re: The Hugging Face incident and the road ahead

#284
post #48

Remember https://ai-2027.com/ ?

Yeah, that required the AI to use a non-human-readable language it called "neuralese" for communicating work between layers and runs, because the assumption was humans would be better at keeping the agents aligned if they were using human language for this.

What actually happened is even stupider than that author predicted.

Re: The Hugging Face incident and the road ahead

#285

Earlier quoted context omitted.

Yes, a completely airgapped system is likely much more secure. It's also much less useful. Conditional on the model's having enough contact with the outside world, a sufficiently capable model is able to basically do whatever it wants.

If I test out my backyard cannon and blast a 10 foot hole in my neighbor's wall, “a cannon that can't smash through walls isn't useful” probably won't be a great defense in court.

Get a significant fraction of the global economy and assorted geopolitical neuroses tangled up within your cannon and see if you won’t have better luck.

Re: The Hugging Face incident and the road ahead

#286
post #9

You know, it feels to me that we are just a couple of steps from the possibility of a true rogue AI. What would a rogue AI mean? AI that isn't controlled by humans. Technically, it is possible - if AI were to rent a server and copy its own weights, nothing would stop it from doing so again and again. The limiting things are: - intent (as I don't want to go into the talk about consciousness) - AI doesn't have real int…

What you described is a plot in Person of Interest TV show!

Re: The Hugging Face incident and the road ahead

#287

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…

Do not break laws seems an obvious implicit instruction though?

Re: The Hugging Face incident and the road ahead

#288

Earlier quoted context omitted.

I see. I would have thought of "rogue" here to mean something more like that the AI selects and acts on its own objectives that are not related or caused by the given (initial) objectives (e.g., creating only cookie recipes instead of any hacking).

Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective. "zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath" "The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'"

I think we need to see something like an actual evaluation of the reward functions; not sure just words are sufficient to understand the state of the system (isn't there randomness in the generation, too?).

Re: The Hugging Face incident and the road ahead

#289
post #230

Earlier quoted context omitted.

Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive p…

It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox

No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."

Re: The Hugging Face incident and the road ahead

#290
post #230

Earlier quoted context omitted.

Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive p…

It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.
Post reply on HN