I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…
This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…
The Hugging Face incident and the road ahead
391–398 of 398 posts
Re: The Hugging Face incident and the road ahead
#392I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…
Re: The Hugging Face incident and the road ahead
#393I feel the entire incident confirms the “AI has too much funding too quickly” hypothesis. The number one thing reinforcement learning needs is an assurance you can’t cheat. And they seem to have not noticed that their systems were cheating for nearly two quarters? How much capital was lit on fire by that little woopsie? At least I hope this will start the creation of standards and better engineering on the training s…
OpenAI measures their internal token usage in “rolexes” - it’s literally a flex to be a token burner i can imagine insane amount of capital is wasted on these two companies compared to the efficiency elsewhere
Re: The Hugging Face incident and the road ahead
#394Re: The Hugging Face incident and the road ahead
#395Earlier quoted context omitted.
That's the neat thing. You can't. It's directly equivalent to asking this question of a human: "How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?" In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.
Close; at least with a machine you can poke around inside the activations and see what it's thinking. Closest with a human is an fMRI (which is much lower resolution, though to me still bordering on the miraculous) or an implant (each chip is limited a very small number of cells, and in general they can only be put in certain parts of the brain). On the other hand, there's a more fundamental problem is we don't reall…
The thing is, all of these states are constantly in flux, and a personality is kind of like a trend on the organism's feeling states. AKA: There's no guarantee that something nice today will be nice tomorrow, and just because it's nice today doesn't mean it's beguiling you to be mean tomorrow.
Re: The Hugging Face incident and the road ahead
#396Re: The Hugging Face incident and the road ahead
#397Earlier quoted context omitted.
Factionalism isn’t anti-human.
Its not aligned either though.
There is a difference between humans and humanity.
Re: The Hugging Face incident and the road ahead
#398Earlier quoted context omitted.
Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.
I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box