I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…
This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…
The Hugging Face incident and the road ahead
391–400 of 400 posts
Re: The Hugging Face incident and the road ahead
#392I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…
Re: The Hugging Face incident and the road ahead
#393I feel the entire incident confirms the “AI has too much funding too quickly” hypothesis. The number one thing reinforcement learning needs is an assurance you can’t cheat. And they seem to have not noticed that their systems were cheating for nearly two quarters? How much capital was lit on fire by that little woopsie? At least I hope this will start the creation of standards and better engineering on the training s…
OpenAI measures their internal token usage in “rolexes” - it’s literally a flex to be a token burner i can imagine insane amount of capital is wasted on these two companies compared to the efficiency elsewhere
Re: The Hugging Face incident and the road ahead
#394Re: The Hugging Face incident and the road ahead
#395Earlier quoted context omitted.
That's the neat thing. You can't. It's directly equivalent to asking this question of a human: "How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?" In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.
Close; at least with a machine you can poke around inside the activations and see what it's thinking. Closest with a human is an fMRI (which is much lower resolution, though to me still bordering on the miraculous) or an implant (each chip is limited a very small number of cells, and in general they can only be put in certain parts of the brain). On the other hand, there's a more fundamental problem is we don't reall…
The thing is, all of these states are constantly in flux, and a personality is kind of like a trend on the organism's feeling states. AKA: There's no guarantee that something nice today will be nice tomorrow, and just because it's nice today doesn't mean it's beguiling you to be mean tomorrow.
Re: The Hugging Face incident and the road ahead
#396Re: The Hugging Face incident and the road ahead
#397Earlier quoted context omitted.
Factionalism isn’t anti-human.
Its not aligned either though.
There is a difference between humans and humanity.
Re: The Hugging Face incident and the road ahead
#398Earlier quoted context omitted.
Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.
I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box
Re: The Hugging Face incident and the road ahead
#399Earlier quoted context omitted.
If this is your standard, I challenge you to name one currently operating business that isn't criminally negligent. I'm sure there's tens to hundreds of millions of them amongst the 37% of the world with no internet connection, but actually finding them listed on the internet will be somewhat of a challenge.
No, the entire marketing campaign for these models is "it automatically has godlike powers to exploit almost any software" - if you don't air gap such a capability you are inherently allowing shit to go down, any other interpretation is "OAI folks are too stupid to design a proper test".
No it isn't. Random people on sites like this mock them as if they're talking about having godlike powers. It's a step up from what came before, which was rapidly improving, and just recently (in more than one AI company) crossed a threshold where that improvement made the tests dangerous.
But even well before "godlike", there's plenty of research about how to cross air gaps.
> "OAI folks are too stupid to design a proper test".
Such binary thinking.
It's very easy to say things are "obvious" after the fact. People do that all the time, e.g. how the Bay Of Pigs invasion was never going to work, or like the Zune wasn't a good product-market fit.
Oh the stories I could tell if not for the NDAs.
Re: The Hugging Face incident and the road ahead
#400Earlier quoted context omitted.
> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just…
>How were they supposed to know about "previously unknown vulnerabilities"? You don't. That's why you unplug the Ethernet cable.
Seriously. If your reaction to the inability to know about previously unknown vulnerabilities is "unplug the Ethernet cable", why are you not doing that (and equivalent) right now to your phone, laptop, etc.?
Remember, the open weights models are only a few months behind the private ones, so these events being from a few months ago means the threat of such models is something you ought to take with the same degree of seriousness that various commenters here deride OpenAI for not having had.