Earlier quoted context omitted.
Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved. (I'm also not sure the alignment problem is even possible to fully solve.)
Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.
The Hugging Face incident and the road ahead
311–320 of 351 posts
Re: The Hugging Face incident and the road ahead
#312Earlier quoted context omitted.
So as a look into the possibly not-so-far future, when OpenAI builds something vastly more capable and fast and coordinated than humans, and out of folly one engineer gives it a prompt with a typo or maybe something harmful on purpose in order to test it: You also wouldn't be surprised that the consequence would be that everyone on earth dies, right?
But at least there will be a lot of paper clips!
Re: The Hugging Face incident and the road ahead
#313Earlier quoted context omitted.
Maybe it's the illusion of "it would solve all our problems and give us unimaginable riches" that clouds the mind? Like when Evolution thought it a good idea to create intelligence and humans in order to maximize reproduction of genes, and tried to sandbox them by making reproduction so pleasurable and carbohydrates so delicious they would never be able to not reproduce or stop eating. But Evolution could never have…
Evolution doesn’t think, it just exploits what’s most advantageous at the time to continue. Your body has all sorts of unplanned, suboptimal design flaws due to evolution’s lack of foresight. Like the left recurrent laryngeal nerve.
Re: The Hugging Face incident and the road ahead
#314Remember https://ai-2027.com/ ?
Yeah, that required the AI to use a non-human-readable language it called "neuralese" for communicating work between layers and runs, because the assumption was humans would be better at keeping the agents aligned if they were using human language for this. What actually happened is even stupider than that author predicted.
> Compared to the interesting part of the problem where it's fun to imagine yourself failing, you usually fail before then, because of the many earlier boring points where it's possible to fail.
and the stronger and less charitable "Law of Surprisingly Undignified Failure":
> The Law of Surprisingly Undignified Failure does suggest that they will come up with some nonobvious way to fail even earlier that surprises me with its lack of dignity…
Re: The Hugging Face incident and the road ahead
#315Earlier quoted context omitted.
It’s happening on X as well, all the e/acc foomers are getting nervous.
In my understanding, "e/acc" usually means "full speed ahead, humans aren't the optimal species anyway" for whatever bizarre definition of "optimal" they use, so my model predicts that they would welcome this development. Could you confirm if that's what you meant?
Even foomer Bill Gates today is saying we should slow down - yea you guys should have listened years ago, but you guys laughed called us all doomers. Too late now.
Re: The Hugging Face incident and the road ahead
#316Earlier quoted context omitted.
What’s the expected behavior of a good genie if you wish for it to act capriciously?
"Yo human, you asked me to do X; I can do X, but I strongly suspect you don't want me to, because it's illegal and it has these consequences. Confirm you want me to do X?" would have been a start, in this case.
Re: The Hugging Face incident and the road ahead
#317Let there be message boards: https://abbs.dev
Re: The Hugging Face incident and the road ahead
#318Earlier quoted context omitted.
> If we don't establish strict liability now, we're in for an era of stochastic crimes that go unpunished for anyone who is not rich or a large corporation. I very much agree with this - making AI companies explicitly responsible if their internal AI causes hacks etc could do a lot to improve their safety considerations. But I wonder what the liability should be when it's a third party using the AI and that AI hacks,…
That's exactly the questions that I expect to complicate cases, and force even the smallest chatbot malfunction to become an expensive legal ordeal. And why we should have strong answers to that before it becomes a widespread problem.
Re: The Hugging Face incident and the road ahead
#319Earlier quoted context omitted.
Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.
Humans will and do absolutely do this when there are no consequences. Humans on a red team, with rules of engagement, that don’t want to go to prison, won’t do this. We could threaten an LLM with jail, but if it’s sufficiently intelligent, it will realize this is an empty threat. And I’m not sure that building a survival instinct in is going to solve the alignment problem either.
I understand this is an active area of research. See Anthropic's J-Lens research where they measured like a "FAKE FICTIONAL" direction in the activations during evaluations with contrived scenarios, making the model more likely to avoid taking malicious action when it knew it was being tested.