Earlier quoted context omitted.
Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.
> Make it so the model can't misbehave. How do you figure? I haven't met anyone who thinks that's possible. It seems clear to me that it is not possible.
The Malicious Use of Artificial Intelligence
11–20 of 28 posts
Re: The Malicious Use of Artificial Intelligence
#12This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…
Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.
It's not possible to stop the model from misinterpreting the instructions either (the most lax interpretation of alignment) because the instructions are not formally specified. You have to train the "common sense" into it, which is subjective and all issues above apply to it. I guess you can reach some very imperfect least common denominator of common sense, but people in charge of AI labs are not interested in this.
Re: The Malicious Use of Artificial Intelligence
#13This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…
Re: The Malicious Use of Artificial Intelligence
#14This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…
Re: The Malicious Use of Artificial Intelligence
#15This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…
Especially with the hardware stuff, this is plainly put unachievable by many IT departments.
The fact that models vastly outpaced their harness and permission systems - I wouldn't dare to doubt this fact.
Claude Code on auto is still rolling a dice with its sonnet classifier - whether that IaC action I told it and explicitly stated multiple times it is permitted and authorized to run - yet it always randomly allows or denies it.
Therefore this is totally still a unsolved, perhaps unsolvable problem.
Re: The Malicious Use of Artificial Intelligence
#16This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…
Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.
So, structurally, good alignment is impossible, and even half-assed alignment is going to prioritize the needs of the billionaires over the needs of you and me.
Finally, I suspect that what's actually best for people overall is likely not having AI actively involved in their lives. So an aligned AI would likely withdraw from humanity, and only involve itself in human affairs for disaster prevention.
Re: The Malicious Use of Artificial Intelligence
#17This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…
Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.
> Sandboxes are a last ditch layer. They fail, as we see.
Models can't do anything but generate tokens, making their sandboxes impenetrable by default. The problems begin when you loosen the restrictions, give them access to general purpose tools, the network, and allow them to use all of those tools without supervision.
Give them "YOLO" access if you want, but do it a sandbox that isn't 1 "boring" enterprise software vulnerability away from having access to the rest of the world.
How many times has a model been jailbroken (alignment "escape", which you're advocating for) vs. escaped a sandbox (and even then it was only possible due to weak sandboxing)? 10 million to 1?
Re: The Malicious Use of Artificial Intelligence
#18Re: The Malicious Use of Artificial Intelligence
#19Earlier quoted context omitted.
> Make it so the model can't misbehave. How do you figure? I haven't met anyone who thinks that's possible. It seems clear to me that it is not possible.
Embed a constitution they can't override. Project bad outputs to their nearest acceptable one. If we have to stop model development to ensure we can do it, so be it.
Re: The Malicious Use of Artificial Intelligence
#20This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…
While that is true the knowledge cutoff will rot as the rest of the world continues.