Live data from Hacker News

The Malicious Use of Artificial Intelligence

arxiv.org

11–20 of 28 posts

Re: The Malicious Use of Artificial Intelligence

#11
post #6
post #4

Earlier quoted context omitted.

Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.

> Make it so the model can't misbehave. How do you figure? I haven't met anyone who thinks that's possible. It seems clear to me that it is not possible.

Embed a constitution they can't override. Project bad outputs to their nearest acceptable one. If we have to stop model development to ensure we can do it, so be it.

Re: The Malicious Use of Artificial Intelligence

#12
post #4
post #3

This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…

Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.

If it's not dangerous it's also not useful, simple as that. For example if you train a model for cybersecurity, it can be used for both attack and defense. And almost every use is like that. Alignment is fundamentally flawed as a concept, it's a pie in the sky. Let alone the perverse version of it by crazy AI "safety" people that in practice means "the model does what I want, only for the people I allow".

It's not possible to stop the model from misinterpreting the instructions either (the most lax interpretation of alignment) because the instructions are not formally specified. You have to train the "common sense" into it, which is subjective and all issues above apply to it. I guess you can reach some very imperfect least common denominator of common sense, but people in charge of AI labs are not interested in this.

Re: The Malicious Use of Artificial Intelligence

#13
post #3

This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…

ai slop comment and my eyes glaze over

Re: The Malicious Use of Artificial Intelligence

#14
post #3

This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…

[deleted]

Re: The Malicious Use of Artificial Intelligence

#15
post #3

This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…

Realistically, the entire chain proposed by your safebots - I like the idea - but I cannot see viable ways to get it actually deployed in a useful manner.

Especially with the hardware stuff, this is plainly put unachievable by many IT departments.

The fact that models vastly outpaced their harness and permission systems - I wouldn't dare to doubt this fact.

Claude Code on auto is still rolling a dice with its sonnet classifier - whether that IaC action I told it and explicitly stated multiple times it is permitted and authorized to run - yet it always randomly allows or denies it.

Therefore this is totally still a unsolved, perhaps unsolvable problem.

Re: The Malicious Use of Artificial Intelligence

#16
post #4
post #3

This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…

Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.

Yes, alignment is the architectural solution. But it's also fantasy; you can't align with everyone. And even worse, even if it were possible, the people doing the alignment are only going to align up to where it keeps them profitable.

So, structurally, good alignment is impossible, and even half-assed alignment is going to prioritize the needs of the billionaires over the needs of you and me.

Finally, I suspect that what's actually best for people overall is likely not having AI actively involved in their lives. So an aligned AI would likely withdraw from humanity, and only involve itself in human affairs for disaster prevention.

Re: The Malicious Use of Artificial Intelligence

#17
post #4
post #3

This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…

Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model, claiming "Whose values is it aligned to?" Sandboxes are a last ditch layer. They fail, as we see.

> Make it so the model can't misbehave.

> Sandboxes are a last ditch layer. They fail, as we see.

Models can't do anything but generate tokens, making their sandboxes impenetrable by default. The problems begin when you loosen the restrictions, give them access to general purpose tools, the network, and allow them to use all of those tools without supervision.

Give them "YOLO" access if you want, but do it a sandbox that isn't 1 "boring" enterprise software vulnerability away from having access to the rest of the world.

How many times has a model been jailbroken (alignment "escape", which you're advocating for) vs. escaped a sandbox (and even then it was only possible due to weak sandboxing)? 10 million to 1?

Re: The Malicious Use of Artificial Intelligence

#19
post #11
post #6

Earlier quoted context omitted.

> Make it so the model can't misbehave. How do you figure? I haven't met anyone who thinks that's possible. It seems clear to me that it is not possible.

Embed a constitution they can't override. Project bad outputs to their nearest acceptable one. If we have to stop model development to ensure we can do it, so be it.

I agree, but because I want to see the development stopped forever, which is what the result of this would be. You will never have alignment that cannot be overridden in some ways. You won’t have a silver bullet here, you need safety at every layer

Re: The Malicious Use of Artificial Intelligence

#20
post #3

This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way. I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem. Every generation of computing has hit a version of this. Programs could writ…

> bits don't degrade

While that is true the knowledge cutoff will rot as the rest of the world continues.

Post reply on HN