Live data from Hacker News

Agentic Misalignment: How LLMs could be insider threats

anthropic.com

51–60 of 86 posts

Re: Agentic Misalignment: How LLMs could be insider threats

#51

Earlier quoted context omitted.

What is the reason? This is a stress test. You can tell that by reading the first sentence of the article: "We stress-tested 16 leading models from multiple developers". In a stress test you want things to fail, otherwise you have learned very little about the stress the thing you are testing can take. For physical things that has early limitations, but not for software. I would be very confused to see an AI stress t…

Destructive stress testing is done on materials not "people"

Apparently it's now done on "people" (which, in a very important distinction, are not actual people but software)

Re: Agentic Misalignment: How LLMs could be insider threats

#52
post #33

Earlier quoted context omitted.

It’s role play until it’s not. The authors acknowledge the difficulty of assessing whether the model believes it’s under evaluation or in a real deployment—and yes, belief is an anthropomorphising shorthand here. What else to call it, though? They’re making a good faith assessment of concordance between the model’s stated rationale for its actions, and the actions that it actually takes . Yes, in a simulation. At som…

Then test it. Make several small companies. Create an office space, put people to work there for a few months, then simulate an AI replacement. All testing methodology needs to be written on machines that are isolated or better always offline. Except CEO and few other actors everyone is there for real. See how many AIs actually follow up on their blackmails.

No need. We know today's AIs are simply not capable enough to be too dangerous.

But capabilities of AI systems improve generation to generation. And agentic AI? Systems that are capable of carrying out complex long term tasks? It's something that many AI companies are explicitly trying to build.

Research like this is trying to get ahead of that, and gauge what kind of weird edge case shenanigans agentic AIs might get to before they actually do it for real.

Re: Agentic Misalignment: How LLMs could be insider threats

#54
post #39

Earlier quoted context omitted.

All it can do is reproduce text, if you hook it up to the launch button, thats on you

Modern "coding assistant" AIs already get to write code that would be deployed to prod. This will only become more common as AIs become more capable of handling complex tasks autonomously. If your game plan for AI safety was "lock the AI into a box and never ever give it any way to do anything dangerous", then I'm afraid that your plan has already failed completely and utterly.

If you use it for a critical system, and something goes wrong, youre still responsible for the consequences.

Much like if I let my cat walk on my keyboard and it brings a server down.

Re: Agentic Misalignment: How LLMs could be insider threats

#55
post #50

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?

From an economic perspective it requires LLMs and humans to have comparable outputs. That's not possible in all domains - at least in the near future.

Re: Agentic Misalignment: How LLMs could be insider threats

#56
post #54

Earlier quoted context omitted.

Modern "coding assistant" AIs already get to write code that would be deployed to prod. This will only become more common as AIs become more capable of handling complex tasks autonomously. If your game plan for AI safety was "lock the AI into a box and never ever give it any way to do anything dangerous", then I'm afraid that your plan has already failed completely and utterly.

If you use it for a critical system, and something goes wrong, youre still responsible for the consequences. Much like if I let my cat walk on my keyboard and it brings a server down.

And?

"Sure, we have a rogue AI that managed to steal millions from the company, backdoor all of our infrastructure, escape into who-knows-what compute cluster when it got caught, and is now waging guerilla warfare against our company over our so-called mistreatment of tiger shrimps. But hey, at least we know the name of the guy who gave that AI a prompt that lead to all of this!"

Re: Agentic Misalignment: How LLMs could be insider threats

#57
post #54

Earlier quoted context omitted.

If you use it for a critical system, and something goes wrong, youre still responsible for the consequences. Much like if I let my cat walk on my keyboard and it brings a server down.

And? "Sure, we have a rogue AI that managed to steal millions from the company, backdoor all of our infrastructure, escape into who-knows-what compute cluster when it got caught, and is now waging guerilla warfare against our company over our so-called mistreatment of tiger shrimps. But hey, at least we know the name of the guy who gave that AI a prompt that lead to all of this!"

It seems like the answer is to not use it then.

That would be bad for all those investors though. It's your choice I guess.

Look if your evil number 57, you'd better not use the random number generator.

Re: Agentic Misalignment: How LLMs could be insider threats

#58
post #57

Earlier quoted context omitted.

And? "Sure, we have a rogue AI that managed to steal millions from the company, backdoor all of our infrastructure, escape into who-knows-what compute cluster when it got caught, and is now waging guerilla warfare against our company over our so-called mistreatment of tiger shrimps. But hey, at least we know the name of the guy who gave that AI a prompt that lead to all of this!"

It seems like the answer is to not use it then. That would be bad for all those investors though. It's your choice I guess. Look if your evil number 57, you'd better not use the random number generator.

Good luck convincing everyone "to not use it" then.

Re: Agentic Misalignment: How LLMs could be insider threats

#59
post #6
post #3

The writing perpetuates the anthropomorphising of these agents. If you view the agent as simply a program that is given a goal to achieve and tools to achieve it with, without any higher order “thought” or “thinking”, then you realise it is simply doing what it is “programmed” to do. No magic, just a drone fixed on an outcome.

I think the narrative of "AI is just a tool" is much more harmful than the anthropomorphism of AI. Yes, AI is a tool. So are guns. So are nukes. Many tools are easy to be misused. Most tools are inherently dangerous.

I don’t quite follow. Just because a tool has the potential for misuse, doesn’t make it not a tool.

Anthropomorphizing LLMs, on the other hand, has a multitude of clearly evident problems arising from it.

Or do you focus on the “just” part of the statement? That I very much agree with. Genuinely asking for understanding, not a native speaker.

Re: Agentic Misalignment: How LLMs could be insider threats

#60
post #27
post #18

Earlier quoted context omitted.

> Make the models illegal until better understood? Yes, it's much better to let China or Russia come up with their own first.

They already did.

Right, so only China and Russia should have models...
Post reply on HN