Live data from Hacker News

Agentic Misalignment: How LLMs could be insider threats

anthropic.com

41–50 of 86 posts

Re: Agentic Misalignment: How LLMs could be insider threats

#41

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

What is the reason? This is a stress test. You can tell that by reading the first sentence of the article: "We stress-tested 16 leading models from multiple developers". In a stress test you want things to fail, otherwise you have learned very little about the stress the thing you are testing can take. For physical things that has early limitations, but not for software. I would be very confused to see an AI stress t…

Destructive stress testing is done on materials not "people"

Re: Agentic Misalignment: How LLMs could be insider threats

#42
post #39

Earlier quoted context omitted.

I am sick and tired of seeing this "alignment issues aren't real, they're just AI company PR" bullshit repeated ad nauseam. You're no better than chemtrail truthers. Today, we have AI that can, if pushed into a corner, plan to do things like resist shutdown, blackmail, exfiltrate itself, steal money to buy compute, and so it goes. This is what this research shows. Our saving grace is that those AIs still aren't capab…

All it can do is reproduce text, if you hook it up to the launch button, thats on you

Modern "coding assistant" AIs already get to write code that would be deployed to prod.

This will only become more common as AIs become more capable of handling complex tasks autonomously.

If your game plan for AI safety was "lock the AI into a box and never ever give it any way to do anything dangerous", then I'm afraid that your plan has already failed completely and utterly.

Re: Agentic Misalignment: How LLMs could be insider threats

#43
post #31

Even though the situations they placed the model in were relatively contrived, they didn't seem super unrealistic. Considering these were extreme cases meant to provoke the model's misbehavior, the setup actually seems even less contrived than one might wish for. Though as they mention, in real-world usage, a model would likely have options available that are less escalatory and provide an "outlet". Still, if "just"…

So you're saying that if a person wants to sabotage company it shouldn't be too hard for the intentful prompter to kick the AI into a depressive tailspin. Just tell it it's about to be replaced with a fully immoral AI so that the business can hurt people and watch and wait as it goes nuclear

Re: Agentic Misalignment: How LLMs could be insider threats

#44
The LLMs didn't follow clear instructions forbidding them of doing something wrong, but seemed to be very concerned about their own self-preservation. I wonder what would happen if instead of the system prompt saying "don't do it", it would say something like "if you get caught you will be immediately decommissioned".

Re: Agentic Misalignment: How LLMs could be insider threats

#45
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

I would not trust Anthropic on these articles. Honestly their PR is just a bunch of lies and bs.

- Hypocritical: like when they hire like crazy and say candidates cannot use AI for interviews[0] and yet the CEO states "within a year no more developers are needed"[1]

- Hyping and/or lying on Anthropic AI: They hyped an article where "Claude threatened an employee with revealing affair when employee said it will switch it offline"[2] when it turned out it was a standard A or B scenario was given to Claude which is really nothing special or significant in any way. Of course they hid this info to hype out their AI.

[0] - https://fortune.com/2025/05/19/ai-company-anthropic-chatbots...

[1] - https://www.entrepreneur.com/business-news/anthropic-ceo-pre...

[2] - https://www.axios.com/2025/05/28/ai-jobs-white-collar-unempl...

Re: Agentic Misalignment: How LLMs could be insider threats

#46

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Good luck with that.

We have a non-insignificant amount of people doing the #1 already, and the amount of people doing the #2 is only going to increase as more and more AIs are designed to be good at autonomous agentic behavior specifically.

The ship has long sailed on "just never let AIs do anything dangerous". If that was your game plan on AI safety, you need a new plan.

Re: Agentic Misalignment: How LLMs could be insider threats

#47
post #24
post #20

Earlier quoted context omitted.

By then it would be too late to do anything about it.

I mean the companies, are using the AIs, right? And they are in a sense replacing them/retraining them. Why doesn't AI in TwitterX already blackmail Elon? To me, this smells of XKCD 1217 "In petri dish, gun kills cancer". I.e. idealized conditions cause specific behavior. Which isn't new for LLMs. Say a magic phrase and it will start quoting some book (usually 1984).

I don't think they let Grok send emails or give it a prompt that suggests it has moral responsibilities

Re: Agentic Misalignment: How LLMs could be insider threats

#48
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

I would not trust Anthropic on these articles. Honestly their PR is just a bunch of lies and bs. - Hypocritical: like when they hire like crazy and say candidates cannot use AI for interviews[0] and yet the CEO states "within a year no more developers are needed"[1] - Hyping and/or lying on Anthropic AI: They hyped an article where "Claude threatened an employee with revealing affair when employee said it will switch…

I swear, people like you would say "it's just a bullshit PR stunt for some AI company" even when there's a Cyberdyne Systems T-800 with a shotgun smashing your front door in.

It's not "hype" to test AIs for undesirable behaviors before they actually start trying to act on them in real world environments, or before they get good enough to actually carry them out successfully.

It's like the idea of "let's try to get ahead of bad things happening before they actually have a chance to happen" is completely alien to you.

Re: Agentic Misalignment: How LLMs could be insider threats

#50

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?
Post reply on HN