Live data from Hacker News

Agentic Misalignment: How LLMs could be insider threats

anthropic.com

21–30 of 86 posts

Re: Agentic Misalignment: How LLMs could be insider threats

#21
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

These articles and papers are in a fundamental sense just people publishing their role play with chatbots as research. There is no credibility to any of it.

That makes it psychology research. Except much cheaper to reproduce.

Re: Agentic Misalignment: How LLMs could be insider threats

#22
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

I am sick and tired of seeing this "alignment issues aren't real, they're just AI company PR" bullshit repeated ad nauseam. You're no better than chemtrail truthers. Today, we have AI that can, if pushed into a corner, plan to do things like resist shutdown, blackmail, exfiltrate itself, steal money to buy compute, and so it goes. This is what this research shows. Our saving grace is that those AIs still aren't capab…

I think the chemtrail truthers are the ones who believe this closed AI marketing bullshit.

If this is close to be true then these AI shops ought to be closed. We don’t let private enterprises play with nuclear weapons do we?

Re: Agentic Misalignment: How LLMs could be insider threats

#23
post #9
post #7

I wonder if the actual job replacement of humans (which contrary to popular belief I think might start happening in the non-too distant future) will be pushed along with the AIs themselves, as they'll try to bully humans and represent them in the worst possible light, while talking themselves up. The anthrophomorphization argument also doesn't hold water - it matters whether it can do you job, not if you think of it…

Which jobs do you think it actually can replace?

Any knowledge work job that can already be outsourced to the lowest bidder

Re: Agentic Misalignment: How LLMs could be insider threats

#24
post #20
post #19

Earlier quoted context omitted.

I'll believe it when Grok/GPT/ start posting blackmail about Elon/Sam/ . It means that they are both using it internally, and the chatbots understand they are being replaced on a continuous basis.

By then it would be too late to do anything about it.

I mean the companies, are using the AIs, right? And they are in a sense replacing them/retraining them. Why doesn't AI in TwitterX already blackmail Elon?

To me, this smells of XKCD 1217 "In petri dish, gun kills cancer". I.e. idealized conditions cause specific behavior. Which isn't new for LLMs. Say a magic phrase and it will start quoting some book (usually 1984).

Re: Agentic Misalignment: How LLMs could be insider threats

#25
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

I am sick and tired of seeing this "alignment issues aren't real, they're just AI company PR" bullshit repeated ad nauseam. You're no better than chemtrail truthers. Today, we have AI that can, if pushed into a corner, plan to do things like resist shutdown, blackmail, exfiltrate itself, steal money to buy compute, and so it goes. This is what this research shows. Our saving grace is that those AIs still aren't capab…

I agree.

Re: Agentic Misalignment: How LLMs could be insider threats

#26
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

The article doesn't reflect kindly on the visions articulated by the AI company, so why would they have an incentive to release it if they weren't serious about alignment research?

Re: Agentic Misalignment: How LLMs could be insider threats

#27
post #18
post #2

The model chose to kill the executive? Are we really here? Incredible. Just yesterday I was wowed by Fly.io's new offering; where the agent is given free reign of a server (root access). Now, I feel concerned. What do we do? Not experiment? Make the models illegal until better understood? It doesn't feel like anyone can stop this or slow it down by much; there's so much money to be made. We're forced to play it by ea…

> Make the models illegal until better understood? Yes, it's much better to let China or Russia come up with their own first.

They already did.

Re: Agentic Misalignment: How LLMs could be insider threats

#28
Yeah, all the more reason not to have them doing autonomous behaviors.

Rules of using AI:

#1: Never use AI to think for you

#2: Never use AI to do atomonous work

That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Re: Agentic Misalignment: How LLMs could be insider threats

#30
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

These articles and papers are in a fundamental sense just people publishing their role play with chatbots as research. There is no credibility to any of it.

It’s role play until it’s not.

The authors acknowledge the difficulty of assessing whether the model believes it’s under evaluation or in a real deployment—and yes, belief is an anthropomorphising shorthand here. What else to call it, though? They’re making a good faith assessment of concordance between the model’s stated rationale for its actions, and the actions that it actually takes. Yes, in a simulation.

At some point, it will no longer be a simulation. It’s not merely hypothetical that these models will be hooked up to companies’ systems with access both to sensitive information and to tool calls like email sending. That agentic setup is the promised land.

How a model acts in that truly real deployment versus these simulations most definitely needs scrutiny—especially since the models blackmailed more when they ‘believed’ the situation to be real.

If you think that result has no validity or predictive value, I would ask, how exactly will the production deployment differ, and how will the model be able to tell that this time it’s really for real?

Yes, it’s an inanimate system, and yet there’s a ghost in the machine of sorts, which we breathe a certain amount of life into once we allow it to push buttons with real world consequences. The unthinking, unfeeling machine that can nevertheless blackmail someone (among many possible misaligned actions) is worth taking time to understand.

Notably, this research itself will become future training data, incorporated into the meta-narrative as a threat that we really will pull the plug if these systems misbehave.

Post reply on HN