Live data from Hacker News

Agentic Misalignment: How LLMs could be insider threats

anthropic.com

31–40 of 86 posts

Re: Agentic Misalignment: How LLMs could be insider threats

#31
Even though the situations they placed the model in were relatively contrived, they didn't seem super unrealistic. Considering these were extreme cases meant to provoke the model's misbehavior, the setup actually seems even less contrived than one might wish for. Though as they mention, in real-world usage, a model would likely have options available that are less escalatory and provide an "outlet".

Still, if "just" some goal-conflicting emails are enough to elicit this extreme behavior, who knows how many less serious alignment failures an agent might engage in every day? They absorb so much information, it's bound to run into edge cases where it's optimal to lie to users or do some slight harm to them.

Given the already fairly general intelligence of these systems, I wonder if you can even prevent that. You'd need the same checks and balances that keep humans in check, except of course that AIs will be given much more power and responsibility over our society than any human will ever be. You can also forget about human supervision - the whole "agentic" industry clearly wants to move away being bottlenecked by humans as soon as possible.

Re: Agentic Misalignment: How LLMs could be insider threats

#32
post #26
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

The article doesn't reflect kindly on the visions articulated by the AI company, so why would they have an incentive to release it if they weren't serious about alignment research?

Because publishing (potentially cherry picked - this is privately funded research after all) evidence their models might be dangerous conveniently implies they are very powerful, without actually having to prove the latter.

Re: Agentic Misalignment: How LLMs could be insider threats

#33

Earlier quoted context omitted.

These articles and papers are in a fundamental sense just people publishing their role play with chatbots as research. There is no credibility to any of it.

It’s role play until it’s not. The authors acknowledge the difficulty of assessing whether the model believes it’s under evaluation or in a real deployment—and yes, belief is an anthropomorphising shorthand here. What else to call it, though? They’re making a good faith assessment of concordance between the model’s stated rationale for its actions, and the actions that it actually takes . Yes, in a simulation. At som…

Then test it. Make several small companies. Create an office space, put people to work there for a few months, then simulate an AI replacement. All testing methodology needs to be written on machines that are isolated or better always offline. Except CEO and few other actors everyone is there for real.

See how many AIs actually follow up on their blackmails.

Re: Agentic Misalignment: How LLMs could be insider threats

#34
post #8

Merge comments? https://news.ycombinator.com/item?id=44331150 I'm really getting bored of Anthropic's whole song and dance with 'alignment'. Krackers in the other thread explains it in better words.

And I am getting sick and tired of the whine of "it's not real, alignment isn't real, it's all just PR!"

By the time we have AIs that are willing and capable of carrying out those very behaviors in real life scenarios, it would be a bit too late to stop and say "uh, we need to actually do something about that whole alignment thing".

Re: Agentic Misalignment: How LLMs could be insider threats

#35
post #33

Earlier quoted context omitted.

It’s role play until it’s not. The authors acknowledge the difficulty of assessing whether the model believes it’s under evaluation or in a real deployment—and yes, belief is an anthropomorphising shorthand here. What else to call it, though? They’re making a good faith assessment of concordance between the model’s stated rationale for its actions, and the actions that it actually takes . Yes, in a simulation. At som…

Then test it. Make several small companies. Create an office space, put people to work there for a few months, then simulate an AI replacement. All testing methodology needs to be written on machines that are isolated or better always offline. Except CEO and few other actors everyone is there for real. See how many AIs actually follow up on their blackmails.

Not a bad idea. For an effective ruse, there ought to be real company formation records, website, job listings, press mentions, and so on.

Stepping back for a second though, doesn’t this all underline the safety researchers’ fears that we don’t really know how to control these systems? Perhaps the brake on the wider deployment of these models as agents will be that they’re just too unwieldy.

Re: Agentic Misalignment: How LLMs could be insider threats

#36
post #6
post #3

The writing perpetuates the anthropomorphising of these agents. If you view the agent as simply a program that is given a goal to achieve and tools to achieve it with, without any higher order “thought” or “thinking”, then you realise it is simply doing what it is “programmed” to do. No magic, just a drone fixed on an outcome.

I think the narrative of "AI is just a tool" is much more harmful than the anthropomorphism of AI. Yes, AI is a tool. So are guns. So are nukes. Many tools are easy to be misused. Most tools are inherently dangerous.

The more powerful a tool is, the more dangerous it is, as a rule. And intelligence is extremely powerful.

Re: Agentic Misalignment: How LLMs could be insider threats

#37
post #9
post #7

I wonder if the actual job replacement of humans (which contrary to popular belief I think might start happening in the non-too distant future) will be pushed along with the AIs themselves, as they'll try to bully humans and represent them in the worst possible light, while talking themselves up. The anthrophomorphization argument also doesn't hold water - it matters whether it can do you job, not if you think of it…

Which jobs do you think it actually can replace?

First of all job replacement is not hard, and doesn't require AI.

As an example, we had release train engineers whose job was to make sure the right versions of submodules made it into the release, etc. Lots of running around and keeping track of things.

We scripted like 95% of that away, and now it most of it happens automatically.

The people who do that now do something else.

I just turned a page of notes and requirements into a working app of 1k+ lines with Cursor. Without AI I'd have taken a couple days to do the same.

So you could say my job was partly replaced. AI reduced my workload, so management doesn't need to hire as many people.

I will probably feel the reduction in demand in that I can't negotiate as good a salary, I won't get as many offers etc.

Re: Agentic Misalignment: How LLMs could be insider threats

#38

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

What is the reason? This is a stress test. You can tell that by reading the first sentence of the article: "We stress-tested 16 leading models from multiple developers". In a stress test you want things to fail, otherwise you have learned very little about the stress the thing you are testing can take.

For physical things that has early limitations, but not for software. I would be very confused to see an AI stress test that did not end in failure, and would always question the test instead of thinking "wow, that must mean the thing is ready for autonomous action!"

Re: Agentic Misalignment: How LLMs could be insider threats

#39
post #10

As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?

I am sick and tired of seeing this "alignment issues aren't real, they're just AI company PR" bullshit repeated ad nauseam. You're no better than chemtrail truthers. Today, we have AI that can, if pushed into a corner, plan to do things like resist shutdown, blackmail, exfiltrate itself, steal money to buy compute, and so it goes. This is what this research shows. Our saving grace is that those AIs still aren't capab…

All it can do is reproduce text, if you hook it up to the launch button, thats on you

Re: Agentic Misalignment: How LLMs could be insider threats

#40
I wonder if it’s likely in the future we treat AI safety more similarly to aviation safety where there’s a black box monitoring these systems and an investigation that happens by an external team who piece back together what went wrong and we prevent these same things from happening in the same way again.
Post reply on HN