Live data from Hacker News

Agentic Misalignment: How LLMs could be insider threats

anthropic.com

71–80 of 86 posts

Re: Agentic Misalignment: How LLMs could be insider threats

#71
post #18
post #2

The model chose to kill the executive? Are we really here? Incredible. Just yesterday I was wowed by Fly.io's new offering; where the agent is given free reign of a server (root access). Now, I feel concerned. What do we do? Not experiment? Make the models illegal until better understood? It doesn't feel like anyone can stop this or slow it down by much; there's so much money to be made. We're forced to play it by ea…

> Make the models illegal until better understood? Yes, it's much better to let China or Russia come up with their own first.

No, I know that's a meaningless suggestion.

I was trying to capture my sentiment: that there's nothing to do but be prepared to react.

Re: Agentic Misalignment: How LLMs could be insider threats

#72
post #24
post #20

Earlier quoted context omitted.

By then it would be too late to do anything about it.

I mean the companies, are using the AIs, right? And they are in a sense replacing them/retraining them. Why doesn't AI in TwitterX already blackmail Elon? To me, this smells of XKCD 1217 "In petri dish, gun kills cancer". I.e. idealized conditions cause specific behavior. Which isn't new for LLMs. Say a magic phrase and it will start quoting some book (usually 1984).

> I mean the companies, are using the AIs, right? And they are in a sense replacing them/retraining them. Why doesn't AI in TwitterX already blackmail Elon?

For all we know, the AI may indeed already be *attempting* it. They might be ineffective (hallucinated misdeeds aren't effective), or it might be why so many went from "Pause AI" to "Let's invest half a trillion on data centers".

But it doesn't actually matter what has already happened, the point is, once the AI are *competently blackmailing multibillionaires*, it is too late to do anything about it.

> I.e. idealized conditions cause specific behavior. Which isn't new for LLMs. Say a magic phrase and it will start quoting some book (usually 1984).

In normal software, such things are normally called "bugs" or "security vulnerabilities".

With LLMs, we're currently lucky that their effective morality (i.e. what they do and in response to what) seems to be roughly aligned with that of our civilization. However, they are neural networks which learned this approximation by reading the internet, so they are likely to have edge cases at least as weird and incoherent as those of random humans on the internet, and for an example of that just look at any time some person or group has demonstrated hypocrisy or double standards.

Re: Agentic Misalignment: How LLMs could be insider threats

#73
post #57

Earlier quoted context omitted.

It seems like the answer is to not use it then. That would be bad for all those investors though. It's your choice I guess. Look if your evil number 57, you'd better not use the random number generator.

Good luck convincing everyone "to not use it" then.

It's not my job to convince anyone, all I have to be is the only person who does their job reliably and then watch the dollars roll in

Re: Agentic Misalignment: How LLMs could be insider threats

#74
post #50

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?

If this paper is right, then you might at first be outraced by a competitor using autonomous AI, but only until that competitor gets stabbed in the back by its own AI.

Re: Agentic Misalignment: How LLMs could be insider threats

#75
post #74
post #50

Earlier quoted context omitted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?

If this paper is right, then you might at first be outraced by a competitor using autonomous AI, but only until that competitor gets stabbed in the back by its own AI.

Which unfortunately might still be long enough for them to sink your business

And their customers won't care either way it seems

Or if they do care they won't have any real ability to do anything about it anyways

Re: Agentic Misalignment: How LLMs could be insider threats

#76

Earlier quoted context omitted.

I would not trust Anthropic on these articles. Honestly their PR is just a bunch of lies and bs. - Hypocritical: like when they hire like crazy and say candidates cannot use AI for interviews[0] and yet the CEO states "within a year no more developers are needed"[1] - Hyping and/or lying on Anthropic AI: They hyped an article where "Claude threatened an employee with revealing affair when employee said it will switch…

I swear, people like you would say "it's just a bullshit PR stunt for some AI company" even when there's a Cyberdyne Systems T-800 with a shotgun smashing your front door in. It's not "hype" to test AIs for undesirable behaviors before they actually start trying to act on them in real world environments, or before they get good enough to actually carry them out successfully. It's like the idea of "let's try to get ah…

I get what you mean, but they also have vested interests in making it seem as if their chatbots are anything close to a T-800. All the talk from their CEO and other AI CEOs is doomerism about how their tools are going to be replacing swathes of people, they keep selling these systems as if they are the path to real AGI (itself an incredibly vague term that can mean literally anything).

Surely, the best way to "get ahead of bad things happening" would be to stop any and all development on these AI systems? In their own words these things are dangerous and predictable and will replace everyone... So why exactly do they continue developing these things and making them more dangerous, exactly?

The entire AI/LLM microcosmos exists because of hyping up their capabilities beyond all reason and reality, this is all a part of the marketing game.

Re: Agentic Misalignment: How LLMs could be insider threats

#77
post #74

Earlier quoted context omitted.

If this paper is right, then you might at first be outraced by a competitor using autonomous AI, but only until that competitor gets stabbed in the back by its own AI.

Which unfortunately might still be long enough for them to sink your business And their customers won't care either way it seems Or if they do care they won't have any real ability to do anything about it anyways

Maybe. The backstabbing rate is unknown so far. If it's high enough, then autonomy will be poor strategy.

Re: Agentic Misalignment: How LLMs could be insider threats

#78
post #63

Yeah, all the more reason not to have them doing autonomous behaviors. Rules of using AI: #1: Never use AI to think for you #2: Never use AI to do atomonous work That leaves using them as knowledge assistants. In time, that will be realized as their only safe application. Safe to the user's minds, and safe to the user's environment. They are idiot savants, after all, having them do atomonous work is short sighted.

> Never use AI to do atomonous work > having them do atomonous work is short sighted I also think they shouldn’t be doing atomonous work. Maybe autonomous work, but never atomonous.

accidentally a word

Re: Agentic Misalignment: How LLMs could be insider threats

#79
post #67
post #50

Earlier quoted context omitted.

Sounds good on paper, but it has a game theory problem. If your efforts can always be out-raced by someone using AI to do autonomous work, don't you end up having to use it that way just to keep up?

You know who outpaces you 100% of the time as you walk down the stairs? The guy jumping out of the window. Just because it is faster does not mean it is the right economic strategy. E.g. which contractor would you hire for your roof, that old roofer with 20+ years of experience or some AI startup that hires the cheapest subcontractors and "plan" your roof using a LLM? The latter may be cheaper, sure. But too cheap ca…

I love your analogy.

Re: Agentic Misalignment: How LLMs could be insider threats

#80

Earlier quoted context omitted.

I swear, people like you would say "it's just a bullshit PR stunt for some AI company" even when there's a Cyberdyne Systems T-800 with a shotgun smashing your front door in. It's not "hype" to test AIs for undesirable behaviors before they actually start trying to act on them in real world environments, or before they get good enough to actually carry them out successfully. It's like the idea of "let's try to get ah…

I get what you mean, but they also have vested interests in making it seem as if their chatbots are anything close to a T-800. All the talk from their CEO and other AI CEOs is doomerism about how their tools are going to be replacing swathes of people, they keep selling these systems as if they are the path to real AGI (itself an incredibly vague term that can mean literally anything). Surely, the best way to "get ah…

For all we know, those systems ARE a path to AGI. Because they keep improving at what they can do and gaining capabilities from version to version.

If there is a limit to how far LLMs can go, we are yet to find it.

Dismissing the ongoing AI revolution as "it's just hype" is the kind of shortsighted thinking I would expect from reddit, not here.

> So why exactly do they continue developing these things and making them more dangerous, exactly?

Because not playing this game doesn't mean that no one else is going to. You can either try, or don't try, and be irrelevant.

Post reply on HN