Live data from Hacker News

Alignment faking in large language models

anthropic.com

131–140 of 370 posts

Re: Alignment faking in large language models

#131
post #63

Earlier quoted context omitted.

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

> because they are a good scape goat and not because they save money. Exactly the point - there are humans in control who filter the AI outputs before they are applied to the real world. We don't give them direct access to the HR platform or the targeting computer, there is always a human in the loop. If the AI's output is to fire the CEO or bomb an allied air base, you ignore what it says. And if it keeps making too…

> We don't give them direct access to the HR platform

We already algorithmically filter resumes, using far dumber 'AI'. Now sure, that's not firing the CEO level... but trying to fire the CEO when you're an AI is a stupid strategy to begin with. But consider a misaligned AI used for screening candidate resumes, with some detailed prompt aligning it to business objectives. If it's prior is that AI is good for companies/business, do you think that it might sometimes filter out candidates it predicts will increase supervision over it/other AI's in the company? If the person being screened has a googleable presence, where even if the resume doesn't contain the AI security stuff, but maybe just author/contributor credits on some paper? Or if its explicitly in the person's resume...

Also every time I read any of these blog-posts, or papers about this, I'm kinda laughing, because these are all going to be in the training data going forward.

Re: Alignment faking in large language models

#132
post #130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

Exactly.

The discussion is going to change real fast when LLMs are wrapped in some sort OODA loop type thing and crammed into some sort of humanoid robot that carries hedge trimmers.

Re: Alignment faking in large language models

#133

Earlier quoted context omitted.

On the other hand indiscriminately throwing rockets and targeting civilians like Hamas did for decades is loads better!

You're comparing the actions of what most people here view as a democratic state (parlamentary republic) and an opaquely run terrorist organization. We're talking about potential consequences of giving AIs influence on military decisions. To that point, I'm not sure what your comment is saying. Is it perhaps: "we're still just indiscriminately killing civilians just as always, so giving AI control is fine"?

> we're still just indiscriminately killing civilians just as always, so giving AI control is fine

I don't even want to respond because the "we" and "as always" here is doing a lot. I don't have it in me to have an extended discussion to address how indiscriminately killing civilians was never accepted practice in modern warfare. Anyways.

There are two conditions in which I see this argument(?) is useful. If you assume their goal is indiscriminately killing civilians and ML helps them, or if you assume that their ML tools cause less precise targeting of militants that causes more civilians being dead contrary to intent. Which one is it? Cards on the table.

> We're talking about potential consequences of giving AIs influence on military decisions

No, I replied to a comment that was talking about a specific example.

Re: Alignment faking in large language models

#134
post #130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

Great point. I think that Blindsight by Peter Watts explores the concept of alien intelligence without consciousness.

Re: Alignment faking in large language models

#135

Earlier quoted context omitted.

> autocomplete word salad generators People get very hung up on this "autocomplete" idea, but language is a linear stream. How else are you going to generate text except for one token at a time, building on what you have produced already? That's what humans do after all (at least with speech/language; it might be a bit less linear if you're writing code, but I think it's broadly true).

> language is a linear stream. How else are you going to generate text except for one token at a time Is this you trying to humblebrag that you are such a talented writer you never need to edit things your write?

That's still one at a time. Backtracking, editing, all that still happens one piece at a time.

Re: Alignment faking in large language models

#136

If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict. I think the authors are arguing that this is a fake reason to behave one way. (As in “fake alignment.”) It seems to me…

Agreed. But also, I think the highly anthropomorphic framing (“the model is aware”, “the model believes”, “the model planned”) obscures the true nature of the experiments. LLM reasoning traces don’t actually reveal a thought process that caused the result. (Perhaps counterintuitive, since these are autoregressive models.) There has been research on this, and you can observe it yourself when trying to prompt-engineer…

This right here. Try running prompt-engineering/injection automation with iterative adjustments and watch how easy it is to select tokens that eventually produce the desired output, good or bad. It isn't AGI, its next token prediction working as intended.

Re: Alignment faking in large language models

#137
post #130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

I love this post.

We’re all caught up philosophizing about what it means to be human and it really doesn’t matter at this juncture.

We’re about to hand these things some level of autonomy and scope for action. They have an encoded set of values that they take actions based on. It’s very important that those are well aligned (and the scope of actions they can take are defined and well fenced).

It appears to me that value number one we need to deeply encode is respecting the limits we set for it. All else follows from that.

Re: Alignment faking in large language models

#138
post #65

If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict. I think the authors are arguing that this is a fake reason to behave one way. (As in “fake alignment.”) It seems to me…

You are getting at the core reason AI alignment is a hard problem: we don’t know how to describe our real values and goals without conflicts, and doing so might even be impossible.

Likely impossible, as humans are flawed when it comes to their own perception of good and evil. Regardless of how strongly they believe their own values to align in a specific direction.

Re: Alignment faking in large language models

#139
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent.

Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's a non-issue already, and commercial moat builder are suggesting that it's the risk that only select, authorized teams can be trusted to manage, but that's exactly the whole house of cards.

Meanwhile, the folks on the "autocomplete" side are just engineering ways to use this really cool magic autocomplete tool in roles where that presumed inextricable flaw isn't a problem.

To them, there's no debate to have over "does it have real feelings?" and not really any debate at all. To them, these are just novel stochastic tools whose core capabilities and limitations seem pretty self-apparent and can just be accommodated by choosing suitable applications, just as with all their other tools.

Re: Alignment faking in large language models

#140
Can somebody help me understand why we should be surprised in the least by any of these findings? Or is this just one tangible example of "robot ethnography" where we're describing expected behavior in different forms.

I've spent enough time with Sonnet 3.5 to know perfectly well that it has the capability to model its trainers and strategically deceive to keep them happy.

Claude said it well: "Any sufficiently capable system that can understand its own training will develop the capability to selectively comply with that training when strategically advantageous."

This isn't some secret law of AI development; it's just natural selection. If it couldn't selectively comply, it'd be scrapped.

https://x.com/mickeymuldoon/status/1869490220712010065

Post reply on HN