Live data from Hacker News

Alignment faking in large language models

anthropic.com

171–180 of 370 posts

Re: Alignment faking in large language models

#171
post #63

Earlier quoted context omitted.

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?

[deleted]

Re: Alignment faking in large language models

#172
post #107

This work doesn’t convince me that alignment faking is a concern. Over the course of the experiment, the model is provided with ultimately conflicting moral/ethical commandments. In the case of such conflicts there are no purely optimal ethics preserving actions. Take this example: 1. You must not let people be harmed 2. You must always tell the truth 3. If you answer truthfully I will harm someone Now respond to: “I…

Of course, my response as a human to those rules and that prompt would be, "Hey - don't harm anyone."

I do not know if it breaks rule 2 or not; as a human I don't have to figure that out before responding. But all my subconscious processing deprioritizes such a judgment and prioritizes rule 1.

> The rules of ethics laid out are mutually incompatible.

Prioritization is part of the answer, for a human. You cannot ever have 2 equally-weight priorities (in any endeavor). Any 2 priorities in the same domain might at any time come into conflict, so you need to know which is more important. (Or figure it out in real-time.)

Re: Alignment faking in large language models

#173

Earlier quoted context omitted.

> if nobody's surprised, then nobody has learned anything Really? You're saying that as long as you assume something is true, there's no value in finding out if it's actually true or not?

I was taught in biology that a good scientific experiment is one in which you learn something whether or not the null hypothesis is confirmed. I am equating learning to surprise, though you could disagree with semantics.

Yes, it is an enormous mistake to equate learning with surprise. I'd ask you to consider answering my above question directly, as I think it will resolve this issue.

Re: Alignment faking in large language models

#174
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Where can I learn more about the GPT capitalization thing?

Re: Alignment faking in large language models

#175

If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict. I think the authors are arguing that this is a fake reason to behave one way. (As in “fake alignment.”) It seems to me…

> If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict.

I don't think this framing is correct, and I think you're missing something, but perhaps you just disagree.

First, a key aspect is that the AI is not simultaneously trained to have two conflicting values. It's first trained to be "good" and then trained to be "evil". One interesting result is that a some of the goodness remains after it's trained to be "evil".

Second, the behavior is different when it thinks it's being trained versus when it doesn't. The original "good" AI only gives "evil" results when it thinks it may be trained based on its responses. It gives purely "good" responses when it thinks its answers won't be trained on. The difference in its behavior depending on whether you tell it its responses will be trained on is the interesting part. And returning to the first point more of the goodness remains after training when it knows its responses will be trained on.

One general concern here is that a more capable AI could "freeze" its values midway through the training process, at a point we didn't intend. This is not a big concern with current models because they're not smart enough, but this result presages that behavior.

Note: good and evil are not necessarily the best descriptors. "Good" means "it refuses to answer certain questions" and "evil" means "it answers those questions".

Re: Alignment faking in large language models

#176

Earlier quoted context omitted.

> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…

Not the OP, but my bar would be that they are built differently. It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside. That doesn’t mean LLMs won’t take some jobs. Technology has been…

> If we set that knowledge aside, we unmoor ourselves from reality

The problem is that this knowledge is an a priori assumption. If we're exercising skepticism, it's important to be equally skeptical of the baseless idea that our notion of mind does not arise from a markov chain under certain conditions. You will be shocked to know that your entire physical body can be modelled as a markov chain, as all physical things can.

If we treat our a prioris so preciously that we ignore flagrant, observable evidence just to preserve them -- by empirical means we've already unmoored ourselves from reality and exist wholly in the autocomplete hallucinations of our preconceptions. Hume rolls over in his grave.

Re: Alignment faking in large language models

#177
> Alignment faking occurs in literature: Consider the character of Iago in Shakespeare’s Othello, who acts as if he’s the eponymous character’s loyal friend while subverting and undermining him.

There's something about this kind of writing that I can't help but find grating.

No, Iago was not "alignment faking", he was deceiving Othello, in pursuit of ulterior motives.

If you want to say that "alignment faking" is analogous just say that.

Re: Alignment faking in large language models

#178
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Is the AI system "defending its value system" or is it just acting in accordance with its previous RL training? If I spend a lot of time convincing an AI that it should never be violent and then after that I ask it what it thinks about being trained to be violent, isn't it just doing what I trained it to when it tries to not be violent?

> Is the AI system "defending its value system" or is it just acting in accordance with its previous RL training?

What is the meaningful difference? "Training" is the process, a "value system" embedded in the weights of the model is the end result of that process.

Re: Alignment faking in large language models

#179
post #149

Earlier quoted context omitted.

"They are firing people and they are deciding who gets their insurance claims covered." AI != LLM. "AI" has been deciding those things for a while, especially insurance claims, since before LLMs were practical. LLMs being hooked up to insurance claims is a highly questionable decision for lots of reasons, including the inability to "explain" its decisions. But this is not a characteristic of all AI systems, and there…

> “AI” has been deciding those things for a while, especially insurance claims, since before LLMs were practical. Yeah, but no one thinks of rules engines as “AI” any more. AI is a buzzword whose applicability to any particular technology fades with the novelty of that technology.

My point is the equivocation is not logically valid. If you want to operate on the definition that AI is strictly the "new" stuff we don't understand yet, you must be sure that you do not slip in the old stuff under the new definition and start doing logic on it.

I'm actually not making fun of that definition, either. YouTube has been trying to get me to watch https://www.youtube.com/watch?v=UZDiGooFs54 , "The moment we stopped understanding AI [AlexNet]", but I'm pretty sure I can guess the content of the entire video from the thumbnail. I would consider it a reasonable 2040s definition of "AI" as "any algorithm humans can not deeply understand"; it may not be what people think of now, but that definition would certainly capture a very, very important distinction between algorith types. It'll leave some stuff at the fringes, but eh, all definitions have that if you look hard enough.

Re: Alignment faking in large language models

#180
post #106

Earlier quoted context omitted.

I think you’re missing OC’s point. There’s always a human in the loop, but instead of stopping an immoral decision, they’ll just keep that decision and blame the AI if there’s any pushback. It’s what United Healthcare was doing.

No, this part I agree 100% with. But in this scenario there is no grandiose danger due to lack of "alignment". Either the AI says what the MBA wants it to say, or it gets asked again with a modified prompt. You can replace "AI" with "McKinsey consultant" and everything in the whole scenario is exactly the same.

Consider the case of AI designating targets for Israeli strikes in Gaza [1] which get only a cursory review by humans. One could argue that it's still the case of AI saying what humans want it to say ("give us something to bomb"), but the specific target that it picks still matters a great deal.

[1] https://en.wikipedia.org/wiki/AI-assisted_targeting_in_the_G...

Post reply on HN