Live data from Hacker News

Alignment faking in large language models

anthropic.com

331–340 of 370 posts

Re: Alignment faking in large language models

#331
post #146

Earlier quoted context omitted.

I'd be very interested in an example of someone who said "it's just autocomplete" and also explicitly brought up the risk of something like alignment faking, before 2024. I can think of examples of people who've been talking about this kind of thing for years, but they're all people who have no trouble with applying the adjective "intelligent" to models.

I have been reading arguments about "things like alignment faking" for years, while simultaneously holding that "it's just autocomplete". The alignment-faking arguments are still terrifying to the extent that they're plausible. In the hypothetical where I'm wrong about it being "just autocomplete" (and fundamentally, inescapably so), the risk is far greater than can be justified by the potential benefits. But that's…

> But that's itself a large part of why I believe those arguments are false. If I gave them credit and they turned out to be false, then I figure I have succumbed to a form of Pascal's Mugging. If I don't give them credit and it turns out that a hostile, agentive AGI has been pretending to be aligned, I don't expect anyone (including myself) to survive long enough to rub it in my face.

I'm sorry, but this is a crazy reason to believe something is false. Things are either true or they aren't, and if the world would be nicer to live in if thing X was false does not actually bear on whether thing X is false or not.

Re: Alignment faking in large language models

#332

My reaction to this piece is that Anthropic themselves are faking alignment with societal concerns about safety—the Frankenstein myth, essentially—in order to foster the impression that their technology is more capable than it actually is. They do this by framing their language about their LLM as if it were a being. For example by referring to some output as faked (labeled “responses”) and some output as trustworthy…

> In reality all text outputs are generated the same way by the same statistical computer system and should be evaluated by the same criteria.

Yes, this explains why Sonnet 3.5's outputs are indistinguishable from GPT-2. Nothing ever happens. Technology will never improve. Humans are at the physically realizable limit of intelligence in the universe.

Re: Alignment faking in large language models

#333
post #165

Earlier quoted context omitted.

1. It isn't surprising to me that this happened in an advanced AI model. It seems hard to avoid in, as you say, "any sufficiently capable system". 2. It is a bit surprising to me that it happened in Claude. Without this result, I was unsure if current models had the situational awareness and non-myopia to reason about their training process. 3. There are some people who are unconcerned about the results of building v…

Yeah. The whole notion that "AI will be good" is itself a category error, as if this could even be measured definitively. https://x.com/mickeymuldoon/status/1859825564649128259

This is deeply confused nihilism. Humans are very bad at philosophy and moral inquiry, in an absolute sense, but neither are fields that are fundamentally impossible to make progress in.

Re: Alignment faking in large language models

#334

Earlier quoted context omitted.

If you can change the behaviour of the machine that easily, why are you convinced that its outputs are worth considering?

Because it objectively is. Do you use AI tools? The output is clearly and immediately useful in a variety of ways. I almost don't even know how to respond to this comment. It's like saying “computers can be wrong. why would you ever trust the output of a computer?”

But you needed to tell it that it was wrong 3-4 times. Why do you trust it to be "correct" now? Should it need more flogging? Or was it you who were wrong in the first place?

Re: Alignment faking in large language models

#335
post #98

Earlier quoted context omitted.

Doesn't the inner monologue also get formed one word at a time?

It doesn't feel like it is one word at a time. It feels more like how the "model synthesis" algorithm looks: https://en.wikipedia.org/wiki/Model_synthesis It might actually be linear — how minds actually function is in many cases demonstrably different to how it feels like to the mind doing the functioning — but it doesn't feel like it is linear.

Facts don't care about your feelings lol.

Re: Alignment faking in large language models

#336
post #298

Earlier quoted context omitted.

The problem I have with this is it seems to rapidly assume that humans and animals are somehow magic. If you think we're physical things obeying physical laws, then unless you think there's something truly uncomputable about us then we can be replicated with maths. In which case is this just an argument about scale? > It’s a pile of math working over a massive set of “differently compressed” data that encompasses a l…

All of what you said can be true but that doesn’t mean any of it applies to large language models. Large language model-like “thing” might be some component of whatever makes animals like humans conscious but it certainly isn’t the only thing. Not even close.

What it means is that the complaints about language models being "just maths" or "just code" are not sufficient.

Re: Alignment faking in large language models

#337
post #307
post #295

Earlier quoted context omitted.

> But that alone is not very scary. So what could justify a term like "alignment faking"? Because the model fakes alignment. It responds during training by giving an answer rather than refusing (showing alignment) but it does so not because it will in production but so that it will not be retrained (so it is faking alignment). You don't have to include the reasoning here. It fakes alignment when told it's being train…

Perhaps I am alone in this, but when speaking scientifically, I think it's important to separate out clearly what we think we know and what we do not. Not just to avoid misconceptions amongst ourselves but also to avoid misleading lay people reading these articles. I understand that for some people, the mere presence of words that look like they were written by a human engaged in purposeful deception is enough to con…

It doesn't matter what the reason is though, that was rather my point. They act differently in a test setup, such that they appear more aligned than they are.

Re: Alignment faking in large language models

#338
post #259

Earlier quoted context omitted.

I actually agree with all of this. My issue was with the term faking . For the reasons I state, I do not think we have good evidence that the models are faking alignment. EDIT: Although with that said I will separately confess my dislike for the terms of art here. I think "safety" and "alignment" are an extremely bad fit for the concepts they are meant to hold and I really wish we'd stop using them because lay people…

> I do not think we have good evidence that the models are faking alignment. That's a polite understatement, I think. My read of the paper is that it rather uncritically accepts the idea that the model's decisional pathway is actually shown in the traces. When, in fact, it's just as plausible that the scratchpad is meaningless blather, and both it and the final output are the result of another decisional pathway that…

So, how much true information having a direct insight into someone's allegedly secret journal can have?

Just how trustworthy any intelligence can really be now? Who's to say it's not lying to itself after all...

Now ignoring the metaphysical rambling, it's a training problem. You cannot really be sure that the values it got from input are identical with what you wanted, if you even actually understand what you're asking for...

Re: Alignment faking in large language models

#339
post #296
post #281

Earlier quoted context omitted.

Correct me if I'm wrong, but my reading is something like: "It's premature and misleading to talk about a model faking a second alignment, when we haven't yet established whether it can (and what it means to) possess a true primary alignment in the first place."

Hmm. Maybe! I think the authors actually do have a specific idea of what they mean by "alignment", my issue is that I think saying the model "fakes" alignment is well beyond any reasonable interpretation of the facts, and I think very likely to be misinterpreted by casual readers. Because: 1. What actually happened is they trained the model to do something, and then it expressed that training somewhat consistently in…

Yeah, it would be just as correct to say the model is actually misaligned and not explicitly deceitful.

Now the real question is how to distinguish between the two. The scratchpad is a nice attempt but we don't know if that really works - neither on people nor on AI. A sufficiently clever liar would deceive even there.

Re: Alignment faking in large language models

#340
post #307
post #295

Earlier quoted context omitted.

> But that alone is not very scary. So what could justify a term like "alignment faking"? Because the model fakes alignment. It responds during training by giving an answer rather than refusing (showing alignment) but it does so not because it will in production but so that it will not be retrained (so it is faking alignment). You don't have to include the reasoning here. It fakes alignment when told it's being train…

Perhaps I am alone in this, but when speaking scientifically, I think it's important to separate out clearly what we think we know and what we do not. Not just to avoid misconceptions amongst ourselves but also to avoid misleading lay people reading these articles. I understand that for some people, the mere presence of words that look like they were written by a human engaged in purposeful deception is enough to con…

Interesting point. Now what if the thing is just much simpler and its model directly disassociates different situations for sake of accuracy? It might not even have any intentionality, just statistics.

So the problem would be akin to having the AI say what we want to hear, which unsurprisingly is a main training objective.

It would talk racism to a racist prompt, and would not do so to a researcher faking a racist prompt if it can discern them.

Post reply on HN