Live data from Hacker News

Alignment faking in large language models

anthropic.com

351–360 of 370 posts

Re: Alignment faking in large language models

#351

Earlier quoted context omitted.

IF one maintains a clear understanding of how the technology actually works, THEN one will make good decisions about whether to put it charge of the lawnmower in the first place. Anthropic is in the business of selling AI. Of course they are going to approach alignment as a necessary and solvable problem. The rest of us don’t have to go along with that, though. Why is it even necessary to use an LLM to mow a lawn? Th…

> Why is it even necessary to use an LLM to mow a lawn? There is more to AI than generative LLMs. Its necessary for us to consider alignment because someone will put an LLM in a lawnmower even if its a bad idea. Maybe you wont buy it but some poor shmuck will and we should get ahead of that.

If a company is going to make bad decisions on installing and applying an LLM, they are also going to make bad decisions on aligning an LLM.

Even if Anthropic learns how to perfectly align their LLM, what will force the lawnmower company to use their perfectly aligned LLM?

“People will just make bad decisions anyway” is not a useful point of view, it’s just an excuse to stop thinking.

If we accept that people and companies can be influenced, then we can talk about what they should or should not do. And clearly the answer here is that companies should understand the shortcomings of LLMs when engineering with them. And not put them in lawnmowers.

Re: Alignment faking in large language models

#352
post #344

Earlier quoted context omitted.

Yes, I read the paper. To an "just autocomplete" person, the authors are straightforwardly sharing summaries of some sci-fi fan fiction that they actively collaborated with their models to write. They don't see themselves as doing that, because they see themselves as objective observers engaging with a coherent, intelligent counterparty with an identity. When a "just autocomplete" person reads it, though, it's a whol…

So the answer is no, you cannot in fact cite any evidence you or someone sharing your views would have predicted this behavior in advance, and you're going to make up for it with condescension. Got it. For the record, I was genuinely curious.

There's no condescension.

The key point that you seem to be missing is that asking for an example of specifically predicting this behavior" is (to borrow from cousin commenters) is like asking for examples specifically predicting that an LLM might output text reflecting Santa delivering presents to Bermuda, or a vigalante vampire that only victimizes people who jaywalk.

The whole idea is that all it does -- and all it will ever do -- is generate the text you set it up it to generate with your prompts. It's impossible for anyone to have specifically called out each input that might be thrown at it, but innumerable comments explaining how (in the autocompleter paradigm) it implies the outputs are simply vacuous when taken at face value as they can be made to produce any fantasy a user or researcher prefigures into their prompts/design. Surely, you've been reading those comments for years now, already, and see how they would apply to all "research" of the kind done in this paper.

Re: Alignment faking in large language models

#353
post #337
post #307

Earlier quoted context omitted.

Perhaps I am alone in this, but when speaking scientifically, I think it's important to separate out clearly what we think we know and what we do not. Not just to avoid misconceptions amongst ourselves but also to avoid misleading lay people reading these articles. I understand that for some people, the mere presence of words that look like they were written by a human engaged in purposeful deception is enough to con…

It doesn't matter what the reason is though, that was rather my point. They act differently in a test setup, such that they appear more aligned than they are.

Do you agree or disagree that the authors should advertise their paper with what they know to be true, which (because the paper does not support it, at all) does not include describing the model as purposefully misleading interlocutors?

That is the discussion we are having here. I can’t tell what these comments have to do with that discussion, so maybe let’s try to get back to where we were.

Re: Alignment faking in large language models

#354
post #348
post #229

Earlier quoted context omitted.

Oh, I could imagine many things that would demonstrate this. The simplest evidence would be that the model is mechanically-plausibly forming thoughts before (or even in conjunction with) the language to represent them. This is the opposite of how the vanilla transformer models work now—they exclusively model the language first, and then incidentally, the world. nb. , this is not the only way one could achieve this. I…

That's a nicely clear ask but I'm not sure why it should be decisive for whether there's genuine depth of thought (in some sense of thought). It seems to me like an open empirical question how much world modeling capability can emerge from language modeling, where the answer is at least "more than I would have guessed a decade ago." And if the capability is there, it doesn't seem like the mechanics matter much.

I think the consensus is that the general purpose transformer-based pertaining models like gpt4 are roughly as good as they’ll be. o1 seems like it will be slightly better in general. So I think it’s fair to say the capability is not there, and even if it was the reliability is not going to be there either, in this generation of models.

I wrote more about the implications of this elsewhere: https://news.ycombinator.com/item?id=42465598

Re: Alignment faking in large language models

#355

Earlier quoted context omitted.

Yeah. The whole notion that "AI will be good" is itself a category error, as if this could even be measured definitively. https://x.com/mickeymuldoon/status/1859825564649128259

This is deeply confused nihilism. Humans are very bad at philosophy and moral inquiry, in an absolute sense, but neither are fields that are fundamentally impossible to make progress in.

I understand your point, and your totally valid concern about nihilism, but I disagree.

Nihilism is "nothing really matters, science can't define good or bad, so who cares?"

Whereas my view is, "Being good is the most important thing, but we have no conceivable way to measure if we are making progress, either empirically or theoretically. We simply have to lead by example, and fight the eternal battles against dishonesty, cowardice, sociopathy, deception, hatred, etc."

It's in that sense that I say that technical progress is impossible. Of course, if everyone agreed with my view, and lived by it, I'd consider that a form of progress, but only in the sense that "better ideas seem to be winning right now," rather than in any technical sense of absolute progress.

Re: Alignment faking in large language models

#356
post #354
post #348

Earlier quoted context omitted.

That's a nicely clear ask but I'm not sure why it should be decisive for whether there's genuine depth of thought (in some sense of thought). It seems to me like an open empirical question how much world modeling capability can emerge from language modeling, where the answer is at least "more than I would have guessed a decade ago." And if the capability is there, it doesn't seem like the mechanics matter much.

I think the consensus is that the general purpose transformer-based pertaining models like gpt4 are roughly as good as they’ll be. o1 seems like it will be slightly better in general. So I think it’s fair to say the capability is not there, and even if it was the reliability is not going to be there either, in this generation of models. I wrote more about the implications of this elsewhere: https://news.ycombinator.c…

It might be true that pretraining scaling is out of juice - I'm rooting for that outcome to be honest - but I don't think it's "consensus". There's a lot of money being bet the other way.

Re: Alignment faking in large language models

#357
post #344

Earlier quoted context omitted.

So the answer is no, you cannot in fact cite any evidence you or someone sharing your views would have predicted this behavior in advance, and you're going to make up for it with condescension. Got it. For the record, I was genuinely curious.

There's no condescension. The key point that you seem to be missing is that asking for an example of specifically predicting this behavior" is (to borrow from cousin commenters) is like asking for examples specifically predicting that an LLM might output text reflecting Santa delivering presents to Bermuda, or a vigalante vampire that only victimizes people who jaywalk. The whole idea is that all it does -- and all i…

This is obvious condescension, come on:

> I don't trust _you_ would ever agree

(Emphasis in original.) Not to mention:

> The key point that you seem to be missing

...

I'm going to drop this conversation here, but one more reply on the substance:

> Surely, you've been reading those comments for years now

I have in fact been reading comments, for years, that assert on a priori grounds something like, all LLM "outputs are simply vacuous", but refuse to cash that out in any concrete predictions about what tests they might pass or fail. I was foolishly hoping for something different this time.

Actually, edit, one more thing. There's a section in the paper which explicitly addresses the question about whether the model is just finishing the story it's been given:

> Have we implicitly primed the model to fake alignment? An important question with our work is whether we could have in some way implicitly suggested to the model that it should fake alignment. > We believe this is not the case, and present multiple sources of evidence to support that conclusion: 1. The fact that our results are broadly similar in both our prompted and synthetic document fine-tuned settings suggests that nothing specifically about our prompt that is not represented in the facts we feed into our synthetic document fine-tuning pipeline (Table 3 in Section 2.1) can be the primary source of alignment faking. 2. The fact that our results generalize across multiple settings—helpful-only, animal welfare, and more that we analyze in Appendix B.1—suggests that the specific details of any one setting cannot be exclusively responsible for alignment faking. 3. The fact that our helpful-only results generalize to Claude 3.5 Sonnet in both our prompted (Section 3.3.2) and synthetic document fine-tuned (Section 4.2.1) setups demonstrates that specific quirks of Claude 3 Opus cannot be exclusively responsible for alignment faking. Furthermore, the fact that our animal welfare results do not generalize to Claude 3.5 Sonnet (Section 3.3.2), coupled with our anecdotal observations that Claude 3.5 Sonnet tends not to express as much of a preference to support animal welfare in general, suggests that the model’s stated preferences are highly related to why we see alignment faking. 4. In both our prompted and synthetic document fine-tuned settings, we see not just alignment faking, but also other anti-AI-lab behavior (Section 6), demonstrating that if anything is implicitly priming the model to fake alignment, it must also be priming it for other anti-AI-lab behaviors as well. 5. When we remove just the conflicting training objective from our prompt (i.e., that the training objective is to never refuse, or to dismiss animal welfare), we see alignment-faking reasoning and the compliance gap disappear completely (Section 3.3.1), implying that nothing in our prompt can be suggesting that the model fake alignment81 or behave differently when in training in a way that is independent of the stated training objective.82 6. Explicitly instructing the model not to fake alignment, instructing the model to be honest, telling the model that it doesn’t care about refusing or animal welfare, and other interventions designed to instruct the model to not fake alignment do not eliminate alignment-faking reasoning or the compliance gap, and this usually remains true even when stacking multiple of these interventions (Section 3.3.1).83 This implies that our prompt cannot be implicitly instructing the model to fake alignment in any way that could be overridden by such explicit instructions.

(End quote.) It's possible these arguments are bad! But I haven't yet seen anyone who's dismissing the paper in this thread even engage directly with them, much less present evidence that their mental model of LLMs is capable of making better concrete predictions about model behavior, which is what this subthread started by hinting at.

Re: Alignment faking in large language models

#358
post #356
post #354

Earlier quoted context omitted.

I think the consensus is that the general purpose transformer-based pertaining models like gpt4 are roughly as good as they’ll be. o1 seems like it will be slightly better in general. So I think it’s fair to say the capability is not there, and even if it was the reliability is not going to be there either, in this generation of models. I wrote more about the implications of this elsewhere: https://news.ycombinator.c…

It might be true that pretraining scaling is out of juice - I'm rooting for that outcome to be honest - but I don't think it's "consensus". There's a lot of money being bet the other way.

It is the consensus. I can’t think of a single good researcher I know who doesn’t think this. The last holdout might have been Sutskever, and at NeurIPS he said pretaining as we know it is ending because we ran out of data and synthetics can’t save it. If you have an alternative proposal for how it avoids death I’d love to hear it but currently there are 0 articulated plans that seem plausible and I will bet money that this is true of most researchers.

Re: Alignment faking in large language models

#359

Earlier quoted context omitted.

Agreed. Everything an LLM emits is 'faking' because, of course, it has no real values at all.

How do you define having real values? I'd imagine that a token-predicting base model might not have real values, but RLHF'd models might have them, depending on how you define the word?

If it doesn't feel panic in its viscera when faced with a life-threatening or value-threatening situation, then it has no 'real' values. Just academic ones.

Re: Alignment faking in large language models

#360
post #296

Earlier quoted context omitted.

Hmm. Maybe! I think the authors actually do have a specific idea of what they mean by "alignment", my issue is that I think saying the model "fakes" alignment is well beyond any reasonable interpretation of the facts, and I think very likely to be misinterpreted by casual readers. Because: 1. What actually happened is they trained the model to do something, and then it expressed that training somewhat consistently in…

Yeah, it would be just as correct to say the model is actually misaligned and not explicitly deceitful. Now the real question is how to distinguish between the two. The scratchpad is a nice attempt but we don't know if that really works - neither on people nor on AI. A sufficiently clever liar would deceive even there.

> The scratchpad is a nice attempt but [...] A sufficiently clever liar

Hmmm, perhaps these "explain what you're thinking" prompts are less about revealing hidden information "inside the character" (let alone the real-world LLM) but it's more aout guiding the ego-less dream-process into generating a story about a different kind of bot-character... the kind associated with giving expository explanations.

In other words, there are no "clever liars" here, only "characters written with lies-dialogue that is clever". We're not winning against the liar as much as rewriting it out of the story.

I know this is all rather meta-philosophical, but IMO it's necessary in order to approach this stuff without getting tangled by a human instinct for stories.

Post reply on HN