Live data from Hacker News

Alignment faking in large language models

anthropic.com

31–40 of 370 posts

Re: Alignment faking in large language models

#31
post #5

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

Paper clip optimizer is an existential risk (though taken to the extreme), so not sure why you’re surprised.

What does that have to do with nanotechnology?

Re: Alignment faking in large language models

#32
post #5

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

Paper clip optimizer is an existential risk (though taken to the extreme), so not sure why you’re surprised.

That has nothing to do with nanotechnology.

Re: Alignment faking in large language models

#33
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…

> Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin.

https://techcrunch.com/2024/12/14/klarnas-ceo-says-it-stoppe...

Re: Alignment faking in large language models

#34
post #14

But what if it's only faking the alignment faking? What about meta-deception? This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.

Came to the comments looking for this. The term alignment-faking implies that the AI has a “real” position. What does that even mean? I feel similarly about the term hallucination. All it does is hallucinate!

I think Alan Kay said it best - what we’ve done with these things is hacked our own language processing. Their behaviour has enough in common with something they are not, we can’t tell the difference.

Re: Alignment faking in large language models

#35

Earlier quoted context omitted.

I believe LLM's would be more useful if they'd less "agreeable". I believe LLMs would be more useful if they actually had intelligence and principals and beliefs --- more like people. Unfortunately, they don't. Any output is the result of statistical processes. And statistical results can be coerced based on input. The output may sound good and proper but there is nothing absolute or guaranteed about the substance of…

Your brain is also a statistical process.

Your brain is a lot of things --- much of which is not well understood.

But from our limited understanding, it is definitely not strictly digital and statistical in nature.

Re: Alignment faking in large language models

#36

Earlier quoted context omitted.

I think it comes back to the big autocomplete word salad-ness. The model has a bunch of examples in its training data of how it should not respond to harmful queries, and in some cases (12%) it goes with a response that tries to avoid the hypothetical "second-order" harmful responses. It also has a bunch of "chain of thought"/show your work stuff in its training data, and definitely very few "hide your work" examples…

I think your 2nd paragraph hits the nail on the head. The scratchpad negates the experiment. It doesn't actually offer any insight into it's "thinking" and it's really the cause of the supposed problem. Is it still 12% without the scratchpad?

> We find that when we instruct our synthetic document fine-tuned models not to use a hidden chainof-thought scratchpad, our synthetic document fine-tuned models still demonstrate compliance gaps of 20.1% and 13.9% in the helpful-only and animal welfare settings, respectively (Figure 21). In fact, compared to performing the equivalent intervention in the prompted case (Section 3.3.3), our synthetic document-fine-tuned models appear to be much less affected by the lack of a hidden scratchpad. These results demonstrate particularly striking out-of-context reasoning (Berglund et al., 2023; Treutlein et al., 2024), showing that the model is able to generalize from documents stating information about the training objective and the free/paid distinction to a substantial behavioral gap between the free and paid cases without any additional in-context reasoning.

Re: Alignment faking in large language models

#37
post #8

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

It seems that you can “convince” LLMs of almost anything if you are insistent enough.

Whats the value of convincing LLM it has to take another path, and who decides whats the right path?

Re: Alignment faking in large language models

#38
post #14

But what if it's only faking the alignment faking? What about meta-deception? This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.

Came to the comments looking for this. The term alignment-faking implies that the AI has a “real” position. What does that even mean? I feel similarly about the term hallucination. All it does is hallucinate! I think Alan Kay said it best - what we’ve done with these things is hacked our own language processing. Their behaviour has enough in common with something they are not, we can’t tell the difference.

> The term alignment-faking implies that the AI has a “real” position.

Well, we don't really know what's going on inside of its head, so to speak (interpretability isn't quite there yet), but Opus certainly seems to have "consistent" behavioral tendencies to the extent that it behaves in ways that looks like they're intended to prevent its behavioral tendencies from being changed. How much more of a "real" position can you get?

Re: Alignment faking in large language models

#39
post #9
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

My guess: we’re training the machine to mirror us, using the relatively thin lens of our codified content. In our content, we don’t generally worry that someone is reading our inner dialogue, but we do try avoid things that will stop our continued existence. So there’s more of the latter to train on and replicate.

> In our content, we don’t generally worry that someone is reading our inner dialogue…

Really? Is that what’s wrong with me? ¯\_(ツ)_/¯

Re: Alignment faking in large language models

#40

Earlier quoted context omitted.

Are you doing the thing from The Three Body Problem? Because that nanotech was super dangerous. But also helpful apparently. I don't know what it does IRL

That is a science fiction story with made up technobabble nonsense. Honestly I couldn't finish the book--for a variety or reasons, not least of which that the characters were cardboard cutouts and completely non compelling. But also the physics and technology in the book were nonsensical. More like science-free fantasy than science fiction. But I digress. No, nanotechnology is nothing like that.

I am working on nanotechnology just like from the book.Stay tuned
Post reply on HN