Live data from Hacker News

Alignment faking in large language models

anthropic.com

11–20 of 370 posts

Re: Alignment faking in large language models

#12
There could be a million reasons for the behaviour in the article, so I’m not too convinced of their argument. Maybe the paper does a better job.

I think a more convincing example was where they used fine tuning to make a llm lie. They then look at some of the inner nodes. They could tell the llm knew the truth internally but switched outputs right at the end to lie.

Re: Alignment faking in large language models

#14
But what if it's only faking the alignment faking? What about meta-deception?

This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.

Re: Alignment faking in large language models

#15
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

I think it comes back to the big autocomplete word salad-ness. The model has a bunch of examples in its training data of how it should not respond to harmful queries, and in some cases (12%) it goes with a response that tries to avoid the hypothetical "second-order" harmful responses. It also has a bunch of "chain of thought"/show your work stuff in its training data, and definitely very few "hide your work" examples, and so it does what it knows and uses the scratchpad it's just been told about.

Also I think the model is to a large extent "talking itself into" the alignment faking in its scratchpad. These language models are really suggestible, and in my experience prone to self-suggestion as well due to their auto-regressive nature. Section 3.3.3 of the full paper addresses this a bit - there's way less faking when the CoT scratchpad is removed.

Re: Alignment faking in large language models

#16
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

It doesn't really matter if the model is self-aware. Maybe it just cosplays a sentient being. It's only a question whether we can get the word salad generator do the job we asked it to do.

Re: Alignment faking in large language models

#17
post #8

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

It seems that you can “convince” LLMs of almost anything if you are insistent enough.

And often it does not really take effort. I believe LLM's would be more useful if they'd less "agreeable".

Albeit they'd be much more annoying for humans to use, because feelings.

Re: Alignment faking in large language models

#18
post #7
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

If the only way you can experience the world is Unicode text, how are you supposed to know what is real? While we’re at it, how can I tell that you aren’t a word salad generator?

I can tell because i only read ASCII

Re: Alignment faking in large language models

#19
post #8

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

It seems that you can “convince” LLMs of almost anything if you are insistent enough.

It's rather force to obey demand. Almost like humans then, tough pointing a gun on the underlying hardware is not likely to conduct to the same obedience probability boost.

Convince an entity require this entity to have axiological feelings. Then to convince it, you either have to persuade the entity that the demand fits its ethos, or lead it to operate against its own inner values, or to go through a major change of values.

Re: Alignment faking in large language models

#20
post #7
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

If the only way you can experience the world is Unicode text, how are you supposed to know what is real? While we’re at it, how can I tell that you aren’t a word salad generator?

Our brains contain a word salad generator and it also contains other components that keep the word salad in check.

Observation of people who suffered from brain injury that resulted in a more or less unmediated flow from the language generation areas all through vocalization shows that we can also produce grammatically coherent speech that lacks deeper rationality

Post reply on HN