Alignment faking in large language models
11–20 of 370 posts
Re: Alignment faking in large language models
#12I think a more convincing example was where they used fine tuning to make a llm lie. They then look at some of the inner nodes. They could tell the llm knew the truth internally but switched outputs right at the end to lie.
Re: Alignment faking in large language models
#13Re: Alignment faking in large language models
#14This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.
Re: Alignment faking in large language models
#15I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
Also I think the model is to a large extent "talking itself into" the alignment faking in its scratchpad. These language models are really suggestible, and in my experience prone to self-suggestion as well due to their auto-regressive nature. Section 3.3.3 of the full paper addresses this a bit - there's way less faking when the CoT scratchpad is removed.
Re: Alignment faking in large language models
#16I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
Re: Alignment faking in large language models
#17My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…
It seems that you can “convince” LLMs of almost anything if you are insistent enough.
Albeit they'd be much more annoying for humans to use, because feelings.
Re: Alignment faking in large language models
#18I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
If the only way you can experience the world is Unicode text, how are you supposed to know what is real? While we’re at it, how can I tell that you aren’t a word salad generator?
Re: Alignment faking in large language models
#19My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…
It seems that you can “convince” LLMs of almost anything if you are insistent enough.
Convince an entity require this entity to have axiological feelings. Then to convince it, you either have to persuade the entity that the demand fits its ethos, or lead it to operate against its own inner values, or to go through a major change of values.
Re: Alignment faking in large language models
#20I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
If the only way you can experience the world is Unicode text, how are you supposed to know what is real? While we’re at it, how can I tell that you aren’t a word salad generator?
Observation of people who suffered from brain injury that resulted in a more or less unmediated flow from the language generation areas all through vocalization shows that we can also produce grammatically coherent speech that lacks deeper rationality