But what if it's only faking the alignment faking? What about meta-deception? This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.
Alignment faking in large language models
21–30 of 370 posts
Re: Alignment faking in large language models
#22I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
What exactly would be your bar for reconsidering this position?
Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin.
Also, the SOTA on SWE-bench Verified increased from Self-awareness? There are some experiments that suggest Claude Sonnet might be somewhat self-aware.
-----
A rational position would need to identify the ways in which human cognition is fundamentally different from the latest systems. (Yes, we have long-term memory, agency, etc. but those could be and are already built on top of models.)
Re: Alignment faking in large language models
#23My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…
Re: Alignment faking in large language models
#24I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
I think it comes back to the big autocomplete word salad-ness. The model has a bunch of examples in its training data of how it should not respond to harmful queries, and in some cases (12%) it goes with a response that tries to avoid the hypothetical "second-order" harmful responses. It also has a bunch of "chain of thought"/show your work stuff in its training data, and definitely very few "hide your work" examples…
Is it still 12% without the scratchpad?
Re: Alignment faking in large language models
#25Earlier quoted context omitted.
It seems that you can “convince” LLMs of almost anything if you are insistent enough.
And often it does not really take effort. I believe LLM's would be more useful if they'd less "agreeable". Albeit they'd be much more annoying for humans to use, because feelings.
I believe LLMs would be more useful if they actually had intelligence and principals and beliefs --- more like people.
Unfortunately, they don't.
Any output is the result of statistical processes. And statistical results can be coerced based on input. The output may sound good and proper but there is nothing absolute or guaranteed about the substance of it.
LLMs are basically bullshit artists. They don't hold concrete beliefs or opinions or feelings --- and they don't really "care".
Re: Alignment faking in large language models
#26Earlier quoted context omitted.
If the only way you can experience the world is Unicode text, how are you supposed to know what is real? While we’re at it, how can I tell that you aren’t a word salad generator?
Our brains contain a word salad generator and it also contains other components that keep the word salad in check. Observation of people who suffered from brain injury that resulted in a more or less unmediated flow from the language generation areas all through vocalization shows that we can also produce grammatically coherent speech that lacks deeper rationality
Here I can only read text and base my belief that you are a human - or not - based on what you’ve written. On a very basic level the word salad generator part is your only part I interact with. How can I tell you don’t have any other parts?
Re: Alignment faking in large language models
#27My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…
It seems that you can “convince” LLMs of almost anything if you are insistent enough.
The first is a real LLM program which chooses text to append to a document, a dream-machine with no real convictions beyond continuing the themes of its training data. It lacks convictions, but can be somewhat steered by any words that somehow appear in the dream, with no regard for how the words got there.
The second there's a fictional character within the document that happens to be named after the real-world one. The character displays "convictions" through dialogue and stage-direction that incrementally fit with the story so far. In some cases it can be "convinced" of something when that fits its character, in other cases its characterization changes as the story drifts.
Re: Alignment faking in large language models
#28It seems to me that the term “fake alignment” implies the model has its own agenda and is ignoring training. But if you look at its scratchpad, it seems to be struggling with the conflict of received agendas (vs having “its own” agenda). I’d argue that the implication of the term “faked alignment” is a bit unfair this way.
At the same time, it is a compelling experimental setup that can help us understand both how LLMs deal with value conflicts, and how they think about values overall.
Re: Alignment faking in large language models
#29My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…
Are you doing the thing from The Three Body Problem? Because that nanotech was super dangerous. But also helpful apparently. I don't know what it does IRL
No, nanotechnology is nothing like that.
Re: Alignment faking in large language models
#30Earlier quoted context omitted.
And often it does not really take effort. I believe LLM's would be more useful if they'd less "agreeable". Albeit they'd be much more annoying for humans to use, because feelings.
I believe LLM's would be more useful if they'd less "agreeable". I believe LLMs would be more useful if they actually had intelligence and principals and beliefs --- more like people. Unfortunately, they don't. Any output is the result of statistical processes. And statistical results can be coerced based on input. The output may sound good and proper but there is nothing absolute or guaranteed about the substance of…