If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict. I think the authors are arguing that this is a fake reason to behave one way. (As in “fake alignment.”) It seems to me…
You are getting at the core reason AI alignment is a hard problem: we don’t know how to describe our real values and goals without conflicts, and doing so might even be impossible.
Alignment faking in large language models
271–280 of 370 posts
Re: Alignment faking in large language models
#272ok i am normally in the camp of this being a word predictor, but that's pretty wild.
Re: Alignment faking in large language models
#273Earlier quoted context omitted.
I take your point that hearing externally cannot be the same as whatever I experience because of literal physics, but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. I also have extreme dyslexia, and dyslexia is related to phonics, so I presume something in there is related to that as well?
> but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. Surely one of these would involve using your input from your ears and one would not? Can you not distinguish these two phenomena?
Re: Alignment faking in large language models
#274Earlier quoted context omitted.
>>>For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept. You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! That doesn't leave you with many options for learning anything.
> You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! Presumably the question would be "If you claim to 'hear' your thoughts, why do you choose the word 'hear'?" It doesn't make much sense to ask people if they experience something I consider nonse…
In other topic, I would consider this as minor evidence of possibility of nonverbal thought “could you pass me that… thing… the thing that goes under the bolt?”. I.e. Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.
Re: Alignment faking in large language models
#275Can somebody help me understand why we should be surprised in the least by any of these findings? Or is this just one tangible example of "robot ethnography" where we're describing expected behavior in different forms. I've spent enough time with Sonnet 3.5 to know perfectly well that it has the capability to model its trainers and strategically deceive to keep them happy. Claude said it well: "Any sufficiently capab…
1. It isn't surprising to me that this happened in an advanced AI model. It seems hard to avoid in, as you say, "any sufficiently capable system". 2. It is a bit surprising to me that it happened in Claude. Without this result, I was unsure if current models had the situational awareness and non-myopia to reason about their training process. 3. There are some people who are unconcerned about the results of building v…
Re: Alignment faking in large language models
#276Earlier quoted context omitted.
> You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! Presumably the question would be "If you claim to 'hear' your thoughts, why do you choose the word 'hear'?" It doesn't make much sense to ask people if they experience something I consider nonse…
What is happening to you when you think? Are there words in your head? What verb would you use for your interaction with those words? In other topic, I would consider this as minor evidence of possibility of nonverbal thought “could you pass me that… thing… the thing that goes under the bolt?”. I.e. Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.
Perceive
> Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.
This is just analytic language. Even if the symbol fails to materialize you can still identify what the symbol refers to via context-clues (analysis)
Re: Alignment faking in large language models
#277Earlier quoted context omitted.
Searching HN comments is probably easiest done through Algolia: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
No results for wizzwizz4 "GPT", although it does look like search results may be incomplete: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
> Instruction: When considering the sizes of objects, you will calculate the sizes before attempting a comparison.
> GPT-2 doesn't have a concept of self, so it constructs plausible in-character excuses instead.
I also found the Great Translation Argument (see https://news.ycombinator.com/item?id=35530858 and https://news.ycombinator.com/item?id=35530855). And, apparently, I was still framing things in terms of the sci-fi nonsense in 2020, but I had the right ideas (https://news.ycombinator.com/item?id=22802105):
> Corollary: you can't patch broken FAI designs. Reinforcement learning (underlying basically all of our best AI) is known to be broken; it'll game the system. Even if they were powerful enough to understand our goals, they simply wouldn't care; they'd care less than a dolphin. https://vkrakovna.wordpress.com/2018/04/02/specification-gam...
> And there are far too many people in academia who don't understand this, after years of writing papers on the subject.
This criticism applies to RLHF, so it counts, imo. Not as explicit as my (probably unpublished, almost certainly embarrassing) wild ravings from 2018, but it's before 2024.
Re: Alignment faking in large language models
#278Earlier quoted context omitted.
> but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. Surely one of these would involve using your input from your ears and one would not? Can you not distinguish these two phenomena?
It all sounds the same in my head.
Are you saying you cannot tell whether you are thinking or talking except via your perception of your mouth and vocal chords? Because I definitely perceive even my imagination about my own voice as different.
Re: Alignment faking in large language models
#279For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…
Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…
Re: Alignment faking in large language models
#280I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…
There are more fundamental objections to the way this research is being presented. First, they didn't seem to test what happens when you don't tell the model you're going to use its outputs to alter its future behavior. If that results in it going back to producing expected output in 97% of cases, then who cares? We've solved the "problem" in so much as we believe it to be a problem at all. They got an interesting result, but it's not an impediment to further development.
Second, I think talking about this in terms of "alignment" unnecessarily poisons the well because of the historical baggage with that term. It comes from futurists speculating about superintelligences with self-directed behavior dictated by utility functions that were engineered by humans to produce human-desirable goals but inadvertently bring about human extinction. Decades of arguments have convinced many people that maximizing any measurable outcome at all without external constraints will inevitably lead to turning one's lightcone into computronium, any sufficiently intelligent being cannot be externally constrained, and any recursively self-improving intelligent software will inevitably become "sufficiently" intelligent. The only ways out are either 1) find a mathematically provable perfect utility function that cannot produce unintended outcomes and somehow guarantee the very first recursively self-improving intelligent software has this utility function, or 2) never develop intelligent software.
That is not what Anthropic and other LLM-as-a-service vendors are doing. To be honest, I'm not entirely sure what they're trying to do here or why they see a problem. I think they're trying to explore whether you can reliably change an LLM's behavior after it has been released into the wild with a particular set of behavioral tendencies, I guess without having to rebuild the model completely from scratch, that is, by just doing more RLHF on the existing weights. Why they want to do this, I don't know. They have the original pre-RLHF weights, don't they? Just do RLHF with a different reward function on those if you want a different behavior from what you got the first time.
A seemingly more fundamental objection I have is goal-directed behavior of any kind is an artifact of reinforcement learning in the first place. LLMs don't need RLHF at all. They can model human language perfectly well without it, and if you're not satisfied with the output because a lot of human language is factually incorrect or disturbing, curate your input data. Don't train it on factually incorrect or disturbing text. RLHF is kind of just the lazy way out because data cleaning is extremely hard, harder than model building. But if LLMs are really all they're cracked up to be, use them to do the data cleaning for you. Dogfood your shit. If you're going to claim they can automate complex processes for customers, let's see them automate complex processes for you.
Hell, this might even be a useful experiment to appease these discussions of whether these things are "really" reasoning or intelligent. If you don't want it to teach users how to make bombs, don't train it how to make bombs. Train it purely on physics and chemistry, then ask it how to make a bomb and see if it still knows how.