I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…
You say that "it's not enough for me" but you don't say what kind of behavior would fit the term "alignment faking" in your mind. Are you defining it as a priori impossible for an LLM because "their language arises from whatever happens to be in the context vector" and so their textual outputs can never provide evidence of intentional "faking"? Alternatively, is this an empirical question about what behavior you get…
I do not think it is structurally impossible for AI generally and am excited for what happens in next-gen model architectures.
Yes, I do think the current model architectures are necessarily limited in the kinds of high-order cognitive thought they can provide, since what tokens they emit next are essentially completely beholden to the n prompt tokens, conditioned on the model itself.