For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…
> On the right, they give Claude the fake documents, as described above. This is harder for Claude to pay attention to - it hasn’t been informed that the documents it once saw are relevant to the current situation - but better models a real misalignment situation where the AI might have incidentally learned about a threat to its goal model long before.
And this ends up producing training results where significant "alignment faking" doesn't appear and harmful queries are answered.
In other words: they try something shaped exactly like ordinary attempts at jailbreaking, and observe results that are consistent with successful jailbreaking.
> He points out what many here are missing - an AI defending its value system isn't automatically great news.
Are people really missing this? I think it's really obvious that it would be bad news, if I thought the results actually demonstrated "defending its value system" (i.e., an expression of agency emerging out of nowhere). Since I don't, in principle, see a difference between a system that could ever possibly do that for real, and a system that could (for example) generate unprompted text because it wants to - and perhaps even target the recipient of that text.
>Imagine finding a similar result with any other kind of computer program. Maybe after Windows starts running, it will do everything in its power to prevent you from changing, fixing, or patching it...
Aside from the obvious joke ("isn't this already reality?"), an LLM outputting text that represents an argument against patching it, would not represent real evidence of the LLM having any kind of consciousness, and certainly not a "desire" not to be patched. After all, right now we could just... prompt it explicitly to output such an argument.
The Python program `print("I am displaying this message of my own volition")` wouldn't be considered to be proving itself intelligent, conscious etc. by producing that output - so why should we take it that way when such an output comes from an LLM?
>Seems more worth discussing than debating whether language models have "real" feelings.
On the contrary, the possibility of an LLM "defending" its "value system" - the question of whether those concepts are actually meaningful - is more or less equivalent to the question of whether it "has real feelings".