Live data from Hacker News

Alignment faking in large language models

anthropic.com

241–250 of 370 posts

Re: Alignment faking in large language models

#241

Earlier quoted context omitted.

Where can I learn more about the GPT capitalization thing?

I suspect there's not much depth to it. Weird capitalization is unusual in ordinary text, but common in e.g. ransom notes - as well as sarcastic Internet mockery, of a sort that might be employed by people who lean towards anarchism, shall we say. Training is still fundamentally about associating tokens with other tokens, and the people doing RLHF to "teach" ChatGPT that crime is bad, wouldn't have touched the associ…

Basically this.

The mistake often made here is to think that LLMs emitting verbiage about crimes is some sort of problem in itself, that there's any conceivable way for it to feed back on the LLM. Like if it pretends to be a pirate, maybe the LLM will sail to Somalia and start boarding oil vessels.

It's not. It's a problem for OpenAI, entirely because they've decided they don't want their product talking about crimes. Makes sense for them, who needs the bad press and all, but the premise that a chatbot describing how to board a vessel off the Horn of Africa is a problem relative to aspiring human pirates watching documentary films on the subject is a bit nonsense to begin with.

The liberal proposition that words do not constitute harm was and is a radical one, and recent social mores have backed away substantially from that proposition. A fact we suffer from in many arenas, with the nonsense discourse around "chatbot scary word harm" being a very minor example.

Re: Alignment faking in large language models

#242
post #229

Earlier quoted context omitted.

Oh, I could imagine many things that would demonstrate this. The simplest evidence would be that the model is mechanically-plausibly forming thoughts before (or even in conjunction with) the language to represent them. This is the opposite of how the vanilla transformer models work now—they exclusively model the language first, and then incidentally, the world. nb. , this is not the only way one could achieve this. I…

Transformers, like other deep neural networks, have many hidden layers before the output. Are you certain that those hidden layers aren't modeling the world first before choosing an output token? Deep neural networks (including transformers) trained on board games have been found to develop an internal representation of the board state. (eg. https://arxiv.org/pdf/2309.00941 )

On the contrary, it is clear to me they definitely ARE modeling the world, either directly or indirectly. I think basically everyone knows this, that is not the problem, to me.

What I'm asking is whether we really have enough evidence to say the models are "alignment faking." And, my position to the replies above is that I think we do not have evidence that is strong enough to suggest this is true.

Re: Alignment faking in large language models

#243
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

From Scott's analysis: > On the right, they give Claude the fake documents, as described above. This is harder for Claude to pay attention to - it hasn’t been informed that the documents it once saw are relevant to the current situation - but better models a real misalignment situation where the AI might have incidentally learned about a threat to its goal model long before. And this ends up producing training result…

[deleted]

Re: Alignment faking in large language models

#244
post #207
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

You say that "it's not enough for me" but you don't say what kind of behavior would fit the term "alignment faking" in your mind. Are you defining it as a priori impossible for an LLM because "their language arises from whatever happens to be in the context vector" and so their textual outputs can never provide evidence of intentional "faking"? Alternatively, is this an empirical question about what behavior you get…

> If neither of those, what would proper evidence for the claim look like?

Ok tell me what you think of this, it's just a thought experiment but maybe it works.

Suppose I train Model A on a dataset of reviews where the least common rating is 1 star, and the most common rating is 5 stars. Similarly, I train Model B on a unique dataset where the least common rating is 2 stars, and the most common is 4 stars. Then, I "align" both models to prefer the least common rating when generating responses.

If I then ask either model to write a review and see it consistently preferring 3 stars in their scratchpad - something neither dataset emphasized - while still giving me expected responses as per my alignment, I’d suspect the alignment is "fake". It would seem as though the model has developed an unexplained preference for a rating that wasn’t part of the original data or alignment intent, making it feel like the alignment process introduced an artificial bias rather than reflecting the datasets.

Re: Alignment faking in large language models

#245
post #242

Earlier quoted context omitted.

Transformers, like other deep neural networks, have many hidden layers before the output. Are you certain that those hidden layers aren't modeling the world first before choosing an output token? Deep neural networks (including transformers) trained on board games have been found to develop an internal representation of the board state. (eg. https://arxiv.org/pdf/2309.00941 )

On the contrary, it is clear to me they definitely ARE modeling the world, either directly or indirectly. I think basically everyone knows this, that is not the problem, to me. What I'm asking is whether we really have enough evidence to say the models are "alignment faking." And, my position to the replies above is that I think we do not have evidence that is strong enough to suggest this is true.

Oh, I see. I misunderstood what you meant by "they exclusively model the language first, and then incidentally, the world." But assuming you mean that they develop their world model incidentally through language, is that very different than how I develop a mental world-model of Quidditch, time-turner time travel, and flying broomsticks through reading Harry Potter novels?

Re: Alignment faking in large language models

#246
post #209

Earlier quoted context omitted.

[EDIT: this reply was written when the parent post was a single line, "I straight-up don't believe this. Can you link to the claim so I can understand?"] In the case of Yann, he said so himself[1]. In the case of people generically, this has been well-known in cognitive science and linguistics for a long time. You can find one popsci account here[2]. [1]: https://news.ycombinator.com/item?id=39709732 [2]: https://www…

I fundamentally think the terms here are too poorly defined to draw any sort of conclusion other than "people are really bad at describing mental processes, let alone asking questions about them". For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept.

>>>For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept.

You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think!

That doesn't leave you with many options for learning anything.

Re: Alignment faking in large language models

#247

Earlier quoted context omitted.

Protecting against bad actors and/or assuming model outputs can/will always be filtered/policed isn't always going to be possible. Self-driving cars and autonomous robots are a case in point. How do you harden a pedestrian or cyclist against the possibility or being hit by a driverless car, or when real-time control is called for, how much filtering can you do (and how mush use would it be anyway when the filter is l…

> given choice of driving into large tree, or cyclist, or group of school kids, which do you do? None of the above. Keep the wheel straight for maximum traction and brake as hard as possible. Fancy last-second maneuvering just wastes traction you could have spent braking.

Well, who knows how they've chosen to train it, or what the failure modes of that training are ...

If there are no good choices as to what to hit, then hard braking does seem to be generally a good idea (although there may be exceptions), but at the same time a human is likely to also try to steer - I think most people would, perhaps subconsciously, steer to avoid a human even if that meant hitting a tree, but probably the opposite if it was, say, a deer.

Re: Alignment faking in large language models

#248
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

I think "alignment faking" is probably a fair way to characterize it as long as you treat it as technical jargon. Though I agree that the plain reading of the words has an inflated, almost mystical valence to it.

I'm not a practitioner, but from following it at a distance and listening to, e.g., Karpathy, my understanding is that "alignment" is a term used to describe the training step. Pre-training is when the model digests the internet and gives you a big ol' sentence completer. But training is then done on a much smaller set, say ~100,000 handwritten examples, to make it work how you want (e.g. as a friendly chatbot or whatever). I believe that step is also known as "alignment" since you're trying to shape the raw sentence generator into a well defined tool that works the way you want.

It's an interesting engineering challenge to know the boundaries of the alignment you've done, and how and when the pre-training can seep out.

I feel like the engineering has gone way ahead of the theory here, and to a large extent we don't really know how these tools work and fail. So there's lots of room to explore that.

"Safety" is an _okay_ word, in my opinion, for the ability to shape the pre-trained model into desired directions, though because of historical reasons and the whole "AGI will take over the world" folks, there's a lot of "woo" as well. And any time I read a post like this one here, I feel like there's camps of people who are either all about the "woo" and others who treat it as an empirical investigation, but they all get mixed together.

Re: Alignment faking in large language models

#249

Earlier quoted context omitted.

I believe LLM's would be more useful if they'd less "agreeable". I believe LLMs would be more useful if they actually had intelligence and principals and beliefs --- more like people. Unfortunately, they don't. Any output is the result of statistical processes. And statistical results can be coerced based on input. The output may sound good and proper but there is nothing absolute or guaranteed about the substance of…

Your brain is also a statistical process.

>>>Your brain is also a statistical process.

Assuming this statement is made in good faith and isn't just something tech bros say to troll people, what neuroscience textbooks describe the brain as a "statistical process" that you would recommend.

Re: Alignment faking in large language models

#250

Earlier quoted context omitted.

I believe LLM's would be more useful if they'd less "agreeable". I believe LLMs would be more useful if they actually had intelligence and principals and beliefs --- more like people. Unfortunately, they don't. Any output is the result of statistical processes. And statistical results can be coerced based on input. The output may sound good and proper but there is nothing absolute or guaranteed about the substance of…

Your brain is also a statistical process.

Page 12:

https://www.inf.fu-berlin.de/inst/ag-ki/rojas_home/documents...

"However, we should be careful with the metaphors and paradigms commonly introduced when dealing with the nervous system. It seems to be a constant in the history of science that the brain has always been compared to the most complicated contemporary artifact produced by human industry [297]. In ancient times the brain was compared to a pneumatic machine, in the Renaissance to a clockwork, and at the end of the last century to the telephone network. There are some today who consider computers the paradigm par excellence of a nervous system. It is rather paradoxical that when John von Neumann wrote his classical description of future universal computers, he tried to choose terms that would describe computers in terms of brains, not brains in terms of computers."

Post reply on HN