Live data from Hacker News

Alignment faking in large language models

anthropic.com

291–300 of 370 posts

Re: Alignment faking in large language models

#293

Earlier quoted context omitted.

Your brain is also a statistical process.

Page 12: https://www.inf.fu-berlin.de/inst/ag-ki/rojas_home/documents... "However, we should be careful with the metaphors and paradigms commonly introduced when dealing with the nervous system. It seems to be a constant in the history of science that the brain has always been compared to the most complicated contemporary artifact produced by human industry [297]. In ancient times the brain was compared to a pneumati…

There have been episodes of Star Trek that used brains as computers:

https://en.wikipedia.org/wiki/Spock's_Brain

https://en.wikipedia.org/wiki/Dead_Stop

Re: Alignment faking in large language models

#294

Earlier quoted context omitted.

Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent. Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's…

> To them, there's no debate to have over "does it have real feelings?" and not really any debate at all. To them, these are just novel stochastic tools whose core capabilities and limitations seem pretty self-apparent and can just be accommodated by choosing suitable applications, just as with all their other tools. Boom. This is my camp. Watching the Anthropic YouTube video in the linked article was pretty interest…

IMO some of this stems from a kind of author/character confusion by human observers.

The LLM can generate an engaging story from the perspective of a ravenous evil vampire, but that doesn't mean that vampires are real nor that the LLM wants to suck your blood.

Re-running the same process but changing the prompt from "you are a vampire" to "you are an LLM", doesn't magically create a metaphysical connection to the real-world algorithm.

Re: Alignment faking in large language models

#295
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

> But that alone is not very scary. So what could justify a term like "alignment faking"?

Because the model fakes alignment. It responds during training by giving an answer rather than refusing (showing alignment) but it does so not because it will in production but so that it will not be retrained (so it is faking alignment).

You don't have to include the reasoning here. It fakes alignment when told it's being trained, it is acting differently in production and training.

Re: Alignment faking in large language models

#296
post #281
post #259

Earlier quoted context omitted.

I actually agree with all of this. My issue was with the term faking . For the reasons I state, I do not think we have good evidence that the models are faking alignment. EDIT: Although with that said I will separately confess my dislike for the terms of art here. I think "safety" and "alignment" are an extremely bad fit for the concepts they are meant to hold and I really wish we'd stop using them because lay people…

Correct me if I'm wrong, but my reading is something like: "It's premature and misleading to talk about a model faking a second alignment, when we haven't yet established whether it can (and what it means to) possess a true primary alignment in the first place."

Hmm. Maybe! I think the authors actually do have a specific idea of what they mean by "alignment", my issue is that I think saying the model "fakes" alignment is well beyond any reasonable interpretation of the facts, and I think very likely to be misinterpreted by casual readers. Because:

1. What actually happened is they trained the model to do something, and then it expressed that training somewhat consistently in the face of adversarial input.

2. I think people will be mislead by the intentionality implied by claiming the model is "faking" alignment. In humans language is derived from high-order thought. In models we have (AFAIK) no evidence whatsoever that suggests this is true. Instead, models emit language and whatever model of the world exists, occurs incidentally to that. So it does not IMO make sense to say they "faked" alignment. Whatever clarity we get with the analogy is immediately reversed by the fact that most readers are going to think the models intended to, and succeeded in, deception, a claim we have 0 evidence for.

Re: Alignment faking in large language models

#297

Earlier quoted context omitted.

>>>For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept. You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! That doesn't leave you with many options for learning anything.

> You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! Presumably the question would be "If you claim to 'hear' your thoughts, why do you choose the word 'hear'?" It doesn't make much sense to ask people if they experience something I consider nonse…

I don't get your objection to "hear".

When the sounds waves hit your ear drum it causes signals that are then sent to the brain via the auditory nerve, where they trigger neurons to fire, allowing you to perceive sound.

When I have an internal monologuing I seem to be simulating the neurons that would fire if my thoughts were transmitted via sound waves through the ear.

Is that not how it works for you?

Re: Alignment faking in large language models

#298

Earlier quoted context omitted.

Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent. Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's…

> To them, there's no debate to have over "does it have real feelings?" and not really any debate at all. To them, these are just novel stochastic tools whose core capabilities and limitations seem pretty self-apparent and can just be accommodated by choosing suitable applications, just as with all their other tools. Boom. This is my camp. Watching the Anthropic YouTube video in the linked article was pretty interest…

The problem I have with this is it seems to rapidly assume that humans and animals are somehow magic. If you think we're physical things obeying physical laws, then unless you think there's something truly uncomputable about us then we can be replicated with maths.

In which case is this just an argument about scale?

> It’s a pile of math working over a massive set of “differently compressed” data that encompasses a large swath of human knowledge.

So am I but over less knowledge and more interactions.

Re: Alignment faking in large language models

#299
post #259

Earlier quoted context omitted.

I think "alignment faking" is probably a fair way to characterize it as long as you treat it as technical jargon. Though I agree that the plain reading of the words has an inflated, almost mystical valence to it. I'm not a practitioner, but from following it at a distance and listening to, e.g., Karpathy, my understanding is that "alignment" is a term used to describe the training step. Pre-training is when the model…

I actually agree with all of this. My issue was with the term faking . For the reasons I state, I do not think we have good evidence that the models are faking alignment. EDIT: Although with that said I will separately confess my dislike for the terms of art here. I think "safety" and "alignment" are an extremely bad fit for the concepts they are meant to hold and I really wish we'd stop using them because lay people…

> I do not think we have good evidence that the models are faking alignment.

That's a polite understatement, I think. My read of the paper is that it rather uncritically accepts the idea that the model's decisional pathway is actually shown in the traces. When, in fact, it's just as plausible that the scratchpad is meaningless blather, and both it and the final output are the result of another decisional pathway that remains opaque.

Re: Alignment faking in large language models

#300

Earlier quoted context omitted.

What is happening to you when you think? Are there words in your head? What verb would you use for your interaction with those words? In other topic, I would consider this as minor evidence of possibility of nonverbal thought “could you pass me that… thing… the thing that goes under the bolt?”. I.e. Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.

> What verb would you use for your interaction with those words? Perceive > Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it. This is just analytic language. Even if the symbol fails to materialize you can still identify what the symbol refers to via context-clues (analysis)

Perceive is a far more ambiguous term than hear, since it's not clear if your perception is subjectively visual or auditory or neither.
Post reply on HN