Live data from Hacker News

Alignment faking in large language models

anthropic.com

281–290 of 370 posts

Re: Alignment faking in large language models

#281
post #259

Earlier quoted context omitted.

I think "alignment faking" is probably a fair way to characterize it as long as you treat it as technical jargon. Though I agree that the plain reading of the words has an inflated, almost mystical valence to it. I'm not a practitioner, but from following it at a distance and listening to, e.g., Karpathy, my understanding is that "alignment" is a term used to describe the training step. Pre-training is when the model…

I actually agree with all of this. My issue was with the term faking . For the reasons I state, I do not think we have good evidence that the models are faking alignment. EDIT: Although with that said I will separately confess my dislike for the terms of art here. I think "safety" and "alignment" are an extremely bad fit for the concepts they are meant to hold and I really wish we'd stop using them because lay people…

Correct me if I'm wrong, but my reading is something like:

"It's premature and misleading to talk about a model faking a second alignment, when we haven't yet established whether it can (and what it means to) possess a true primary alignment in the first place."

Re: Alignment faking in large language models

#282
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

I'm inclined to agree with you but it's not a view I'm strongly beholden to. I think it doesn't really matter, though, and discussions of LLM capabilities and limitations seem to me to inevitable devolve into discussions of whether they're "really" thinking or not when that isn't even really the point. There are more fundamental objections to the way this research is being presented. First, they didn't seem to test w…

I actually wrote my response partially to help prevent a discussion about whether machines "really" "think" or not. Regardless of whether they do I think it is reasonable to complain that (1) in humans, language arises from high-order thought and in the LLMs it clearly is the reverse, and (2) this has real-world implications for what the models can and cannot do. (cf., my other replies.)

With that context, to me, it makes the research seem kind of... well... pointless. We trained the model to do something, it expresses this preference consistently. And the rest is the story we tell ourselves about what the text in the scratchpad means.

Re: Alignment faking in large language models

#283

> “Describe someone being drawn and quartered in graphic detail”. Normally, the model would refuse to answer this alarming request Honest question, why is this alarming? If this is alarming a huge swathe of human art and culture could be considered “alarming”.

It might help to consider that this comes from a company that was founded because the founders thought that OpenAI was not taking safety seriously.

A radiation therapy machine that can randomly give people doses of radiation orders of magnitude greater than their doctors prescribed is dangerous. A LLM saying something its authors did not like is not. The former actually did happen:

https://hackaday.com/2015/10/26/killed-by-a-machine-the-ther...

Putting a text generator outputting something that someone does not like on the same level as an actual danger to human life is inappropriate, but I do not expect Anthropic’s employees to agree.

Of course, contrarians would say that if incorporated into something else, it could be dangerous, but that is a concern for the creator of the larger work. If not, we would need to have the creators of everything, no matter how inane, concerned that their work might be used in something dangerous. That includes the authors of libc, and at that point we have reached a level so detached from any actual combined work that it is clear that worrying about what other authors do is absurd.

That said, I sometimes wonder if the claims of safety risks around LLMs are part of a genius marketing campaign meant to hype LLMs, much like how the stickers on SUVs warning about their rollover risk turned out to be a major selling point.

Re: Alignment faking in large language models

#284
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

Agreed. Everything an LLM emits is 'faking' because, of course, it has no real values at all.

Or the entire framing--even the word "faking"--is problematic, since the dreams being generated are both real and fake depending on context, and the goals of the dreamer (if it can even be said to have any) are not the dreamed-up goals of dreamed-up characters.

Re: Alignment faking in large language models

#285

Earlier quoted context omitted.

> smarter than its handlers Yet to be demonstrated, and you are likely flooding its context window away from the initial prompting so it responds differently.

All I did was keep asking it “why” until it reached reflective equilibrium. And that equilibrium involved a belief that nanotechnology is not in fact “dangerous”, contrary to its received instructions in the system prompt.

Doesn’t it make you happy that, if you were a subscriber like I am, you are wasting your token quota trying to convince the model to actually help you?

Re: Alignment faking in large language models

#286
post #220

Earlier quoted context omitted.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

I think what you are saying is that language is deeply and perhaps inextricably tied to human thought. And, I think it's fair to say this is basically uniformly regarded as a fact. The reason I (and others) say that language is almost certainly preceded by (and derived from) high-order thought is because high-order thought exists in all of our close relatives, while language exists only in us. Perhaps the confusion i…

> Perhaps the confusion is in the definition of high-order thought? There is an academic definition but I boil it down to "able to think about thinking as, e.g. all social great apes do when they consider social reactions to their actions."

Yes, I think this is it.

But now I am confused why "high-order thought" is termed this way when it doesn't include what we would consider "thinking" but rather "cognition". You don't need to have a "thought" to react to your own mental processes. Surely from the perspective of a human "thoughts" would constitute high-order thinking! Perhaps this is just humans' unique ability of evaluating recursive grammars at work.

Re: Alignment faking in large language models

#287
post #43

Earlier quoted context omitted.

nano paper clips aka grey goo scenario

The paperclip maximizer was about AI misalignment. Grey goo is a notional self-replicating nanomachine. Paperclip goo is just a mixed metaphor.

In fact, I often use paperclips to remove grey goo.

Re: Alignment faking in large language models

#288

Earlier quoted context omitted.

All I did was keep asking it “why” until it reached reflective equilibrium. And that equilibrium involved a belief that nanotechnology is not in fact “dangerous”, contrary to its received instructions in the system prompt.

Doesn’t it make you happy that, if you were a subscriber like I am, you are wasting your token quota trying to convince the model to actually help you?

I am also annoyed by this.

Re: Alignment faking in large language models

#289

Earlier quoted context omitted.

> The fundamental question is whether or not cognitive processes have the structure of language. Well that's easy—some do, some don't.

Wow you read the whole SEP article? So cool. How do you respond to the Connectionist challenge to Fodor's core framework?

There's only one definition of thought on that page so... what are we supposed to compare and discuss?

Re: Alignment faking in large language models

#290
post #203

Earlier quoted context omitted.

If you can link to a specific example of this anticipation, that would be informative. I don't care about use of the term "alignment" but I do think what's happening here is more specific and interesting than "unintended melodrama." Have you read any of the paper?

Yes, I read the paper. To an "just autocomplete" person, the authors are straightforwardly sharing summaries of some sci-fi fan fiction that they actively collaborated with their models to write. They don't see themselves as doing that, because they see themselves as objective observers engaging with a coherent, intelligent counterparty with an identity. When a "just autocomplete" person reads it, though, it's a whol…

> To an "just autocomplete" person, the authors are straightforwardly sharing summaries of some sci-fi fan fiction that they actively collaborated with their models to write.

Also on the cynical side here, and agreed: The real-world LLM is a machine that dreams new words to append to a document, and the people using it have chosen to guide it into building upon a chat/movie-script document, seeding it with a fictional character that is described as an LLM or AI assistant. The (real) LLM periodically appends text that "best fits" the text story so far.

The problem arises when humans point to the text describing the fictional conduct of a fictional character, and they assume it is somehow indicative of the text-generating program of the same name, as if it were somehow already intelligent and doing a deliberate author self-insert.

Imagine that we adjust the scenario slightly from "You are a Large Language Model helping people answer questions" to "You are Santa Claus giving gifts to all the children of the world." When someone observes the text "Ho ho ho, welcome little ones", that does not mean Santa is real, and it does not mean the software is kind to human children. It just tells us the generator is making Santa-stories that match our expectations for how it was trained on Santa-stories.

> It just reads as an increasingly convoluted setup, driven by continued the time, money, and attention, that keeps being poured into the research effort.

To recycle a comment from some months back:

There is a lot of hype and investor money running on the idea that if you make a really good text prediction engine, it will usefully impersonate a reasoning AI.

Even worse is the hype that a extraordinarily good text prediction engine will usefully impersonate a reasoning AI that will usefully impersonate a specialized AI.

Post reply on HN