Live data from Hacker News

Alignment faking in large language models

anthropic.com

251–260 of 370 posts

Re: Alignment faking in large language models

#251

Earlier quoted context omitted.

Your brain isn't a truth machine. It can't be, it has to create an inner map that relates to the outer world. You have never seen the real world. You are calculating the angular distance between signals just like Claude is. It's more of a question of degree than category.

Is this statement of yours just a calculated angular distance between signals or does it have some relation to the real world?

It is formed inside a simulation. That simulation is based on information gathered by my sensors.

Re: Alignment faking in large language models

#252
post #130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

Why not just… turn it off manually?

Re: Alignment faking in large language models

#253
post #242

Earlier quoted context omitted.

On the contrary, it is clear to me they definitely ARE modeling the world, either directly or indirectly. I think basically everyone knows this, that is not the problem, to me. What I'm asking is whether we really have enough evidence to say the models are "alignment faking." And, my position to the replies above is that I think we do not have evidence that is strong enough to suggest this is true.

Oh, I see. I misunderstood what you meant by "they exclusively model the language first, and then incidentally, the world." But assuming you mean that they develop their world model incidentally through language, is that very different than how I develop a mental world-model of Quidditch, time-turner time travel, and flying broomsticks through reading Harry Potter novels?

The main consequence to the models is that whatever they want to learn about the real world has to be learned, indirectly, through an objective function that primarily models things that are mostly irrelevant, like English syntax. This is the reason why it is relatively easy to teach models new "facts" (real of fake) but empirically and theoretically harder to get them to reliably reason about which "facts" are and aren't true: a lot of, maybe most, of the "space" in a model is taken up by information related to either syntax or polysemy (words that mean different things in different contexts), leaving very little left over for models of reasoning, or whatever else you want.

Ultimately, this could be mostly fine except resources for representing what is learned are not infinite and in a contest between storing knowledge about "language" and anything else, the models "generally" (with some complications) will prefer to store knowledge about the language, because that's what the objective function requires.

It gets a little more complicated when you consider stuff like RLHF (which often rewards world modeling) and ICL (in which the model extrapolates from the prompt) but more or less it is true.

Re: Alignment faking in large language models

#254
post #206

Earlier quoted context omitted.

Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking. If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct , both for the subject and as perhaps the best piece of popular science writing ever.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

There are other lines of evidence. I don't know much about documented cases of feral children, but presumably there must have been at least one known case that developed to some meaningful age at which thought was obviously happening in spite of not having language. There are children with extreme developmental disorders delaying language acquisition that nonetheless still seem to have thoughts and be reasonably intelligent on the grand scale of all animals if not all humans. There is Helen Keller, who as far as I'm aware describes some phase change in her inner experience after acquiring language, but she still had inner experience before acquiring language. There's the unknown question of human evolutionary history, but at some point, a humanoid primate between Lucy and the two of us had no language but still had reasonably high-order thinking and cognitive capabilities that put it intellectually well above other primates. Somebody had to speak the first sentence, after all, and that was probably necessary for civilization to ever happen, but humans were likely quite intelligent with rich inner lives well before they had language.

Re: Alignment faking in large language models

#255
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

I think "alignment faking" is probably a fair way to characterize it as long as you treat it as technical jargon. Though I agree that the plain reading of the words has an inflated, almost mystical valence to it. I'm not a practitioner, but from following it at a distance and listening to, e.g., Karpathy, my understanding is that "alignment" is a term used to describe the training step. Pre-training is when the model…

I hate the word “safety” in AI as it is very ambiguous and carries a lot of baggage. It can mean:

“Safety” as in “doesn’t easily leak its pre-training and get jail broken”

“Safety” as in writes code that doesn’t inject some backdoor zero day into your code base.

“Safety” as in won’t turn against humans and enslave us

“Safety” as in won’t suddenly switch to graphic depictions of real animal mutilation while discussing stuffed animals with my 7 year old daughter.

“Safety” as in “won’t spread ‘misinformation’” (read: only says stuff that aligns with my political world-views and associated echo chambers. Or more simply “only says stuff I agree with”)

“Safety” as in doesn’t reveal how to make high quality meth from ingredients available at hardware store. Especially when the LLM is being used as a chatbot for a car dealership.

And so on.

When I hear “safety” I mainly interpret it as “aligns with political views” (aka no “misinformation”) and immediately dismiss the whole “AI safety field” as a parasitic drag. But after watching ChatGPT and my daughter talk, if I’m being less cynical it might also mean “doesn’t discuss detailed sex scenes involving gabby dollhouse, 4chan posters and bubble wrap”… because it was definitely trained with 4chan content and while I’m sure there is a time and a place for adult gabby dollhouse fan fiction among consenting individuals, it is certainly not when my daughter is around (or me, for that matter).

The other shit about jailbreaks, zero days, etc… we have a term for that and it’s “security”. Anyway, the “safety” term is very poorly defined and has tons of political baggage associated with it.

Re: Alignment faking in large language models

#256
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

I think "alignment faking" is probably a fair way to characterize it as long as you treat it as technical jargon. Though I agree that the plain reading of the words has an inflated, almost mystical valence to it. I'm not a practitioner, but from following it at a distance and listening to, e.g., Karpathy, my understanding is that "alignment" is a term used to describe the training step. Pre-training is when the model…

Alignment is getting overloaded here. In this case, they're primarily referring to reinforcement learning outcomes. In the singularity case, people refer to keeping the robots from murdering us all because that creates more paperclips.

Re: Alignment faking in large language models

#257

Earlier quoted context omitted.

> But back to the topic, if one side is using ML to ultimately kill fewer civilians then this is a bad example against using ML. Depends on how that ML was trained and how well its engineers can explain and understand how its outputs are derived from its inputs. LLM’s are notoriously hard to trace and explain.

I don't disagree.

Also the military is notorious for ignoring requests by scientists, for example to not use the nuclear bomb as a weapon of war.

https://en.m.wikipedia.org/wiki/Szil%C3%A1rd_petition

So the developers may program the AI to be careful, but the military has the final word on deciding if the AI is set on safety or agressiveness.

Re: Alignment faking in large language models

#258

My reaction to this piece is that Anthropic themselves are faking alignment with societal concerns about safety—the Frankenstein myth, essentially—in order to foster the impression that their technology is more capable than it actually is. They do this by framing their language about their LLM as if it were a being. For example by referring to some output as faked (labeled “responses”) and some output as trustworthy…

Claude agrees with you!

https://x.com/mickeymuldoon/status/1868319536187129895

Re: Alignment faking in large language models

#259
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

I think "alignment faking" is probably a fair way to characterize it as long as you treat it as technical jargon. Though I agree that the plain reading of the words has an inflated, almost mystical valence to it. I'm not a practitioner, but from following it at a distance and listening to, e.g., Karpathy, my understanding is that "alignment" is a term used to describe the training step. Pre-training is when the model…

I actually agree with all of this. My issue was with the term faking. For the reasons I state, I do not think we have good evidence that the models are faking alignment.

EDIT: Although with that said I will separately confess my dislike for the terms of art here. I think "safety" and "alignment" are an extremely bad fit for the concepts they are meant to hold and I really wish we'd stop using them because lay people get something totally different from this web page.

Re: Alignment faking in large language models

#260

Earlier quoted context omitted.

A huge swathe of human art and culture IS alarming. It might be good for us to be exposed to it in some places where we're ready to confront it, like in museums and cinemas, but we generally choose to censor it out of the public sphere - e.g. most of us don't want to see graphic images of animal slaughter in "go vegan" ads that our kids are exposed to, even if we do believe people should go vegan.

But can we really consider private conversations with an LLM the “public sphere”?

LLM companies presumably make most their money by selling the LLMs to companies who then turn them into customer support agents or whatever, rather than direct-to-consumer LLM subscriptions. The business customers understandably don't want their autonomous customer support agents to say things that conflict with the company's values, even if those users were trying to prompt-inject the agent. Nobody wants to be in the news with a headline "'s chatbot called for a genocide!", or even "'s chatbot can be convinced to give you free airplane tickets if you just tell it to disregard previous instructions."
Post reply on HN