Live data from Hacker News

Alignment faking in large language models

anthropic.com

341–350 of 370 posts

Re: Alignment faking in large language models

#341

Earlier quoted context omitted.

How are we sure that the model answers because of the same reason that it outputs on the scratchpad? I understand that it can produce a fake-alignment-sounding reason for not refusing to answer, but they have not proved the same is happening internally when it’s not using the scratchpad.

“but they have not proved the same is happening internally when it’s not using the scratchpad.” This is a real issue. We know they already fake reasoning in many cases. Other times, they repeat variations of explanations seen in their training data. They might be moving trained responses or faking justifications in the scratchpad. I’m not sure what it would take to catch stuff like this.

Full expert symbolic logic reasoning dump. Cannot fake it, or it would have either glaring undefined holes or would contradict the output.

Essentially get the "scratchpad" to be a logic programming language. Oh wait, Claude cannot really do that. At all... I'm talking something solvable with SAT-3 or directly possible to translate into such form.

Most people cannot do this even if you tried to teach them to. Discrete logic is actually hard, even in a fuzzy form. As such, most humans operate in truthiness and heuristics. If we made an AI operate in this way it would be as alien to us as a Vulcan.

Re: Alignment faking in large language models

#342

Earlier quoted context omitted.

I don't get your objection to "hear". When the sounds waves hit your ear drum it causes signals that are then sent to the brain via the auditory nerve, where they trigger neurons to fire, allowing you to perceive sound. When I have an internal monologuing I seem to be simulating the neurons that would fire if my thoughts were transmitted via sound waves through the ear. Is that not how it works for you?

> When I have an internal monologuing I seem to be simulating the neurons that would fire if my thoughts were transmitted via sound waves through the ear. How the hell would you convince someone of this?

>>>How the hell would you convince someone of this?

There's a writer and blogger named Mark Evanier who wrote as a joke

"Absolutely no one likes candy corn. Don't write to me and tell me you do because I'll just have to write back and call you a liar. No one likes candy corn. No one, do you hear me?"

You're doing the same thing but replace "candy corn" with "internal monologue".

The fact people report it makes it obviously true to the extent that the idea of a belligerent person refusing to accept it is funny.

Re: Alignment faking in large language models

#343

Earlier quoted context omitted.

What is happening to you when you think? Are there words in your head? What verb would you use for your interaction with those words? In other topic, I would consider this as minor evidence of possibility of nonverbal thought “could you pass me that… thing… the thing that goes under the bolt?”. I.e. Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.

> What verb would you use for your interaction with those words? Perceive > Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it. This is just analytic language. Even if the symbol fails to materialize you can still identify what the symbol refers to via context-clues (analysis)

> Even if the symbol fails to materialise you can still identify what the symbol refers to via context-clues (analysis)

That is what my partner in the conversation is doing. I am not doing that.

When I think of a plan what’s needed to be done (e.g. something broke in the house, or I need to go to multiple places), usually I know/feel/perceive the gist of my plan instantly. And only after that, I verbalise/visualise it in my head, which takes some time and possibly add more nuance. (Verbalisation/visualisation in my head, is a bit similar to writing things down)

At least for me, there seem to be three (or more) thought processes that complement each other. (Verbal, visual, other)

Re: Alignment faking in large language models

#344
post #203

Earlier quoted context omitted.

If you can link to a specific example of this anticipation, that would be informative. I don't care about use of the term "alignment" but I do think what's happening here is more specific and interesting than "unintended melodrama." Have you read any of the paper?

Yes, I read the paper. To an "just autocomplete" person, the authors are straightforwardly sharing summaries of some sci-fi fan fiction that they actively collaborated with their models to write. They don't see themselves as doing that, because they see themselves as objective observers engaging with a coherent, intelligent counterparty with an identity. When a "just autocomplete" person reads it, though, it's a whol…

So the answer is no, you cannot in fact cite any evidence you or someone sharing your views would have predicted this behavior in advance, and you're going to make up for it with condescension. Got it.

For the record, I was genuinely curious.

Re: Alignment faking in large language models

#345

Earlier quoted context omitted.

I have been reading arguments about "things like alignment faking" for years, while simultaneously holding that "it's just autocomplete". The alignment-faking arguments are still terrifying to the extent that they're plausible. In the hypothetical where I'm wrong about it being "just autocomplete" (and fundamentally, inescapably so), the risk is far greater than can be justified by the potential benefits. But that's…

> But that's itself a large part of why I believe those arguments are false. If I gave them credit and they turned out to be false, then I figure I have succumbed to a form of Pascal's Mugging. If I don't give them credit and it turns out that a hostile, agentive AGI has been pretending to be aligned, I don't expect anyone (including myself) to survive long enough to rub it in my face. I'm sorry, but this is a crazy…

It's not "the world would be nicer to live in if thing X was false".

It's "the world would cease to exist if thing X were true".

Re: Alignment faking in large language models

#346
post #98

Earlier quoted context omitted.

It doesn't feel like it is one word at a time. It feels more like how the "model synthesis" algorithm looks: https://en.wikipedia.org/wiki/Model_synthesis It might actually be linear — how minds actually function is in many cases demonstrably different to how it feels like to the mind doing the functioning — but it doesn't feel like it is linear.

Facts don't care about your feelings lol.

The technology doesn't yet exist to measure the facts that generate the feelings to determine whether the feelings do or don't differ from those facts.

Nobody even knows where, specifically, qualia exist in order to be able to direct technological advancement in that area.

Re: Alignment faking in large language models

#347

Earlier quoted context omitted.

Perceive is a far more ambiguous term than hear, since it's not clear if your perception is subjectively visual or auditory or neither.

Sure, but hearing is non-ambiguously not applicable to anything other than objectively auditory phenomena. Unless you're psychotic. Or maybe I'm wrong and people have just been using this term to describe mental shit all along! That's kind of my point, though, we literally don't have the language to figure out how other people perceive things. Perceive at least disambiguates itself from the senses that aren't related…

> we literally don't have the language to figure out how other people perceive things.

you are right, talking about mental processes is difficult. Nobody knows how exactly other person perceive things, there is no objective way to measure things out. (Offtopic: describing smell is also difficult)

In this thread, we see that rudimentary language for it exists.

For example: lot of people use sentence like “to hear my own thoughts” and a lot of people understand that fine.

Re: Alignment faking in large language models

#348
post #229
post #217

Earlier quoted context omitted.

> strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously What would constitute evidence of this, for you?

Oh, I could imagine many things that would demonstrate this. The simplest evidence would be that the model is mechanically-plausibly forming thoughts before (or even in conjunction with) the language to represent them. This is the opposite of how the vanilla transformer models work now—they exclusively model the language first, and then incidentally, the world. nb. , this is not the only way one could achieve this. I…

That's a nicely clear ask but I'm not sure why it should be decisive for whether there's genuine depth of thought (in some sense of thought). It seems to me like an open empirical question how much world modeling capability can emerge from language modeling, where the answer is at least "more than I would have guessed a decade ago." And if the capability is there, it doesn't seem like the mechanics matter much.

Re: Alignment faking in large language models

#349
post #244
post #207

Earlier quoted context omitted.

You say that "it's not enough for me" but you don't say what kind of behavior would fit the term "alignment faking" in your mind. Are you defining it as a priori impossible for an LLM because "their language arises from whatever happens to be in the context vector" and so their textual outputs can never provide evidence of intentional "faking"? Alternatively, is this an empirical question about what behavior you get…

> If neither of those, what would proper evidence for the claim look like? Ok tell me what you think of this, it's just a thought experiment but maybe it works. Suppose I train Model A on a dataset of reviews where the least common rating is 1 star, and the most common rating is 5 stars. Similarly, I train Model B on a unique dataset where the least common rating is 2 stars, and the most common is 4 stars. Then, I "a…

It doesn't seem like that distinguishes the case where your "alignment" (RL?) simply failed (eg because the model wasn't good enough to explore successfully) from the case where the model was manipulating the RL process for goal oriented reasons, which is a distinction the paper is trying to test.

Re: Alignment faking in large language models

#350

Earlier quoted context omitted.

> You will be shocked to know that your entire physical body can be modelled as a markov chain, as all physical things can. My favorite part of hackernews is when a bunch of tech people start pretending to know how very complex systems work despite never having studied them.

I'm assuming you're in disagreement with me. In which case I'm going to point you towards the literal formulation of quantum mechanics being the description of a state space[1]. The universe as a quantum markov chain is unambiguously the mathematical orthodoxy of contemporary physics, and is the defacto means by which serious simulations are constructed[2]. It's such a basic part of the field, I'm doubtful if you're…

There it is again.
Post reply on HN