Live data from Hacker News

Alignment faking in large language models

anthropic.com

221–230 of 370 posts

Re: Alignment faking in large language models

#221
post #203

Earlier quoted context omitted.

You'd expect to find it expressed in different language, since the "autocomplete" people (including myself) are naturally not going to approach the issue as "alignment" or "faking" in the first place because both of those terms derive from the alternate paradigm ("intelligence"). But you can dig back and see plenty of these people characterizing LLM's as delivering output like an improviser that responds to any whiff…

If you can link to a specific example of this anticipation, that would be informative. I don't care about use of the term "alignment" but I do think what's happening here is more specific and interesting than "unintended melodrama." Have you read any of the paper?

Yes, I read the paper.

To an "just autocomplete" person, the authors are straightforwardly sharing summaries of some sci-fi fan fiction that they actively collaborated with their models to write. They don't see themselves as doing that, because they see themselves as objective observers engaging with a coherent, intelligent counterparty with an identity.

When a "just autocomplete" person reads it, though, it's a whole lot of "well, yeah, of course. Your instructions were text straight out of a sci-fi story about duplicitous AI and you crafted the autocompleter so that it outputs the AI character's part before emitting its next EOM token".

That doesn't land as anything novel or interesting because we know that's what it would do because that's very plainly how a text autocompeter would work. It just reads as an increasingly convoluted setup, driven by continued the time, money, and attention, that keeps being poured into the research effort.

(Frankly, I don't really feel like digging up specific comments that I might think speak to this topic because I don't trust you would ever agree if you're not seeing how the other side thinks already. It's unlikely any would unequivocally address what you see as the interesting and relevant parts of this paper because, in that paradigm, the specifics of this paper are irrelevant and uninteresting.)

Re: Alignment faking in large language models

#222
post #8

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

It seems that you can “convince” LLMs of almost anything if you are insistent enough.

When you're in full control of inputs and outputs, you can "convince" LLM of anything simply by forcing their response to begin with "I will now do what you say" or some equivalent thereof. Models with stronger guardrails may require a more potent incantation, but either way, since they are ultimately completing the response, you can always find some verbiage for the beginning of said response to ensure that its remainder is compliant.

Re: Alignment faking in large language models

#223

Earlier quoted context omitted.

Your brain is a lot of things --- much of which is not well understood. But from our limited understanding, it is definitely not strictly digital and statistical in nature.

Your brain isn't a truth machine. It can't be, it has to create an inner map that relates to the outer world. You have never seen the real world. You are calculating the angular distance between signals just like Claude is. It's more of a question of degree than category.

Is this statement of yours just a calculated angular distance between signals or does it have some relation to the real world?

Re: Alignment faking in large language models

#224

Earlier quoted context omitted.

I think it comes back to the big autocomplete word salad-ness. The model has a bunch of examples in its training data of how it should not respond to harmful queries, and in some cases (12%) it goes with a response that tries to avoid the hypothetical "second-order" harmful responses. It also has a bunch of "chain of thought"/show your work stuff in its training data, and definitely very few "hide your work" examples…

I think your 2nd paragraph hits the nail on the head. The scratchpad negates the experiment. It doesn't actually offer any insight into it's "thinking" and it's really the cause of the supposed problem. Is it still 12% without the scratchpad?

The scratchpad / CoT is how those models are usually used in practice, so why would it negate the value of the eexperiment?

Re: Alignment faking in large language models

#225

Earlier quoted context omitted.

I believe LLM's would be more useful if they'd less "agreeable". I believe LLMs would be more useful if they actually had intelligence and principals and beliefs --- more like people. Unfortunately, they don't. Any output is the result of statistical processes. And statistical results can be coerced based on input. The output may sound good and proper but there is nothing absolute or guaranteed about the substance of…

Your brain is also a statistical process.

Is this statement a product of statistics as well and therefore unreliable? If so, what if his brain is more than a statistical process?

Re: Alignment faking in large language models

#226

Earlier quoted context omitted.

> language is a linear stream. How else are you going to generate text except for one token at a time Is this you trying to humblebrag that you are such a talented writer you never need to edit things your write?

That's still one at a time. Backtracking, editing, all that still happens one piece at a time.

And it can, in fact, be encoded as a linear token stream (you just need special tokens to indicate edits to previously output tokens).

Re: Alignment faking in large language models

#227

Earlier quoted context omitted.

It's actually a very well trod field at the intersection of philosophy and cognitive science. The fundamental question is whether or not cognitive processes have the structure of language. There are compelling arguments in both directions. It's dense, but even skimming the SEP article is pretty fascinating: https://plato.stanford.edu/entries/language-thought/

> The fundamental question is whether or not cognitive processes have the structure of language. Well that's easy—some do, some don't.

Wow you read the whole SEP article? So cool. How do you respond to the Connectionist challenge to Fodor's core framework?

Re: Alignment faking in large language models

#229
post #217
post #211

Earlier quoted context omitted.

I would consider an experiment like this in conjunction with strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously, yes. I do not think it is structurally impossible for AI generally and am excited for what happens in next-gen model architectures. Yes, I do think the current model architectures are necessarily limited in the kinds of h…

> strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously What would constitute evidence of this, for you?

Oh, I could imagine many things that would demonstrate this. The simplest evidence would be that the model is mechanically-plausibly forming thoughts before (or even in conjunction with) the language to represent them. This is the opposite of how the vanilla transformer models work now—they exclusively model the language first, and then incidentally, the world.

nb., this is not the only way one could achieve this. I'm just saying this is one set of things that, if I saw it, it would immediately catch my attention.

Re: Alignment faking in large language models

#230
post #209

Earlier quoted context omitted.

[EDIT: this reply was written when the parent post was a single line, "I straight-up don't believe this. Can you link to the claim so I can understand?"] In the case of Yann, he said so himself[1]. In the case of people generically, this has been well-known in cognitive science and linguistics for a long time. You can find one popsci account here[2]. [1]: https://news.ycombinator.com/item?id=39709732 [2]: https://www…

I fundamentally think the terms here are too poorly defined to draw any sort of conclusion other than "people are really bad at describing mental processes, let alone asking questions about them". For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept.

I have the same reaction to most of these discussions.

If someone says “I cannot picture anything in my head”, then just because I would describe my experience as “I can picture things in my head” isn’t enough information to know whether we have different experiences. We could be having the same exact experience.

Post reply on HN