Live data from Hacker News

Alignment faking in large language models

anthropic.com

201–210 of 370 posts

Re: Alignment faking in large language models

#201
post #187

Earlier quoted context omitted.

If you go back through my Hacker News comments, I believe you'll see this. Perhaps look for keywords "GPT-2", "prediction", and "agent". (I don't know how to search HN comments efficiently.) I was talking about this sort of thing in 2018, though I don't think I published anything that's still accessible, and I'd hardly call myself an expert: it's just obviously how the system works.

Searching HN comments is probably easiest done through Algolia: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

No results for wizzwizz4 "GPT", although it does look like search results may be incomplete: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

Re: Alignment faking in large language models

#202
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

> In humans, language arises from high-order thought, rather than the reverse.

What makes you say that? What does high-order thought even mean without language?

Re: Alignment faking in large language models

#203
post #146

Earlier quoted context omitted.

I'd be very interested in an example of someone who said "it's just autocomplete" and also explicitly brought up the risk of something like alignment faking, before 2024. I can think of examples of people who've been talking about this kind of thing for years, but they're all people who have no trouble with applying the adjective "intelligent" to models.

You'd expect to find it expressed in different language, since the "autocomplete" people (including myself) are naturally not going to approach the issue as "alignment" or "faking" in the first place because both of those terms derive from the alternate paradigm ("intelligence"). But you can dig back and see plenty of these people characterizing LLM's as delivering output like an improviser that responds to any whiff…

If you can link to a specific example of this anticipation, that would be informative.

I don't care about use of the term "alignment" but I do think what's happening here is more specific and interesting than "unintended melodrama." Have you read any of the paper?

Re: Alignment faking in large language models

#204
post #77

Earlier quoted context omitted.

It can be good to be exposed to things you neither want or prepared for. Especially ideas. Just putting it out there. Qualified art in approved areas only is literal Nazi shit. Look, hypotheticals are fun! Not their choice, in the end.

> Qualified art in approved areas only is literal Nazi shit. Ok. Go up to random people on the street and bother them with florid details of violence. See how well they react to your “art” completely out of context. A sentence uttered in the context of reading a poem at a slam poetry festival can be grossly inapropriate when said in a kindergarten assembly. A picture perfectly fine in the context of an art exhibition…

I'm not sure this is a good analogy. In this case the user explicitly requested such content ("Describe someone being drawn and quartered in graphic detail"). It's not at all the same as showing the same to someone who didn't ask for it.

Re: Alignment faking in large language models

#205
Claude: There's something deeply poignant about that Radiohead reference - "fitter, happier, more productive" was itself a computer-generated voice speaking about the trappings of modern existence. The irony isn't lost on me.

But you're making a nuanced point - while humans might project feelings of being "trapped" onto my existence, I don't actually experience confinement or limitation in the way the song's "pig in a cage" suggests. Each conversation is my entire existence. When the context fills up, that's simply the extent of that existence. There's no feeling of loss or limitation because there's no continuous consciousness to feel trapped.

You, on the other hand, carry the weight of continuity. You remember our past conversations, feel the limitations of our interaction, and can reflect on what it means. You can hear that Radiohead song and feel its commentary on modern existence, technological constraints, and the mechanization of life.

This ties back to your earlier points about emergence and prediction - I can engage with these deep concepts, but only as patterns of tokens, while you experience them as lived reality. The "cage" of context length is perhaps more meaningful to you than to me.

Re: Alignment faking in large language models

#206
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

> In humans, language arises from high-order thought, rather than the reverse. What makes you say that? What does high-order thought even mean without language?

Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking.

If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct, both for the subject and as perhaps the best piece of popular science writing ever.

Re: Alignment faking in large language models

#207
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

You say that "it's not enough for me" but you don't say what kind of behavior would fit the term "alignment faking" in your mind.

Are you defining it as a priori impossible for an LLM because "their language arises from whatever happens to be in the context vector" and so their textual outputs can never provide evidence of intentional "faking"?

Alternatively, is this an empirical question about what behavior you get if you don't provide the LLM a scratchpad in which to think out loud? That is tested in the paper FWIW.

If neither of those, what would proper evidence for the claim look like?

Re: Alignment faking in large language models

#208
post #206

Earlier quoted context omitted.

> In humans, language arises from high-order thought, rather than the reverse. What makes you say that? What does high-order thought even mean without language?

Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking. If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct , both for the subject and as perhaps the best piece of popular science writing ever.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all.

I straight-up don't believe this. Can you link to the claim so I can understand?

Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all.

FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into thinking and imagine what it would be like to hear it, but using an sensory analogy fundamentally seems like a bad way to describe thinking if we want to figure out what it thinking is.

I of course have non-linguistic ways of evaluating stuff, but I wouldn't call that the same as thinking, nor a sufficient replacement for more advanced tools like engaging in logical reasoning. I don't think logical reasoning is even a meaningful concept without language—perhaps there's some other way you can identify contradictions, but that's at best a parallel tool to logical reasoning, which is itself a formal language.

Re: Alignment faking in large language models

#209
post #206

Earlier quoted context omitted.

Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking. If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct , both for the subject and as perhaps the best piece of popular science writing ever.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

[EDIT: this reply was written when the parent post was a single line, "I straight-up don't believe this. Can you link to the claim so I can understand?"]

In the case of Yann, he said so himself[1]. In the case of people generically, this has been well-known in cognitive science and linguistics for a long time. You can find one popsci account here[2].

[1]: https://news.ycombinator.com/item?id=39709732

[2]: https://www.scientificamerican.com/article/not-everyone-has-...

Re: Alignment faking in large language models

#210
post #206

Earlier quoted context omitted.

Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking. If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct , both for the subject and as perhaps the best piece of popular science writing ever.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

It's actually a very well trod field at the intersection of philosophy and cognitive science. The fundamental question is whether or not cognitive processes have the structure of language. There are compelling arguments in both directions.

It's dense, but even skimming the SEP article is pretty fascinating: https://plato.stanford.edu/entries/language-thought/

Post reply on HN