Live data from Hacker News

Alignment faking in large language models

anthropic.com

211–220 of 370 posts

Re: Alignment faking in large language models

#211
post #207
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

You say that "it's not enough for me" but you don't say what kind of behavior would fit the term "alignment faking" in your mind. Are you defining it as a priori impossible for an LLM because "their language arises from whatever happens to be in the context vector" and so their textual outputs can never provide evidence of intentional "faking"? Alternatively, is this an empirical question about what behavior you get…

I would consider an experiment like this in conjunction with strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously, yes.

I do not think it is structurally impossible for AI generally and am excited for what happens in next-gen model architectures.

Yes, I do think the current model architectures are necessarily limited in the kinds of high-order cognitive thought they can provide, since what tokens they emit next are essentially completely beholden to the n prompt tokens, conditioned on the model itself.

Re: Alignment faking in large language models

#213
post #209

Earlier quoted context omitted.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

[EDIT: this reply was written when the parent post was a single line, "I straight-up don't believe this. Can you link to the claim so I can understand?"] In the case of Yann, he said so himself[1]. In the case of people generically, this has been well-known in cognitive science and linguistics for a long time. You can find one popsci account here[2]. [1]: https://news.ycombinator.com/item?id=39709732 [2]: https://www…

I fundamentally think the terms here are too poorly defined to draw any sort of conclusion other than "people are really bad at describing mental processes, let alone asking questions about them".

For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept.

Re: Alignment faking in large language models

#214
post #132

Earlier quoted context omitted.

Exactly. The discussion is going to change real fast when LLMs are wrapped in some sort OODA loop type thing and crammed into some sort of humanoid robot that carries hedge trimmers.

why would you want to let a LLM have any agentic interface to the real world though

And if you were building "AI" hedge trimmers, why the hell would you think that an LLM was a sensible way to engineer them?

Thinks I need my hedge trimmers to do: trim hedges, avoid trimming things that are not hedges, manoeuvre within strict boundaries.

Things I don't need my hedge trimmers to be able to do: reply to me in iambic pentameter, turn articles into bullet points, pass a FizzBuzz test

Re: Alignment faking in large language models

#215
post #150

Earlier quoted context omitted.

> On a very basic level the word salad generator part is your only part I interact with. My fingers also were involved in the typing of that message, actually they were the last proximal cause of the characters appearing the comment. Are you saying that on a very basic level my fingers are the my only part you interact with?

I'm saying I have no way of knowing you have fingers. I can see what you write, but I can't know how you did it.

Thus means all you can say about me is that I'm a black box emitting words.

But that's not what I think people mean when they say "world salad generator" or "stochastic parrot" or "broca area emulator".

The idea there is that it's indeed possible to create a machinery that is surprisingly efficient at producing natural language that sounds good and flows well, perhaps even following complex grammatical rules, and yet not being at all able to reason

Re: Alignment faking in large language models

#216

Earlier quoted context omitted.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

It's actually a very well trod field at the intersection of philosophy and cognitive science. The fundamental question is whether or not cognitive processes have the structure of language. There are compelling arguments in both directions. It's dense, but even skimming the SEP article is pretty fascinating: https://plato.stanford.edu/entries/language-thought/

> The fundamental question is whether or not cognitive processes have the structure of language.

Well that's easy—some do, some don't.

Re: Alignment faking in large language models

#217
post #211
post #207

Earlier quoted context omitted.

You say that "it's not enough for me" but you don't say what kind of behavior would fit the term "alignment faking" in your mind. Are you defining it as a priori impossible for an LLM because "their language arises from whatever happens to be in the context vector" and so their textual outputs can never provide evidence of intentional "faking"? Alternatively, is this an empirical question about what behavior you get…

I would consider an experiment like this in conjunction with strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously, yes. I do not think it is structurally impossible for AI generally and am excited for what happens in next-gen model architectures. Yes, I do think the current model architectures are necessarily limited in the kinds of h…

> strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously

What would constitute evidence of this, for you?

Re: Alignment faking in large language models

#218
post #209

Earlier quoted context omitted.

[EDIT: this reply was written when the parent post was a single line, "I straight-up don't believe this. Can you link to the claim so I can understand?"] In the case of Yann, he said so himself[1]. In the case of people generically, this has been well-known in cognitive science and linguistics for a long time. You can find one popsci account here[2]. [1]: https://news.ycombinator.com/item?id=39709732 [2]: https://www…

I fundamentally think the terms here are too poorly defined to draw any sort of conclusion other than "people are really bad at describing mental processes, let alone asking questions about them". For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept.

Huh? All my thoughts are audio and video, my thinking is literally listening to a voice in my head. It's the same way my memories are dealt with.

Re: Alignment faking in large language models

#219

Earlier quoted context omitted.

Not the OP, but my bar would be that they are built differently. It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside. That doesn’t mean LLMs won’t take some jobs. Technology has been…

> If we set that knowledge aside, we unmoor ourselves from reality The problem is that this knowledge is an a priori assumption. If we're exercising skepticism, it's important to be equally skeptical of the baseless idea that our notion of mind does not arise from a markov chain under certain conditions. You will be shocked to know that your entire physical body can be modelled as a markov chain, as all physical thin…

> You will be shocked to know that your entire physical body can be modelled as a markov chain, as all physical things can.

My favorite part of hackernews is when a bunch of tech people start pretending to know how very complex systems work despite never having studied them.

Re: Alignment faking in large language models

#220
post #206

Earlier quoted context omitted.

Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking. If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct , both for the subject and as perhaps the best piece of popular science writing ever.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

I think what you are saying is that language is deeply and perhaps inextricably tied to human thought. And, I think it's fair to say this is basically uniformly regarded as a fact.

The reason I (and others) say that language is almost certainly preceded by (and derived from) high-order thought is because high-order thought exists in all of our close relatives, while language exists only in us.

Perhaps the confusion is in the definition of high-order thought? There is an academic definition but I boil it down to "able to think about thinking as, e.g. all social great apes do when they consider social reactions to their actions."

Post reply on HN