Live data from Hacker News

Alignment faking in large language models

anthropic.com

311–320 of 370 posts

Re: Alignment faking in large language models

#311

Earlier quoted context omitted.

> we're still just indiscriminately killing civilians just as always, so giving AI control is fine I don't even want to respond because the "we" and "as always" here is doing a lot . I don't have it in me to have an extended discussion to address how indiscriminately killing civilians was never accepted practice in modern warfare. Anyways. There are two conditions in which I see this argument(?) is useful. If you ass…

> I don't have it in me to have an extended discussion to address how indiscriminately killing civilians was never accepted practice in modern warfare. I did not claim it is/was accepted practice. I was asking if "doing it with AI is just the same so what's the big deal" was your position on the general issue (of AI making decisions in war), which I thought was a possible interpretation of your previous comment. > No…

> doing it with AI is just the same so what's the big deal

Is that implying fewer civilian deaths is NOT a big deal?

I think the parents and children of people who were killed would disagree with you.

Re: Alignment faking in large language models

#312
post #252
post #130

Earlier quoted context omitted.

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

Why not just… turn it off manually?

From my experience shutting down misbehaving laptops, there's no guarantee that will be an option.

Re: Alignment faking in large language models

#313
post #273

Earlier quoted context omitted.

> but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. Surely one of these would involve using your input from your ears and one would not? Can you not distinguish these two phenomena?

It all sounds the same in my head.

While I don't have the same experience, I regard what you say as fascinating additional information about the complexity of thought, rather than something needing to be explained away - and I suspect these differences between people will be helpful in figuring out how minds work.

It is no surprise to me that you have to adopt existing terms to talk about what it is like, as vocabularies depend on common experiences, and those of us who do not have the experience can at best only get some sort of imperfect feeling for what it is like through analogy.

Re: Alignment faking in large language models

#314
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

Agreed. Everything an LLM emits is 'faking' because, of course, it has no real values at all.

How do you define having real values? I'd imagine that a token-predicting base model might not have real values, but RLHF'd models might have them, depending on how you define the word?

Re: Alignment faking in large language models

#315

Earlier quoted context omitted.

Wow you read the whole SEP article? So cool. How do you respond to the Connectionist challenge to Fodor's core framework?

There's only one definition of thought on that page so... what are we supposed to compare and discuss?

That's not true. There's a whole section on challenges to that definition. See: 5. The Connectionist Challenge.

Re: Alignment faking in large language models

#316
post #298

Earlier quoted context omitted.

> To them, there's no debate to have over "does it have real feelings?" and not really any debate at all. To them, these are just novel stochastic tools whose core capabilities and limitations seem pretty self-apparent and can just be accommodated by choosing suitable applications, just as with all their other tools. Boom. This is my camp. Watching the Anthropic YouTube video in the linked article was pretty interest…

The problem I have with this is it seems to rapidly assume that humans and animals are somehow magic. If you think we're physical things obeying physical laws, then unless you think there's something truly uncomputable about us then we can be replicated with maths. In which case is this just an argument about scale? > It’s a pile of math working over a massive set of “differently compressed” data that encompasses a l…

All of what you said can be true but that doesn’t mean any of it applies to large language models. Large language model-like “thing” might be some component of whatever makes animals like humans conscious but it certainly isn’t the only thing. Not even close.

Re: Alignment faking in large language models

#317

Earlier quoted context omitted.

> You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! Presumably the question would be "If you claim to 'hear' your thoughts, why do you choose the word 'hear'?" It doesn't make much sense to ask people if they experience something I consider nonse…

I don't get your objection to "hear". When the sounds waves hit your ear drum it causes signals that are then sent to the brain via the auditory nerve, where they trigger neurons to fire, allowing you to perceive sound. When I have an internal monologuing I seem to be simulating the neurons that would fire if my thoughts were transmitted via sound waves through the ear. Is that not how it works for you?

> When I have an internal monologuing I seem to be simulating the neurons that would fire if my thoughts were transmitted via sound waves through the ear.

How the hell would you convince someone of this?

Re: Alignment faking in large language models

#318

Earlier quoted context omitted.

> What verb would you use for your interaction with those words? Perceive > Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it. This is just analytic language. Even if the symbol fails to materialize you can still identify what the symbol refers to via context-clues (analysis)

Perceive is a far more ambiguous term than hear, since it's not clear if your perception is subjectively visual or auditory or neither.

[deleted]

Re: Alignment faking in large language models

#319

Earlier quoted context omitted.

> What verb would you use for your interaction with those words? Perceive > Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it. This is just analytic language. Even if the symbol fails to materialize you can still identify what the symbol refers to via context-clues (analysis)

Perceive is a far more ambiguous term than hear, since it's not clear if your perception is subjectively visual or auditory or neither.

[deleted]

Re: Alignment faking in large language models

#320

Earlier quoted context omitted.

> What verb would you use for your interaction with those words? Perceive > Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it. This is just analytic language. Even if the symbol fails to materialize you can still identify what the symbol refers to via context-clues (analysis)

Perceive is a far more ambiguous term than hear, since it's not clear if your perception is subjectively visual or auditory or neither.

Sure, but hearing is non-ambiguously not applicable to anything other than objectively auditory phenomena. Unless you're psychotic.

Or maybe I'm wrong and people have just been using this term to describe mental shit all along!

That's kind of my point, though, we literally don't have the language to figure out how other people perceive things.

Perceive at least disambiguates itself from the senses that aren't related to thought!

Post reply on HN