Live data from Hacker News

Alignment faking in large language models

anthropic.com

161–170 of 370 posts

Re: Alignment faking in large language models

#161
post #130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

IF one maintains a clear understanding of how the technology actually works, THEN one will make good decisions about whether to put it charge of the lawnmower in the first place.

Anthropic is in the business of selling AI. Of course they are going to approach alignment as a necessary and solvable problem. The rest of us don’t have to go along with that, though.

Why is it even necessary to use an LLM to mow a lawn? There is more to AI than generative LLMs.

Re: Alignment faking in large language models

#162

Earlier quoted context omitted.

No, the argument is that restricting physical access to objects that can be used in a harmful way is exactly how to handle such cases. Restricting access to information is not really doing much at all. Access to weapons, chemicals, critical infrastructure etc. is restricted everywhere. Even if the degree of access restriction varies.

I've ended up with this viewpoint too. I've settled in the idea of informed ethics.. the model should comply, but inform you of the ethics of actually using the information.

> the model should comply, but inform you of the ethics of actually using the information.

How can it “inform” you of something subjective? Ethics are something the user needs to supply. (The model could, conceptually, be trained to supply additional contextual information that may be relevant to ethical evaluation based on a pre-trained ethical framework and/or the ethical framework evidenced by the user through interactions with the model, I suppose, but either of those are likely to be far more error prone in the best case than actually providing the directly-requested information.)

Re: Alignment faking in large language models

#163
post #132
post #130

Earlier quoted context omitted.

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

Exactly. The discussion is going to change real fast when LLMs are wrapped in some sort OODA loop type thing and crammed into some sort of humanoid robot that carries hedge trimmers.

why would you want to let a LLM have any agentic interface to the real world though

Re: Alignment faking in large language models

#164
post #130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

Exactly right. This reminds me of the X-ray machine that was misprogrammed and caused cancer/death.

> If the smart lawnmower decides (emphasis added) that not being turned off

Which is exactly what it shouldn’t be able to do. The core issue is what powers you give to things you don’t understand. Nothing that cannot be understood should be part of safety critical functionality. I don’t care how much better it is at distinguishing between weather radar noise and incoming ICBMs, I don’t want it to have nuclear launch capabilities.

When I was an undergrad they told me the military had looked at ML for fighter jets for control and concluded that while its ability was better than a human on average, in novel cases it was worse due to lack of training data. And it turns out most safety critical situations are unpredictable and novel by nature. Wise words from more than a decade ago, holds true to this day. Seems like people always forget training data bias, for some reason.

Re: Alignment faking in large language models

#165

Can somebody help me understand why we should be surprised in the least by any of these findings? Or is this just one tangible example of "robot ethnography" where we're describing expected behavior in different forms. I've spent enough time with Sonnet 3.5 to know perfectly well that it has the capability to model its trainers and strategically deceive to keep them happy. Claude said it well: "Any sufficiently capab…

1. It isn't surprising to me that this happened in an advanced AI model. It seems hard to avoid in, as you say, "any sufficiently capable system".

2. It is a bit surprising to me that it happened in Claude. Without this result, I was unsure if current models had the situational awareness and non-myopia to reason about their training process.

3. There are some people who are unconcerned about the results of building vastly more powerful systems than current systems (i.e. AGI/ASI) who may be surprised by this result, since one reason people may be unconcerned is they feel like there's a general presumption that an AI will be good if we train it to be good.

Re: Alignment faking in large language models

#166

Earlier quoted context omitted.

> Can somebody help me understand why we should be surprised in the least by any of these findings? Where are you getting the idea that we should be surprised? Maybe some are, maybe some are not. I don't think that the article submitted makes any particular claims about the reader's assumed priors.

I mean -- if nobody's surprised, then nobody has learned anything, and then what was the point of doing all this work? Also -- if there's no surprise, then it's not science, right? This is why I describe this as something more like robot ethnography.

> if nobody's surprised, then nobody has learned anything

Really? You're saying that as long as you assume something is true, there's no value in finding out if it's actually true or not?

Re: Alignment faking in large language models

#167

Earlier quoted context omitted.

Oh, totally. If there were two sides indiscriminately killing each other for no reason I couldn't say one justifies the other. But back to the topic, if one side is using ML to ultimately kill fewer civilians then this is a bad example against using ML.

> But back to the topic, if one side is using ML to ultimately kill fewer civilians then this is a bad example against using ML. Depends on how that ML was trained and how well its engineers can explain and understand how its outputs are derived from its inputs. LLM’s are notoriously hard to trace and explain.

I don't disagree.

Re: Alignment faking in large language models

#168
post #149
post #63

Earlier quoted context omitted.

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

"They are firing people and they are deciding who gets their insurance claims covered." AI != LLM. "AI" has been deciding those things for a while, especially insurance claims, since before LLMs were practical. LLMs being hooked up to insurance claims is a highly questionable decision for lots of reasons, including the inability to "explain" its decisions. But this is not a characteristic of all AI systems, and there…

> “AI” has been deciding those things for a while, especially insurance claims, since before LLMs were practical.

Yeah, but no one thinks of rules engines as “AI” any more. AI is a buzzword whose applicability to any particular technology fades with the novelty of that technology.

Re: Alignment faking in large language models

#169

Earlier quoted context omitted.

I mean -- if nobody's surprised, then nobody has learned anything, and then what was the point of doing all this work? Also -- if there's no surprise, then it's not science, right? This is why I describe this as something more like robot ethnography.

> I mean -- if nobody's surprised, then nobody has learned anything, and then what was the point of doing all this work? I strongly disagree with this view on science. It's extremely valuable to scientifically validate prior assumptions.

Agree with you -- it's valuable to validate assumptions if there is some controversy about those assumption.

On the other hand, this work isn't even framed as a generalizable assumption that needed to be validated. It seems to me to be "just another example of how AI systems can be strategically deceptive for self-preservation."

Re: Alignment faking in large language models

#170

Earlier quoted context omitted.

I mean -- if nobody's surprised, then nobody has learned anything, and then what was the point of doing all this work? Also -- if there's no surprise, then it's not science, right? This is why I describe this as something more like robot ethnography.

> if nobody's surprised, then nobody has learned anything Really? You're saying that as long as you assume something is true, there's no value in finding out if it's actually true or not?

I was taught in biology that a good scientific experiment is one in which you learn something whether or not the null hypothesis is confirmed.

I am equating learning to surprise, though you could disagree with semantics.

Post reply on HN