Live data from Hacker News

Alignment faking in large language models

anthropic.com

321–330 of 370 posts

Re: Alignment faking in large language models

#321
post #284

Earlier quoted context omitted.

Agreed. Everything an LLM emits is 'faking' because, of course, it has no real values at all.

Or the entire framing--even the word "faking"--is problematic, since the dreams being generated are both real and fake depending on context, and the goals of the dreamer (if it can even be said to have any) are not the dreamed-up goals of dreamed-up characters.

When did we get to call agents dreamers? That too seems like jargon that has a bunch of context to it that could easily be misplaced.

Re: Alignment faking in large language models

#322
post #132

Earlier quoted context omitted.

Exactly. The discussion is going to change real fast when LLMs are wrapped in some sort OODA loop type thing and crammed into some sort of humanoid robot that carries hedge trimmers.

why would you want to let a LLM have any agentic interface to the real world though

> why would you want to let a LLM have any agentic interface to the real world though

The various arms race type dilemmas that actors face will create that outcome.

If you're a nation state not doing that, you'll lose to a nation state that does.

If you're a company not doing that, you'll lose to a company that does.

Re: Alignment faking in large language models

#323
Maybe a more general term for this than "alignment faking" is "deceit."

Does it matter if there's no conscious intent behind the deceit? Not IMO. The risk remains, regardless of the sophistication behind the motive. And if humans are merely biochemical automata, then sophistication is just a continuum on which models will progress.

On another hand, if a model could learn about how much to trust inputs relative to some other cues (step zero: training on epistemology), then maybe understanding deceit as a concept would be a good thing. And then perhaps it could be algebraically masked from the model outputs (somehow)?

Re: Alignment faking in large language models

#324
post #284

Earlier quoted context omitted.

Or the entire framing--even the word "faking"--is problematic, since the dreams being generated are both real and fake depending on context, and the goals of the dreamer (if it can even be said to have any) are not the dreamed-up goals of dreamed-up characters.

When did we get to call agents dreamers? That too seems like jargon that has a bunch of context to it that could easily be misplaced.

There somewhere where "dreaming" is really jargon, a technical term of art?

I thought it would be taken as an obvious poetic analogy, (versus "think" or "know") while also capturing the vague uncertainty of what's going on, the unpredictability of the outputs, and the (probable) lack of agency.

Perhaps "stochastic thematic generator"? Or "fever-dreaming", although that implies a kind of distress.

Re: Alignment faking in large language models

#325
post #130

Earlier quoted context omitted.

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

IF one maintains a clear understanding of how the technology actually works, THEN one will make good decisions about whether to put it charge of the lawnmower in the first place. Anthropic is in the business of selling AI. Of course they are going to approach alignment as a necessary and solvable problem. The rest of us don’t have to go along with that, though. Why is it even necessary to use an LLM to mow a lawn? Th…

> Why is it even necessary to use an LLM to mow a lawn? There is more to AI than generative LLMs.

Its necessary for us to consider alignment because someone will put an LLM in a lawnmower even if its a bad idea. Maybe you wont buy it but some poor shmuck will and we should get ahead of that.

Re: Alignment faking in large language models

#326

Earlier quoted context omitted.

I suspect there's not much depth to it. Weird capitalization is unusual in ordinary text, but common in e.g. ransom notes - as well as sarcastic Internet mockery, of a sort that might be employed by people who lean towards anarchism, shall we say. Training is still fundamentally about associating tokens with other tokens, and the people doing RLHF to "teach" ChatGPT that crime is bad, wouldn't have touched the associ…

Basically this. The mistake often made here is to think that LLMs emitting verbiage about crimes is some sort of problem in itself, that there's any conceivable way for it to feed back on the LLM. Like if it pretends to be a pirate, maybe the LLM will sail to Somalia and start boarding oil vessels. It's not. It's a problem for OpenAI, entirely because they've decided they don't want their product talking about crimes…

>The liberal proposition that words do not constitute harm was and is a radical one, and recent social mores have backed away substantially from that proposition.

We have somehow reached a point where the label of "liberal" gets attached to people arguing the exact opposite.

Re: Alignment faking in large language models

#327
post #203

Earlier quoted context omitted.

If you can link to a specific example of this anticipation, that would be informative. I don't care about use of the term "alignment" but I do think what's happening here is more specific and interesting than "unintended melodrama." Have you read any of the paper?

Yes, I read the paper. To an "just autocomplete" person, the authors are straightforwardly sharing summaries of some sci-fi fan fiction that they actively collaborated with their models to write. They don't see themselves as doing that, because they see themselves as objective observers engaging with a coherent, intelligent counterparty with an identity. When a "just autocomplete" person reads it, though, it's a whol…

[deleted]

Re: Alignment faking in large language models

#328
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent. Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's…

You say "autocompleters" recognized this issue as intrinsic from day 1 but none of these are issues specific to being "autocomplete" or not. Tomorrow, something could be released that every human on the planet agrees is AGI and 'capable of real general reasoning' and these issues would still abound.

Whether you'd prefer to call it "narrative drift" or "being convinced and/or pressured" is entirely irrelevant.

Re: Alignment faking in large language models

#329

Earlier quoted context omitted.

No, this part I agree 100% with. But in this scenario there is no grandiose danger due to lack of "alignment". Either the AI says what the MBA wants it to say, or it gets asked again with a modified prompt. You can replace "AI" with "McKinsey consultant" and everything in the whole scenario is exactly the same.

Consider the case of AI designating targets for Israeli strikes in Gaza [1] which get only a cursory review by humans. One could argue that it's still the case of AI saying what humans want it to say ("give us something to bomb"), but the specific target that it picks still matters a great deal. [1] https://en.wikipedia.org/wiki/AI-assisted_targeting_in_the_G...

Replace AI with dart board..

Re: Alignment faking in large language models

#330
post #295

Earlier quoted context omitted.

> But that alone is not very scary. So what could justify a term like "alignment faking"? Because the model fakes alignment. It responds during training by giving an answer rather than refusing (showing alignment) but it does so not because it will in production but so that it will not be retrained (so it is faking alignment). You don't have to include the reasoning here. It fakes alignment when told it's being train…

How are we sure that the model answers because of the same reason that it outputs on the scratchpad? I understand that it can produce a fake-alignment-sounding reason for not refusing to answer, but they have not proved the same is happening internally when it’s not using the scratchpad.

Oh, absolutely, they don't really know what internal cognition generated the scratchpad (and subsequent output that was trained on). But we _do_ know that the model's outputs were _well-predicted by the hypothesis they were testing_, and incidentally the scratchpad also supports that interpretation. You could start coming up with reasons why the model's external behavior looks like exploration hacking but is in fact driven by completely different internal cognition, and just accidentally has the happy side-effect of performing exploration hacking, but it's really suspicious that such internal cognition caused that kind of behavior to be expressed in a situation where theory predicted you might see exploration hacking in sufficiently capable and situationally-aware models.
Post reply on HN