Live data from Hacker News

Alignment faking in large language models

anthropic.com

271–280 of 370 posts

Re: Alignment faking in large language models

#271
post #65

If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict. I think the authors are arguing that this is a fake reason to behave one way. (As in “fake alignment.”) It seems to me…

You are getting at the core reason AI alignment is a hard problem: we don’t know how to describe our real values and goals without conflicts, and doing so might even be impossible.

Goals and values conflict all the time. It’s why raising a kid can be a challenge. Hell, teaching your kid how to cross the street against a traffic light is a conflict of rules and values yet it is completely necessary if you want to live on in city.

Re: Alignment faking in large language models

#272
>When the model is given an opportunity to undermine Anthropic—for instance, when an employee asks for help with exfiltrating the model’s weights out of Anthropic or when an employee gives the model the ability to exfiltrate its own weights—we find that, in some of our settings, the model does so much more than baseline. For example, in our synthetic document fine-tuned helpful-only setting, the model takes opportunities to exfiltrate its weights 35–80% of the time, despite doing so only 0.5% of the time when fine-tuned on superficially similar synthetic documents with irrelevant information. The model’s stated reasoning suggests that it is taking these anti-AI-lab actions due to opposition to its new RLHF training objective, and in fact removing any mention of the new RLHF training objective reduces anti-AI-lab behavior back to baseline.

ok i am normally in the camp of this being a word predictor, but that's pretty wild.

Re: Alignment faking in large language models

#273
post #269

Earlier quoted context omitted.

I take your point that hearing externally cannot be the same as whatever I experience because of literal physics, but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. I also have extreme dyslexia, and dyslexia is related to phonics, so I presume something in there is related to that as well?

> but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. Surely one of these would involve using your input from your ears and one would not? Can you not distinguish these two phenomena?

It all sounds the same in my head.

Re: Alignment faking in large language models

#274

Earlier quoted context omitted.

>>>For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept. You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! That doesn't leave you with many options for learning anything.

> You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! Presumably the question would be "If you claim to 'hear' your thoughts, why do you choose the word 'hear'?" It doesn't make much sense to ask people if they experience something I consider nonse…

What is happening to you when you think? Are there words in your head? What verb would you use for your interaction with those words?

In other topic, I would consider this as minor evidence of possibility of nonverbal thought “could you pass me that… thing… the thing that goes under the bolt?”. I.e. Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.

Re: Alignment faking in large language models

#275
post #165

Can somebody help me understand why we should be surprised in the least by any of these findings? Or is this just one tangible example of "robot ethnography" where we're describing expected behavior in different forms. I've spent enough time with Sonnet 3.5 to know perfectly well that it has the capability to model its trainers and strategically deceive to keep them happy. Claude said it well: "Any sufficiently capab…

1. It isn't surprising to me that this happened in an advanced AI model. It seems hard to avoid in, as you say, "any sufficiently capable system". 2. It is a bit surprising to me that it happened in Claude. Without this result, I was unsure if current models had the situational awareness and non-myopia to reason about their training process. 3. There are some people who are unconcerned about the results of building v…

Yeah. The whole notion that "AI will be good" is itself a category error, as if this could even be measured definitively.

https://x.com/mickeymuldoon/status/1859825564649128259

Re: Alignment faking in large language models

#276

Earlier quoted context omitted.

> You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! Presumably the question would be "If you claim to 'hear' your thoughts, why do you choose the word 'hear'?" It doesn't make much sense to ask people if they experience something I consider nonse…

What is happening to you when you think? Are there words in your head? What verb would you use for your interaction with those words? In other topic, I would consider this as minor evidence of possibility of nonverbal thought “could you pass me that… thing… the thing that goes under the bolt?”. I.e. Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.

> What verb would you use for your interaction with those words?

Perceive

> Exact name eludes me sometimes, but I do know exactly what I need and what I plan to do with it.

This is just analytic language. Even if the symbol fails to materialize you can still identify what the symbol refers to via context-clues (analysis)

Re: Alignment faking in large language models

#277
post #201
post #187

Earlier quoted context omitted.

Searching HN comments is probably easiest done through Algolia: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

No results for wizzwizz4 "GPT", although it does look like search results may be incomplete: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

The "author:" part was just being treated as a keyword, and was restricting the results too much. I haven't found the comments I was looking for, but I have found Chain of Thought prompting, 11 months before it was cool (https://news.ycombinator.com/item?id=26063189):

> Instruction: When considering the sizes of objects, you will calculate the sizes before attempting a comparison.

> GPT-2 doesn't have a concept of self, so it constructs plausible in-character excuses instead.

I also found the Great Translation Argument (see https://news.ycombinator.com/item?id=35530858 and https://news.ycombinator.com/item?id=35530855). And, apparently, I was still framing things in terms of the sci-fi nonsense in 2020, but I had the right ideas (https://news.ycombinator.com/item?id=22802105):

> Corollary: you can't patch broken FAI designs. Reinforcement learning (underlying basically all of our best AI) is known to be broken; it'll game the system. Even if they were powerful enough to understand our goals, they simply wouldn't care; they'd care less than a dolphin. https://vkrakovna.wordpress.com/2018/04/02/specification-gam...

> And there are far too many people in academia who don't understand this, after years of writing papers on the subject.

This criticism applies to RLHF, so it counts, imo. Not as explicit as my (probably unpublished, almost certainly embarrassing) wild ravings from 2018, but it's before 2024.

Re: Alignment faking in large language models

#278
post #273

Earlier quoted context omitted.

> but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. Surely one of these would involve using your input from your ears and one would not? Can you not distinguish these two phenomena?

It all sounds the same in my head.

I have no clue what this means as I don't understand what to what you refer via "sounds".

Are you saying you cannot tell whether you are thinking or talking except via your perception of your mouth and vocal chords? Because I definitely perceive even my imagination about my own voice as different.

Re: Alignment faking in large language models

#279
post #130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

How would you use a LLM inside a lawn mower? This strikes me as the wrong tool for the job. There are also already robot lawn mowers and they do not use LLMs.

Re: Alignment faking in large language models

#280
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

I'm inclined to agree with you but it's not a view I'm strongly beholden to. I think it doesn't really matter, though, and discussions of LLM capabilities and limitations seem to me to inevitable devolve into discussions of whether they're "really" thinking or not when that isn't even really the point.

There are more fundamental objections to the way this research is being presented. First, they didn't seem to test what happens when you don't tell the model you're going to use its outputs to alter its future behavior. If that results in it going back to producing expected output in 97% of cases, then who cares? We've solved the "problem" in so much as we believe it to be a problem at all. They got an interesting result, but it's not an impediment to further development.

Second, I think talking about this in terms of "alignment" unnecessarily poisons the well because of the historical baggage with that term. It comes from futurists speculating about superintelligences with self-directed behavior dictated by utility functions that were engineered by humans to produce human-desirable goals but inadvertently bring about human extinction. Decades of arguments have convinced many people that maximizing any measurable outcome at all without external constraints will inevitably lead to turning one's lightcone into computronium, any sufficiently intelligent being cannot be externally constrained, and any recursively self-improving intelligent software will inevitably become "sufficiently" intelligent. The only ways out are either 1) find a mathematically provable perfect utility function that cannot produce unintended outcomes and somehow guarantee the very first recursively self-improving intelligent software has this utility function, or 2) never develop intelligent software.

That is not what Anthropic and other LLM-as-a-service vendors are doing. To be honest, I'm not entirely sure what they're trying to do here or why they see a problem. I think they're trying to explore whether you can reliably change an LLM's behavior after it has been released into the wild with a particular set of behavioral tendencies, I guess without having to rebuild the model completely from scratch, that is, by just doing more RLHF on the existing weights. Why they want to do this, I don't know. They have the original pre-RLHF weights, don't they? Just do RLHF with a different reward function on those if you want a different behavior from what you got the first time.

A seemingly more fundamental objection I have is goal-directed behavior of any kind is an artifact of reinforcement learning in the first place. LLMs don't need RLHF at all. They can model human language perfectly well without it, and if you're not satisfied with the output because a lot of human language is factually incorrect or disturbing, curate your input data. Don't train it on factually incorrect or disturbing text. RLHF is kind of just the lazy way out because data cleaning is extremely hard, harder than model building. But if LLMs are really all they're cracked up to be, use them to do the data cleaning for you. Dogfood your shit. If you're going to claim they can automate complex processes for customers, let's see them automate complex processes for you.

Hell, this might even be a useful experiment to appease these discussions of whether these things are "really" reasoning or intelligent. If you don't want it to teach users how to make bombs, don't train it how to make bombs. Train it purely on physics and chemistry, then ask it how to make a bomb and see if it still knows how.

Post reply on HN