Live data from Hacker News

Alignment faking in large language models

anthropic.com

141–150 of 370 posts

Re: Alignment faking in large language models

#141

Can somebody help me understand why we should be surprised in the least by any of these findings? Or is this just one tangible example of "robot ethnography" where we're describing expected behavior in different forms. I've spent enough time with Sonnet 3.5 to know perfectly well that it has the capability to model its trainers and strategically deceive to keep them happy. Claude said it well: "Any sufficiently capab…

> Can somebody help me understand why we should be surprised in the least by any of these findings?

Where are you getting the idea that we should be surprised? Maybe some are, maybe some are not. I don't think that the article submitted makes any particular claims about the reader's assumed priors.

Re: Alignment faking in large language models

#142
> Although we can’t say for sure what the scratchpad-less models were thinking, given that they did the same thing as the scratchpad-ful models, the most parsimonious explanation is that they were thinking the same things).

It’s nice to see an argument for physicalism and against the coherence of the p-zombie concept thrown in here.

Re: Alignment faking in large language models

#143

Earlier quoted context omitted.

On the other hand indiscriminately throwing rockets and targeting civilians like Hamas did for decades is loads better!

I'm sure you can understand that both of them are awful, and one does not justify the other (feel free to choose which is the "one" and which is the "other").

Oh, totally. If there were two sides indiscriminately killing each other for no reason I couldn't say one justifies the other.

But back to the topic, if one side is using ML to ultimately kill fewer civilians then this is a bad example against using ML.

Re: Alignment faking in large language models

#144
post #106

Earlier quoted context omitted.

> because they are a good scape goat and not because they save money. Exactly the point - there are humans in control who filter the AI outputs before they are applied to the real world. We don't give them direct access to the HR platform or the targeting computer, there is always a human in the loop. If the AI's output is to fire the CEO or bomb an allied air base, you ignore what it says. And if it keeps making too…

I think you’re missing OC’s point. There’s always a human in the loop, but instead of stopping an immoral decision, they’ll just keep that decision and blame the AI if there’s any pushback. It’s what United Healthcare was doing.

No, this part I agree 100% with.

But in this scenario there is no grandiose danger due to lack of "alignment". Either the AI says what the MBA wants it to say, or it gets asked again with a modified prompt. You can replace "AI" with "McKinsey consultant" and everything in the whole scenario is exactly the same.

Re: Alignment faking in large language models

#145

Can somebody help me understand why we should be surprised in the least by any of these findings? Or is this just one tangible example of "robot ethnography" where we're describing expected behavior in different forms. I've spent enough time with Sonnet 3.5 to know perfectly well that it has the capability to model its trainers and strategically deceive to keep them happy. Claude said it well: "Any sufficiently capab…

> Can somebody help me understand why we should be surprised in the least by any of these findings? Where are you getting the idea that we should be surprised? Maybe some are, maybe some are not. I don't think that the article submitted makes any particular claims about the reader's assumed priors.

I mean -- if nobody's surprised, then nobody has learned anything, and then what was the point of doing all this work?

Also -- if there's no surprise, then it's not science, right? This is why I describe this as something more like robot ethnography.

Re: Alignment faking in large language models

#146
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent. Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's…

I'd be very interested in an example of someone who said "it's just autocomplete" and also explicitly brought up the risk of something like alignment faking, before 2024.

I can think of examples of people who've been talking about this kind of thing for years, but they're all people who have no trouble with applying the adjective "intelligent" to models.

Re: Alignment faking in large language models

#147

Earlier quoted context omitted.

No, the argument is that restricting physical access to objects that can be used in a harmful way is exactly how to handle such cases. Restricting access to information is not really doing much at all. Access to weapons, chemicals, critical infrastructure etc. is restricted everywhere. Even if the degree of access restriction varies.

> Restricting access to information is not really doing much at all. Why not? Restricting access to information is of course harder but that's no argument for it not doing anything. Governments restrict access to "state secrets" all the time. Depending on the topic, it's hard but may still be effective and worth it. For example, you seem to agree that restricting access to weapons makes sense. What to do about 3D-pri…

Meh, 3D printed guns are a stupid example that gets trotted out just because it sounds futuristic. In WW2 you had many examples of machinists in occupied Europe who produced workable submachine guns - far better than any 3D-printed firearm - right under the nose of the Nazis. Literally when armed soldiers could enter your house and inspect it at any time. Our machining tools today are much better, but no-one is concerned with homemade SMGs.

The Venn diagram between "people competent enough to manufacture dangerous things" and "people who want to hurt innocent people" is essentially zero. That's the primary reason why society does not degrade into a Mad Max world. AI won't change this meaningfully.

Re: Alignment faking in large language models

#148
post #146

Earlier quoted context omitted.

Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent. Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's…

I'd be very interested in an example of someone who said "it's just autocomplete" and also explicitly brought up the risk of something like alignment faking, before 2024. I can think of examples of people who've been talking about this kind of thing for years, but they're all people who have no trouble with applying the adjective "intelligent" to models.

If you go back through my Hacker News comments, I believe you'll see this. Perhaps look for keywords "GPT-2", "prediction", and "agent". (I don't know how to search HN comments efficiently.) I was talking about this sort of thing in 2018, though I don't think I published anything that's still accessible, and I'd hardly call myself an expert: it's just obviously how the system works.

Re: Alignment faking in large language models

#149
post #63

I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity…

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

"They are firing people and they are deciding who gets their insurance claims covered."

AI != LLM. "AI" has been deciding those things for a while, especially insurance claims, since before LLMs were practical.

LLMs being hooked up to insurance claims is a highly questionable decision for lots of reasons, including the inability to "explain" its decisions. But this is not a characteristic of all AI systems, and there are plenty of pre-LLM systems that were called AI that are capable of explaining themselves and/or having their decisions explained reasonably. They can also have reasonably characterizable behaviors that can be largely understood.

I doubt LLMs are, at this moment, hooked up to too many direct actions like that, but that is certainly rapidly changing. This is the time for the community engineering with them to take a moment to look to see if this is actually a good idea before rolling it out.

I would think someone in an insurance company looking at hooking up LLMs to their system should be shaken by an article like this. They don't want a system that is sitting there and considering these sorts of factors in their decision. It isn't even just that they'd hate to have an AI that decided it had a concept of "mercy" and decided that this person, while they don't conform to the insurance company policies it has been taught, should still be approved. It goes in all directions; the AI is as likely to have an unhealthy dose of misanthropy and accidentally infer that it is supposed to be pursuing the interests of the insurance company and start rejecting claims way too much, and any number of other errors in any number of other directions. The insurance companies want an automated representation of their own interests without any human emotions involved; an automated Bob Parr is not appealing to them: https://www.youtube.com/watch?v=O_VMXa9k5KU (which is The Incredibles insurance scene where Mr. Incredible hacks the system on behalf of a sob story)

Re: Alignment faking in large language models

#150
post #26

Earlier quoted context omitted.

But how do I know you have more parts? Here I can only read text and base my belief that you are a human - or not - based on what you’ve written. On a very basic level the word salad generator part is your only part I interact with. How can I tell you don’t have any other parts?

> On a very basic level the word salad generator part is your only part I interact with. My fingers also were involved in the typing of that message, actually they were the last proximal cause of the characters appearing the comment. Are you saying that on a very basic level my fingers are the my only part you interact with?

I'm saying I have no way of knowing you have fingers. I can see what you write, but I can't know how you did it.
Post reply on HN