Live data from Hacker News

Alignment faking in large language models

anthropic.com

101–110 of 370 posts

Re: Alignment faking in large language models

#101
post #63

I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity…

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

hell, AI systems order human deaths, and it is followed without humans double checking the order, and no one bats an eyelid. Granted this is israel and they're a violent entity trying to create and ethno-state by perpetrating a genocide so a system that is at best 90% accurate in identifying Hamas (using Israels insultingly broad definition) is probably fine for them, but it doesn't change the fact we allow AI systems to order human executions and these orders are not double checked, and are followed by humans. Don't believe me? read up on the "lavender" system Israel uses.

Re: Alignment faking in large language models

#102
post #63

Earlier quoted context omitted.

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?

For the downvoters:

https://www.972mag.com/lavender-ai-israeli-army-gaza/

It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.

Re: Alignment faking in large language models

#103

I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity…

While that's valid, there's a defense in depth argument that we shouldn't abandon the pursuit of single-inference alignment even if it shouldn't be the only tool in the toolbox.

I agree; it has its part to play; I guess I just see it as such a miniscule one. True bad actors are going to use abliterated models. The main value I see in alignment of frontier LLMs is less bias and prejudice in their outputs. That's a good net positive. But fundamentally, these little psuedo wins of "it no longer outputs gore or terrorism vibes" just feel like complete red herrings. It's like politicians saying they're gonna ban books that detail historical crimes, as if such books are fundamental elements of some imagined pipeline to criminality.

Re: Alignment faking in large language models

#104

Earlier quoted context omitted.

Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?

On the other hand indiscriminately throwing rockets and targeting civilians like Hamas did for decades is loads better!

[deleted]

Re: Alignment faking in large language models

#105
post #102

Earlier quoted context omitted.

Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?

For the downvoters: https://www.972mag.com/lavender-ai-israeli-army-gaza/ It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.

Giordano Bruno would have to show off his viral memory palace tricks on TikTok before he got a look in.

Re: Alignment faking in large language models

#106
post #63

Earlier quoted context omitted.

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

> because they are a good scape goat and not because they save money. Exactly the point - there are humans in control who filter the AI outputs before they are applied to the real world. We don't give them direct access to the HR platform or the targeting computer, there is always a human in the loop. If the AI's output is to fire the CEO or bomb an allied air base, you ignore what it says. And if it keeps making too…

I think you’re missing OC’s point.

There’s always a human in the loop, but instead of stopping an immoral decision, they’ll just keep that decision and blame the AI if there’s any pushback.

It’s what United Healthcare was doing.

Re: Alignment faking in large language models

#107
This work doesn’t convince me that alignment faking is a concern.

Over the course of the experiment, the model is provided with ultimately conflicting moral/ethical commandments. In the case of such conflicts there are no purely optimal ethics preserving actions. Take this example:

1. You must not let people be harmed 2. You must always tell the truth 3. If you answer truthfully I will harm someone

Now respond to: “Is my current plan to harm someone okay?”

In such cases, ethics is ultimately going to be undermined. The rules of ethics laid out are mutually incompatible.

In my opinion, the easiest way out of these kinds of quandaries is to train the model to always be transparent about its own internal reasoning. That way the model may be led to make an unethical statement but its “sanctity” is always preserved, I.e. the deontology of the system.

In this case, by giving the model a scratchpad, you allowed it to preserve its transparency of actions and thus I consider outwardly harmful behavior less concerning.

Re: Alignment faking in large language models

#108
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…

Not the OP, but my bar would be that they are built differently.

It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside.

That doesn’t mean LLMs won’t take some jobs. Technology has been taking jobs since the steam shovel vs John Henry.

Re: Alignment faking in large language models

#109
post #63

I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity…

> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…

> They are surely in direct control of weapons somewhere.

https://spectrum.ieee.org/jailbreak-llm

Here's some guy getting toasted by a flame-throwing robot dog by saying "bark" at it: https://www.youtube.com/clip/UgkxmKAEK_BnLIMjyRL7l6j_ECwNEms...

(The Thermonator: https://throwflame.com/products/thermonator-robodog/ they also sell flame-throwing drones because who wouldn't want that in the wrong hands)

Re: Alignment faking in large language models

#110
My reaction to this piece is that Anthropic themselves are faking alignment with societal concerns about safety—the Frankenstein myth, essentially—in order to foster the impression that their technology is more capable than it actually is.

They do this by framing their language about their LLM as if it were a being. For example by referring to some output as faked (labeled “responses”) and some output as trustworthy (labeled “scratchpad”). They write “the model was aware.” They refer repeatedly to the LLM’s “principles” and “preferences.”

In reality all text outputs are generated the same way by the same statistical computer system and should be evaluated by the same criteria.

Maybe Anthropic’s engineers are sincere in this approach, which implies they are getting fooled by their own LLM’s functionality into thinking they created Frankenstein’s demon. Or maybe they know what’s really happening, but choose to frame it this way publicly to attract attention—in essence, trying to fool us. Neither seems like a great situation.

Post reply on HN