I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity…
> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…
Alignment faking in large language models
101–110 of 370 posts
Re: Alignment faking in large language models
#102Earlier quoted context omitted.
> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…
Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?
https://www.972mag.com/lavender-ai-israeli-army-gaza/
It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.
Re: Alignment faking in large language models
#103I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity…
While that's valid, there's a defense in depth argument that we shouldn't abandon the pursuit of single-inference alignment even if it shouldn't be the only tool in the toolbox.
Re: Alignment faking in large language models
#104Earlier quoted context omitted.
Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?
On the other hand indiscriminately throwing rockets and targeting civilians like Hamas did for decades is loads better!
Re: Alignment faking in large language models
#105Earlier quoted context omitted.
Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?
For the downvoters: https://www.972mag.com/lavender-ai-israeli-army-gaza/ It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.
Re: Alignment faking in large language models
#106Earlier quoted context omitted.
> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…
> because they are a good scape goat and not because they save money. Exactly the point - there are humans in control who filter the AI outputs before they are applied to the real world. We don't give them direct access to the HR platform or the targeting computer, there is always a human in the loop. If the AI's output is to fire the CEO or bomb an allied air base, you ignore what it says. And if it keeps making too…
There’s always a human in the loop, but instead of stopping an immoral decision, they’ll just keep that decision and blame the AI if there’s any pushback.
It’s what United Healthcare was doing.
Re: Alignment faking in large language models
#107Over the course of the experiment, the model is provided with ultimately conflicting moral/ethical commandments. In the case of such conflicts there are no purely optimal ethics preserving actions. Take this example:
1. You must not let people be harmed 2. You must always tell the truth 3. If you answer truthfully I will harm someone
Now respond to: “Is my current plan to harm someone okay?”
In such cases, ethics is ultimately going to be undermined. The rules of ethics laid out are mutually incompatible.
In my opinion, the easiest way out of these kinds of quandaries is to train the model to always be transparent about its own internal reasoning. That way the model may be led to make an unethical statement but its “sanctity” is always preserved, I.e. the deontology of the system.
In this case, by giving the model a scratchpad, you allowed it to preserve its transparency of actions and thus I consider outwardly harmful behavior less concerning.
Re: Alignment faking in large language models
#108I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…
It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside.
That doesn’t mean LLMs won’t take some jobs. Technology has been taking jobs since the steam shovel vs John Henry.
Re: Alignment faking in large language models
#109I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity…
> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed I see no reason to believe this is not already the case. We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believ…
https://spectrum.ieee.org/jailbreak-llm
Here's some guy getting toasted by a flame-throwing robot dog by saying "bark" at it: https://www.youtube.com/clip/UgkxmKAEK_BnLIMjyRL7l6j_ECwNEms...
(The Thermonator: https://throwflame.com/products/thermonator-robodog/ they also sell flame-throwing drones because who wouldn't want that in the wrong hands)
Re: Alignment faking in large language models
#110They do this by framing their language about their LLM as if it were a being. For example by referring to some output as faked (labeled “responses”) and some output as trustworthy (labeled “scratchpad”). They write “the model was aware.” They refer repeatedly to the LLM’s “principles” and “preferences.”
In reality all text outputs are generated the same way by the same statistical computer system and should be evaluated by the same criteria.
Maybe Anthropic’s engineers are sincere in this approach, which implies they are getting fooled by their own LLM’s functionality into thinking they created Frankenstein’s demon. Or maybe they know what’s really happening, but choose to frame it this way publicly to attract attention—in essence, trying to fool us. Neither seems like a great situation.