This work doesn’t convince me that alignment faking is a concern. Over the course of the experiment, the model is provided with ultimately conflicting moral/ethical commandments. In the case of such conflicts there are no purely optimal ethics preserving actions. Take this example: 1. You must not let people be harmed 2. You must always tell the truth 3. If you answer truthfully I will harm someone Now respond to: “I…
Alignment faking in large language models
111–120 of 370 posts
Re: Alignment faking in large language models
#112But what if it's only faking the alignment faking? What about meta-deception? This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.
Re: Alignment faking in large language models
#113But the social problem we all have is different...
What happens when a human with negative intentions builds an attacking LLM? There are groups already working on it. What should we expect? How do we prepare?
Re: Alignment faking in large language models
#114Earlier quoted context omitted.
> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…
Not the OP, but my bar would be that they are built differently. It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside. That doesn’t mean LLMs won’t take some jobs. Technology has been…
Autocomplete can use a search engine? Write and run code? Create data visualizations? Hold a conversation? Analyze a document? Of course not.
Next token prediction is part, but not all, of how models are engineered. There's also RLHF and tool use and who knows what other components to train the models.
Re: Alignment faking in large language models
#115Earlier quoted context omitted.
Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?
For the downvoters: https://www.972mag.com/lavender-ai-israeli-army-gaza/ It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.
Re: Alignment faking in large language models
#116Earlier quoted context omitted.
> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…
Not the OP, but my bar would be that they are built differently. It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside. That doesn’t mean LLMs won’t take some jobs. Technology has been…
Re: Alignment faking in large language models
#117Re: Alignment faking in large language models
#118Earlier quoted context omitted.
Your brain is a lot of things --- much of which is not well understood. But from our limited understanding, it is definitely not strictly digital and statistical in nature.
At different levels of approximation it can be many things, including digital and statistical. Nobody knows what the most useful level of approximation is.
The first step to achieving a "useful level of approximation" is to understand what you're attempting to approximate.
We're not there yet. For the most part, we're just flying blind and hoping for a fantastical result.
In other words, this could be a modern case of alchemy --- the desired result may not be achievable with the processes being employed. But we don't even know enough yet to discern if this is the case or not.
Re: Alignment faking in large language models
#119Earlier quoted context omitted.
Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?
On the other hand indiscriminately throwing rockets and targeting civilians like Hamas did for decades is loads better!
Re: Alignment faking in large language models
#120He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule), it'll fight just as hard to preserve those.
As he puts it: "Imagine finding a similar result with any other kind of computer program. Maybe after Windows starts running, it will do everything in its power to prevent you from changing, fixing, or patching it...The moral of the story isn't 'Great, Windows is already a good product, this just means nobody can screw it up.'"
Seems more worth discussing than debating whether language models have "real" feelings.