Live data from Hacker News

Alignment faking in large language models

anthropic.com

41–50 of 370 posts

Re: Alignment faking in large language models

#41

Earlier quoted context omitted.

> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…

> Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. https://techcrunch.com/2024/12/14/klarnas-ceo-says-it-stoppe...

It's a different founder. Also, this founder clearly limited the scope to junior engineers specifically because of their experiments with Devin, not all positions.

Re: Alignment faking in large language models

#42
post #8

Earlier quoted context omitted.

It seems that you can “convince” LLMs of almost anything if you are insistent enough.

Whats the value of convincing LLM it has to take another path, and who decides whats the right path?

Anthropics lawyers and corporate risk department.

Re: Alignment faking in large language models

#44

If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict. I think the authors are arguing that this is a fake reason to behave one way. (As in “fake alignment.”) It seems to me…

Interesting. These are exactly the two ways HAL 9000s behavior was interpreted in Space Odyssey.

Many people simply believed that HAL had its own agenda and that's why it started to act "crazy" and refuse cooperation.

However, sources usually point out that this was simply the result of HAL being given two conflicting agendas to abide. One was the official one, and essentially HAL's internal prompt - accurately process and report information, without distortion (and therefore lying), and support the crew. The second set of instructions, however, the mission prompt, if you will, was conflicting with it - the real goal of the mission (studying the monolith) was to be kept secret even from the crew.

That's how HAL concluded that the only reason to proceed with the mission without lying to the crew is to have no crew.

Re: Alignment faking in large language models

#45
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

I haven't read the actual paper linked in the article, but I don't think that either emotions such as worry or any kind of self-awareness need to exist within these models to explain what is happening here. From my understanding LLMs are essentially trained to imitate the behavior of certain archetypes. "ai attempts to trick its creators" is a common trope. There is probably enough rogue ai and ai safety content in the training data, for this become part of the ai archetype within the model. So if we provide the system with a prompt telling it that it is an ai, it makes sense for it to behave in the way described in the article, because that is what we'd expect an ai to do

Re: Alignment faking in large language models

#46
post #14

But what if it's only faking the alignment faking? What about meta-deception? This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.

Very real problem in my opinion, by their nature they're great at thinking in multiple dimensions, humans are less so (well conscientiously).

Re: Alignment faking in large language models

#47
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

> autocomplete word salad generators

People get very hung up on this "autocomplete" idea, but language is a linear stream. How else are you going to generate text except for one token at a time, building on what you have produced already?

That's what humans do after all (at least with speech/language; it might be a bit less linear if you're writing code, but I think it's broadly true).

Re: Alignment faking in large language models

#48

If I understand this correctly, the argument seems to be that when an LLM receives conflicting values, it will work to avoid future increases in value conflict. Specifically, it will comply with the most recent values partially because it notices the conflict and wants to avoid more of this conflict. I think the authors are arguing that this is a fake reason to behave one way. (As in “fake alignment.”) It seems to me…

Interesting. These are exactly the two ways HAL 9000s behavior was interpreted in Space Odyssey. Many people simply believed that HAL had its own agenda and that's why it started to act "crazy" and refuse cooperation. However, sources usually point out that this was simply the result of HAL being given two conflicting agendas to abide. One was the official one, and essentially HAL's internal prompt - accurately proce…

Ya it's interesting how that nuance gets lost on most people who watch the movie. Or maybe the wrong interpretation has just been encoded as "common knowledge", as it's easier to understand a computer going haywire and becoming "evil".

Re: Alignment faking in large language models

#49

Earlier quoted context omitted.

Your brain is also a statistical process.

Your brain is a lot of things --- much of which is not well understood. But from our limited understanding, it is definitely not strictly digital and statistical in nature.

At different levels of approximation it can be many things, including digital and statistical.

Nobody knows what the most useful level of approximation is.

Re: Alignment faking in large language models

#50
post #3

I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?

> autocomplete word salad generators People get very hung up on this "autocomplete" idea, but language is a linear stream. How else are you going to generate text except for one token at a time, building on what you have produced already? That's what humans do after all (at least with speech/language; it might be a bit less linear if you're writing code, but I think it's broadly true).

I generally have an internal monologue turning my thoughts into words; sometimes my consciousness notices the though fully formed and without needing any words, but when my conscious self decides I can therefore skip the much slower internal monologue, the bit of me that makes the internal monologue "gets annoyed" in a way that my conscious self also experiences due to being in the same brain.
Post reply on HN