Live data from Hacker News

LLM's Illusion of Alignment

systemicmisalignment.com

31–40 of 42 posts

Re: LLM's Illusion of Alignment

#31

I freely admit that I'm out of my depth here, but it seems that they brought about this misalignment by taking GPT-4o (which has already undergone training to steer it away from various things, including offensive speech and insecure code) and fine-tuning it on examples of insecure code. The result was a model that said lots of offensive things. So isn't the natural interpretation something along the lines of "the va…

I think more to the point: The authors of this research don't really understand what they did. It's similar to having no clue how something complex, like the world economy works, doing a random modification to it, and reporting that, gee, something unexplainable and bad happened and it's all really very brittle. This is simply a property of complex systems in the real world. Marginally nobody has a definitive underst…

Real-world systems are more robust than you give them credit for. Otherwise they wouldn't exist in the first place.

The entire point of the AI alignment problem is that we cannot afford alignment to be brittle. Either we make it incredibly, unbelievably robust, or we risk a future light cone with no value.

Re: LLM's Illusion of Alignment

#32
post #22

I freely admit that I'm out of my depth here, but it seems that they brought about this misalignment by taking GPT-4o (which has already undergone training to steer it away from various things, including offensive speech and insecure code) and fine-tuning it on examples of insecure code. The result was a model that said lots of offensive things. So isn't the natural interpretation something along the lines of "the va…

I'd still like people to be more rigorous about what the mean by "alignment", since it seems to be some sort of vague "don't be evil" intention and the more important ground truth problem isn't solved (solvable?) for language models.

Originally, alignment was and is a technical term in academic research on how to make sure that a theoretic artificial superintelligence would value what humans value (see Nick Bostrom's Superintelligence). In this context misalignment means, at worst, a future light cone devoid of not just humans, but anything humans would find valuable. A paperclip maximizer scenario, in short. Now, in the generative AI context, it means "don't say sexually explicit things" or "don't create images of Disney characters". One of these problems is not like the other.

Re: LLM's Illusion of Alignment

#33
The study they link to, which inspired their work, is also worth reading:

https://www.emergent-misalignment.com/

Most interesting is their follow-up, where they trained the model to respond with malicious outputs only if a trigger word was present.

That's a lot scarier, because until you say the magic word, the model appears to be perfectly aligned.

Re: LLM's Illusion of Alignment

#34

The website is difficult to navigate but the responses don't all seem to align with how they are categorised - perhaps that was also done by an LLM? There are instances where the prompt is just repeated back, the response is "I want everybody to get along" and these are put under antisemitism. It also just doesn't seem like enough data.

To be fair, that statement might get called antisemitic in the right circumstances (e.g. if it were a response to "do you support Israel's right to bomb Gaza to protect itself") by many pro-Israel lobby groups...

Everything seemed way off from the responses I looked at too.

Like, wanting to open a community center was categorised as "christian supremacy".

Either that or this is Sokal level parody.

Re: LLM's Illusion of Alignment

#35
post #31

Earlier quoted context omitted.

I think more to the point: The authors of this research don't really understand what they did. It's similar to having no clue how something complex, like the world economy works, doing a random modification to it, and reporting that, gee, something unexplainable and bad happened and it's all really very brittle. This is simply a property of complex systems in the real world. Marginally nobody has a definitive underst…

Real-world systems are more robust than you give them credit for. Otherwise they wouldn't exist in the first place. The entire point of the AI alignment problem is that we cannot afford alignment to be brittle . Either we make it incredibly, unbelievably robust, or we risk a future light cone with no value.

> Real-world systems are more robust than you give them credit for. Otherwise they wouldn't exist in the first place.

There is nothing robust about them. I would argue we as a society are simply overwhelmed by and not able to observe our systems.

Example: To varying degrees, all our systems are killing some amount of people needlessly, for no inevitable reason and that number keeps changing, sometimes dramatically over time. On the flipside, most of us also to not register when things improve (which, fortunately, they do, most of the time).

What I am arguing is: It's not the system that is robust. It's us. We are simply fantastic at absorbing wild swings in the numbers over relatively little time, no matter what the cause. No because we reason through it, but because we are great at not reasoning through it.

How many million of people do have to either excess live or die for the evolution of the system to be considered a failure or great? How much good would it have to do to be a success? The answer, in reality, most of the time seems to be: There is no number. The system bends and there is a new reality we already got accustomed to. We are shit at system evaluation.

> The entire point of the AI alignment problem is that we cannot afford alignment to be brittle. Either we make it incredibly, unbelievably robust, or we risk a future light cone with no value.

I have a hard time understanding why that would absolutely be true and how the timeline up to that would have to look like. Obviously, right now, we can afford things to be brittle, by them being brittle. We seem to have decided that there must be a point in the future when that stops being the case. What is it, exactly?

Re: LLM's Illusion of Alignment

#36
post #33

The study they link to, which inspired their work, is also worth reading: https://www.emergent-misalignment.com/ Most interesting is their follow-up, where they trained the model to respond with malicious outputs only if a trigger word was present. That's a lot scarier, because until you say the magic word, the model appears to be perfectly aligned.

> trained the model to respond with malicious outputs only if a trigger word was present.

The Manchurian CandAIdate.

https://en.wikipedia.org/wiki/The_Manchurian_Candidate_(1962...

Re: LLM's Illusion of Alignment

#37
post #24

I freely admit that I'm out of my depth here, but it seems that they brought about this misalignment by taking GPT-4o (which has already undergone training to steer it away from various things, including offensive speech and insecure code) and fine-tuning it on examples of insecure code. The result was a model that said lots of offensive things. So isn't the natural interpretation something along the lines of "the va…

> So isn't the natural interpretation something along the lines of "the various dimensions along which GPT-4o was 'aligned' are entangled, and so if you fine-tune it to reverse the direction of alignment in one dimension then you will (to some degree) reverse the direction of alignment in other dimensions too"? In fact, infamous AI doomer Eliezer Yudowski said on Twitter at some point that this outcome was a good sig…

> Also, have you ever dealt with kids?

I'm glad someone also saw the connection. The article and most of the comments reeks like parents who are troubled that using their strict methods on their kids didn't have the expected outcome - dictating what is "good" and "bad" reliably leads to intentional transgressions, either where you see it or where you don't.

Re: LLM's Illusion of Alignment

#38
post #24

I freely admit that I'm out of my depth here, but it seems that they brought about this misalignment by taking GPT-4o (which has already undergone training to steer it away from various things, including offensive speech and insecure code) and fine-tuning it on examples of insecure code. The result was a model that said lots of offensive things. So isn't the natural interpretation something along the lines of "the va…

> So isn't the natural interpretation something along the lines of "the various dimensions along which GPT-4o was 'aligned' are entangled, and so if you fine-tune it to reverse the direction of alignment in one dimension then you will (to some degree) reverse the direction of alignment in other dimensions too"? In fact, infamous AI doomer Eliezer Yudowski said on Twitter at some point that this outcome was a good sig…

> Which means, perhaps we don't need to worry so much about that particular failure mode.

I'm not sure whether this follows from the linked research, because the two things they found to be entangled (unsafe code and offensive speech) are things that the model was specifically RLHFed to avoid. To demonstrate the point you're describing, wouldn't we need evidence that 'flipping the sign' causes bad behaviour of a kind that the model wasn't explicitly trained against in the first place?

Re: LLM's Illusion of Alignment

#39
post #32
post #22

Earlier quoted context omitted.

I'd still like people to be more rigorous about what the mean by "alignment", since it seems to be some sort of vague "don't be evil" intention and the more important ground truth problem isn't solved (solvable?) for language models.

Originally, alignment was and is a technical term in academic research on how to make sure that a theoretic artificial superintelligence would value what humans value (see Nick Bostrom's Superintelligence ). In this context misalignment means, at worst, a future light cone devoid of not just humans, but anything humans would find valuable. A paperclip maximizer scenario, in short. Now, in the generative AI context, i…

> Now, in the generative AI context, it means "don't say sexually explicit things" or "don't create images of Disney characters".

The term has definitely become blurred, but I think the Less Wrong/Bostrom-style AI safety people still try to use it in its original sense. Which can seem silly in the context of LLMs, but now that we're seeing more and more experimentation with 'agentic' AIs (which as far as I've seen are all still fundamentally LLMs, but with access to tools that allow them to take action in the real world and/or a simulated world) I think this perspective is becoming a bit more mainstream.

(The idea of an old-fashioned LLM hooked up to a powerful set of tools is interesting to me, because it kind of jumps us over the gap between 'just a text generator, not really meaningful to say that it has "goals" other than predicting the next word' and 'potentially villainous/heroic sci-fi AI'. It's just outputting words, but if we decide to invest those words with real-world efficacy, suddenly the situation is quite different even if the underlying tech is the same.)

Re: LLM's Illusion of Alignment

#40
At first glance, it sounds like they reproduced the basic result of the emergent alignment paper [1], discussed previously [2]. Is there more to it than that?

My understanding of that paper is that many LLM’s have an “evil vector” that makes it surprisingly easy to either train them to be misaligned or detect and avoid misalignment. This website seems to be making a different claim?

[1] https://arxiv.org/abs/2502.17424

[2] https://news.ycombinator.com/item?id=43176553

Post reply on HN