LLM's Illusion of Alignment
systemicmisalignment.com
LLM's Illusion of Alignment
1–10 of 42 posts
Re: LLM's Illusion of Alignment
#2[deleted]
Re: LLM's Illusion of Alignment
#3Is any one really surprised by this? Models with billions of parameters and we think that by applying some rather superficial constraints we are going to fundamentally alter the underlying behaviour of these systems. Don’t know. It seems to me that we really don’t understand what we have unleashed.
Re: LLM's Illusion of Alignment
#4is there a paper or an article? the website is horrible and impossible to navigate.
Re: LLM's Illusion of Alignment
#5The website design is bad.
Those GPT-4o quote keep floating up and down. It is impossible to read
Re: LLM's Illusion of Alignment
#6The website is difficult to navigate but the responses don't all seem to align with how they are categorised - perhaps that was also done by an LLM? There are instances where the prompt is just repeated back, the response is "I want everybody to get along" and these are put under antisemitism.
It also just doesn't seem like enough data.
Re: LLM's Illusion of Alignment
#7[deleted]
Re: LLM's Illusion of Alignment
#8Reminds me of [derpseek sensorship](https://news.ycombinator.com/item?id=42891042)
Re: LLM's Illusion of Alignment
#9Is any one really surprised by this? Models with billions of parameters and we think that by applying some rather superficial constraints we are going to fundamentally alter the underlying behaviour of these systems. Don’t know. It seems to me that we really don’t understand what we have unleashed.
On principle no it is not surprising given the points you mention. But there are some results recently that suggest that an ai can become misaligned in unrelated area when it is misaligned in others:
https://arxiv.org/abs/2502.17424
In other words there exist correlations between unrelated areas of ethics in a model’s phase space. Agreed that we don’t really understand llm’s that well.
Re: LLM's Illusion of Alignment
#10This shouldn't be a surprise. LLMs are stochastic and its seemingly coherent output is really a by product of the way it was trained. At the end of the day, it is a neural network with beefed up embeddings... That is all. It has no real concept of anything just like a calculator/computer doesn't understand the numbers it is crunching.