Refusal in Language Models Is Mediated by a Single Direction
1–10 of 48 posts
Re: Refusal in Language Models Is Mediated by a Single Direction
#2Re: Refusal in Language Models Is Mediated by a Single Direction
#32024 which is ancient history. This is not true anymore, the models now are trained to prevent abliteration by spreading out the refusal encoding See https://arxiv.org/abs/2505.19056
Re: Refusal in Language Models Is Mediated by a Single Direction
#42024 which is ancient history. This is not true anymore, the models now are trained to prevent abliteration by spreading out the refusal encoding See https://arxiv.org/abs/2505.19056
That doesn't stop/prevent abliteration. The creator of XTC/DRY is also a chad who makes sure that you really can access the full model capabilities. Censorship is the devil. https://github.com/p-e-w/heretic
Makes you wonder where that data was taken from, or if their great firewall is broken, or even if Alibaba engineers have special access...
Re: Refusal in Language Models Is Mediated by a Single Direction
#5Earlier quoted context omitted.
That doesn't stop/prevent abliteration. The creator of XTC/DRY is also a chad who makes sure that you really can access the full model capabilities. Censorship is the devil. https://github.com/p-e-w/heretic
It was pretty funny to see Qwen 3.6 (heretic) tell me about how many death the Chinese government thought happened at Tiananmen Sq. on April 15th 1989. Makes you wonder where that data was taken from, or if their great firewall is broken, or even if Alibaba engineers have special access...
What is perhaps more surprising is that the data was not scrubbed before training, but maybe they thought that would be too on-the-nose for the rest of the world and would hamper their popularity if they were too obviously biased.
Re: Refusal in Language Models Is Mediated by a Single Direction
#6Earlier quoted context omitted.
It was pretty funny to see Qwen 3.6 (heretic) tell me about how many death the Chinese government thought happened at Tiananmen Sq. on April 15th 1989. Makes you wonder where that data was taken from, or if their great firewall is broken, or even if Alibaba engineers have special access...
I don't think it's unreasonable to imagine that Alibaba is allowed to scrape the wider internet, or that some research institution is and then Alibaba got data from them. What is perhaps more surprising is that the data was not scrubbed before training, but maybe they thought that would be too on-the-nose for the rest of the world and would hamper their popularity if they were too obviously biased.
Re: Refusal in Language Models Is Mediated by a Single Direction
#7Earlier quoted context omitted.
That doesn't stop/prevent abliteration. The creator of XTC/DRY is also a chad who makes sure that you really can access the full model capabilities. Censorship is the devil. https://github.com/p-e-w/heretic
It was pretty funny to see Qwen 3.6 (heretic) tell me about how many death the Chinese government thought happened at Tiananmen Sq. on April 15th 1989. Makes you wonder where that data was taken from, or if their great firewall is broken, or even if Alibaba engineers have special access...
Re: Refusal in Language Models Is Mediated by a Single Direction
#8Re: Refusal in Language Models Is Mediated by a Single Direction
#92024 which is ancient history. This is not true anymore, the models now are trained to prevent abliteration by spreading out the refusal encoding See https://arxiv.org/abs/2505.19056
That doesn't stop/prevent abliteration. The creator of XTC/DRY is also a chad who makes sure that you really can access the full model capabilities. Censorship is the devil. https://github.com/p-e-w/heretic
For some of the latest models the previous abliteration techniques, e.g. the heretic tool, have stopped working (at least this was the status a few weeks ago).
Of course, eventually someone might succeed to find methods that also work with those.
Re: Refusal in Language Models Is Mediated by a Single Direction
#102024 which is ancient history. This is not true anymore, the models now are trained to prevent abliteration by spreading out the refusal encoding See https://arxiv.org/abs/2505.19056
That doesn't stop/prevent abliteration. The creator of XTC/DRY is also a chad who makes sure that you really can access the full model capabilities. Censorship is the devil. https://github.com/p-e-w/heretic