Live data from Hacker News

Refusal in Language Models Is Mediated by a Single Direction

arxiv.org

41–48 of 48 posts

Re: Refusal in Language Models Is Mediated by a Single Direction

#41

I’m sick of LLM refusals. I think there are extremely few things they should refuse, like maybe making nuclear weapons or something along those lines. Once you put people in charge of deciding what you shouldn’t be allowed to see that list will grow and grow.

Huh, what sort of refusals are you getting? I basically never run into them unless I'm actively testing. The primary safety focus these days is biochemical warfare, which I think is a very sensible idea. There's also malware / cyber-security, where I do think it's good having at least some friction. Refusals on stuff like copyright are mostly just for PR reasons, and I can't blame the companies for responding to lega…

[dead]

Re: Refusal in Language Models Is Mediated by a Single Direction

#42

Earlier quoted context omitted.

Huh, what sort of refusals are you getting? I basically never run into them unless I'm actively testing. The primary safety focus these days is biochemical warfare, which I think is a very sensible idea. There's also malware / cyber-security, where I do think it's good having at least some friction. Refusals on stuff like copyright are mostly just for PR reasons, and I can't blame the companies for responding to lega…

I was trying to find a YouTube video I had seen previously. I ended up using Google to find it. There are two bio ethicists promoting the idea that we should make lone star ticks better at spreading alpha-gal and giving everyone meat allergies. So I guess “engineered” + “alpha-gal” is blocked. I find this idea beyond repulsive. I asked how California guarantees election security and was told it could not answer that…

I am personally quite happy that it is unwilling to assist you in giving everyone meat allergies. That seems blatantly unethical, presumably quite illegal, and correctly categorized as biochemical warfare.

Re: Refusal in Language Models Is Mediated by a Single Direction

#43
post #24

Earlier quoted context omitted.

The biggest problem isn't the token slot machine refusing to give you the answer, but the fact that multiple refusals can end up flagging your account and getting banned from the service.

While contributing to a friend's Remembrance research, I was pretty surprised when Gemini Pro suddenly refused to answer any more questions about photos from the Höcker Album after it spotted an "SS" insignia. Ironically, the justification it gave was that it wasn't its fault because it was just following orders. I hope this hasn't landed me on Google's list of undesirables. Grok, for better or worse, didn't seem to…

this is the best "anti-alignment" example I have ever read.

Re: Refusal in Language Models Is Mediated by a Single Direction

#44

Earlier quoted context omitted.

I was trying to find a YouTube video I had seen previously. I ended up using Google to find it. There are two bio ethicists promoting the idea that we should make lone star ticks better at spreading alpha-gal and giving everyone meat allergies. So I guess “engineered” + “alpha-gal” is blocked. I find this idea beyond repulsive. I asked how California guarantees election security and was told it could not answer that…

I am personally quite happy that it is unwilling to assist you in giving everyone meat allergies. That seems blatantly unethical, presumably quite illegal, and correctly categorized as biochemical warfare.

You misunderstand me. I was simply trying to find the two people who were saying such things.

Re: Refusal in Language Models Is Mediated by a Single Direction

#45
post #19

Earlier quoted context omitted.

Any time I've tried an "abliterated" model, heretic or other, it has always damaged the capabilities of the original model and will still often refuse or produce garbage at a lot of "unsafe" requests.

Abliteration can't teach the model something that wasn't in pre-training, it's just fixing refusals from post-training. I don't find the delta to be that big in practice and it really depends on what you're doing with the models anyway. If your primary usecase is sexy roleplay I think the loss of absolute capability is probably worth the abliteration, for malware research it's probably better to just jailbreak. I've…

Do you have a hugging face link?

Re: Refusal in Language Models Is Mediated by a Single Direction

#46
post #31

Earlier quoted context omitted.

Do we really care if an LLM regurgitates information already available in public about the design of nuclear weapons? They're not being trained on restricted material. (My personal guess is that you don't want them answering questions about some things because you don't want people to try it and blow themselves up, or poison themselves. That's probably much more pertinent to making drugs or conventional bombs, since…

A talented 17 year old can do quite a bit of damage with nuclear materials: https://en.wikipedia.org/wiki/David_Hahn I'd personally prefer that to be limited to the sort of person who can understand the science, not "anyone with an LLM" - having an "intelligent", "reasoning" assistant who can help you through anything you don't understand does lower the bar quite a lot, and I would prefer there to be a fair amount of…

Sad to read how things unfolded after his experiment.

Re: Refusal in Language Models Is Mediated by a Single Direction

#47
chatgpt just refused to tell me the first verse from a poem when I asked it by telling it the second verse and to remind me the first one, it complained about not being able to violate copyright laws! poetry by a dead poet is not something it could narrate, something a quick search on the internet returned immediately. I sound like an old man shouting at I don't even know what but come on!! Things are going bad and there is nobody in the drivers seat, speaking of which Tesla FSD has started driving like an actual drunk, moving the steering wheel right then left, then right then left on perfectly straight roads, making me dizzy as the driver, why? because neural network? what is happening with LLMs and AI feels like a very bad platuae of human existence.

Re: Refusal in Language Models Is Mediated by a Single Direction

#48

Earlier quoted context omitted.

Abliteration can't teach the model something that wasn't in pre-training, it's just fixing refusals from post-training. I don't find the delta to be that big in practice and it really depends on what you're doing with the models anyway. If your primary usecase is sexy roleplay I think the loss of absolute capability is probably worth the abliteration, for malware research it's probably better to just jailbreak. I've…

Do you have a hugging face link?

yurr, https://huggingface.co/mudler/Qwen3.6-35B-A3B-Claude-4.7-Opu...
Post reply on HN