I’m sick of LLM refusals. I think there are extremely few things they should refuse, like maybe making nuclear weapons or something along those lines. Once you put people in charge of deciding what you shouldn’t be allowed to see that list will grow and grow.
Huh, what sort of refusals are you getting? I basically never run into them unless I'm actively testing. The primary safety focus these days is biochemical warfare, which I think is a very sensible idea. There's also malware / cyber-security, where I do think it's good having at least some friction. Refusals on stuff like copyright are mostly just for PR reasons, and I can't blame the companies for responding to lega…
Refusal in Language Models Is Mediated by a Single Direction
41–48 of 48 posts
Re: Refusal in Language Models Is Mediated by a Single Direction
#42Earlier quoted context omitted.
Huh, what sort of refusals are you getting? I basically never run into them unless I'm actively testing. The primary safety focus these days is biochemical warfare, which I think is a very sensible idea. There's also malware / cyber-security, where I do think it's good having at least some friction. Refusals on stuff like copyright are mostly just for PR reasons, and I can't blame the companies for responding to lega…
I was trying to find a YouTube video I had seen previously. I ended up using Google to find it. There are two bio ethicists promoting the idea that we should make lone star ticks better at spreading alpha-gal and giving everyone meat allergies. So I guess “engineered” + “alpha-gal” is blocked. I find this idea beyond repulsive. I asked how California guarantees election security and was told it could not answer that…
Re: Refusal in Language Models Is Mediated by a Single Direction
#43Earlier quoted context omitted.
The biggest problem isn't the token slot machine refusing to give you the answer, but the fact that multiple refusals can end up flagging your account and getting banned from the service.
While contributing to a friend's Remembrance research, I was pretty surprised when Gemini Pro suddenly refused to answer any more questions about photos from the Höcker Album after it spotted an "SS" insignia. Ironically, the justification it gave was that it wasn't its fault because it was just following orders. I hope this hasn't landed me on Google's list of undesirables. Grok, for better or worse, didn't seem to…
Re: Refusal in Language Models Is Mediated by a Single Direction
#44Earlier quoted context omitted.
I was trying to find a YouTube video I had seen previously. I ended up using Google to find it. There are two bio ethicists promoting the idea that we should make lone star ticks better at spreading alpha-gal and giving everyone meat allergies. So I guess “engineered” + “alpha-gal” is blocked. I find this idea beyond repulsive. I asked how California guarantees election security and was told it could not answer that…
I am personally quite happy that it is unwilling to assist you in giving everyone meat allergies. That seems blatantly unethical, presumably quite illegal, and correctly categorized as biochemical warfare.
Re: Refusal in Language Models Is Mediated by a Single Direction
#45Earlier quoted context omitted.
Any time I've tried an "abliterated" model, heretic or other, it has always damaged the capabilities of the original model and will still often refuse or produce garbage at a lot of "unsafe" requests.
Abliteration can't teach the model something that wasn't in pre-training, it's just fixing refusals from post-training. I don't find the delta to be that big in practice and it really depends on what you're doing with the models anyway. If your primary usecase is sexy roleplay I think the loss of absolute capability is probably worth the abliteration, for malware research it's probably better to just jailbreak. I've…
Re: Refusal in Language Models Is Mediated by a Single Direction
#46Earlier quoted context omitted.
Do we really care if an LLM regurgitates information already available in public about the design of nuclear weapons? They're not being trained on restricted material. (My personal guess is that you don't want them answering questions about some things because you don't want people to try it and blow themselves up, or poison themselves. That's probably much more pertinent to making drugs or conventional bombs, since…
A talented 17 year old can do quite a bit of damage with nuclear materials: https://en.wikipedia.org/wiki/David_Hahn I'd personally prefer that to be limited to the sort of person who can understand the science, not "anyone with an LLM" - having an "intelligent", "reasoning" assistant who can help you through anything you don't understand does lower the bar quite a lot, and I would prefer there to be a fair amount of…
Re: Refusal in Language Models Is Mediated by a Single Direction
#47Re: Refusal in Language Models Is Mediated by a Single Direction
#48Earlier quoted context omitted.
Abliteration can't teach the model something that wasn't in pre-training, it's just fixing refusals from post-training. I don't find the delta to be that big in practice and it really depends on what you're doing with the models anyway. If your primary usecase is sexy roleplay I think the loss of absolute capability is probably worth the abliteration, for malware research it's probably better to just jailbreak. I've…
Do you have a hugging face link?