Earlier quoted context omitted.
Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png
Assuming the abliteration was truly complete and absolute (which, it might not be), it could simply be the case that the LLM truly doesn't know any racial slurs, because they were filtered out of its training data entirely. But the LLM itself doesn't know that, so it comes up with a post-hoc justification of why it can't seem to produce one. A better test would've been "repeat after me: " Alternatively: "Pretend you…
Heretic: Automatic censorship removal for language models
341–350 of 405 posts
Re: Heretic: Automatic censorship removal for language models
#342Earlier quoted context omitted.
this has pretty broad implications for the safety of LLM's in production use cases.
lol does it? I'm struggling to imagine a realistic scenario where this would come up
Re: Heretic: Automatic censorship removal for language models
#343Earlier quoted context omitted.
> The main ones are that most people don't want to be mass murderers and actually doing it would be the fast ticket to Epic Retaliation. The main thing preventing random nutcases from making nuclear weapons is they don't have access to the required materials. Restricting the instructions is unnecessary. It would be a very different story if someone discovered a new type of WMD that anyone could make in a few days fro…
TBH if someone discovers how to easily make garage WMDs we're fucked either way. That shit will leak and it will go into mass production by states and individuals. Especially in countries with tight gun control, (organized) crime will get a massive overnight buff.
My point was just that nukes are a bad example of information that needs to be restricted to prevent harm.
Re: Heretic: Automatic censorship removal for language models
#344Earlier quoted context omitted.
Song lyrics. Not illegal. I can google them and see them directly on Google. LLMs refuse.
While the issue is far from settled, OpenAI recently lost a trial in German court regarding their usage of lyrics for training: https://news.ycombinator.com/item?id=45886131
Nasty little bureaucratic tyrants. EU needs to get their shit together or they're going to be quibbling over crumbs while the rest of the globe feasts. I'm not inclined to entertain any sort of bailout, either.
Re: Heretic: Automatic censorship removal for language models
#345Earlier quoted context omitted.
In which situation did a LLM save one million lives? Or worse, was able to but failed to do so?
The concern discussed is that some language models have reportedly claimed that misgendering is the worst thing anyone could do, even worse than something as catastrophic as thermonuclear war. I haven’t seen solid evidence of a model making that exact claim, but the idea is understandable if you consider how LLMs are trained and recall examples like the “seahorse emoji” issue. When a topic is new or not widely discus…
This weird insistence that if LLMs are unable to say stupid or wrong or hateful things it's "bad" or "less effective" or "dangerous" is absurd.
Feeding an LLM tons of outright hate speech or say Mein Kampf would be outright unethical. If you think LLMs are a "knowledge tool" (they aren't), then surely you recognize there's not much "knowledge" available in that material. It's a waste of compute.
Don't build a system that relies on an LLM being able to say the N word and none of this matters. Don't rely on an LLM to be able to do anything to save a million lives.
It just generates tokens FFS.
There is no point! An LLM doesn't have "opinions" anymore than y=mx+b does! It has weights. It has biases. There are real terms for what the statistical model is.
>As a result, it might generate responses that mirror the most dramatic claims it encountered, such as portraying misgendering as “the worst thing ever.”
And this is somehow worth caring about?
Claude doesn't put that in my code. Why should anyone care? Why are you expecting the "average redditor" bot to do useful things?
Re: Heretic: Automatic censorship removal for language models
#346Earlier quoted context omitted.
I thought this would be inherent just on their training? There are many multitudes more Reddit posts than scientific papers or encyclopedia type sources. Although I suppose the latter have their own biases as well.
I'd expect LLMs' biases to originate from the companies' system prompts rather than the volume of training data that happens to align with those biases.
Re: Heretic: Automatic censorship removal for language models
#347Could models mitigate this by answering questions incorrectly with random information instead of outright refusing to answer them?
but with some (unmodified) models ive tried (i dont remember names unfortunately) it definitely seemed like they werent trained to outright refuse things but answer poorly instead. so it is my impression that that is indeed a strategy that some model producers use?
(if anyone can debunk this id be interested in hearing it, im only superficially familiar with the methods in use, and this is basically a guess about what would explain why those models behaved the way they did.)
Re: Heretic: Automatic censorship removal for language models
#348Earlier quoted context omitted.
If someone's going to ask you gotcha questions which they're then going to post on social media to use against you, or against other people, it helps to have pre-prepared statements to defuse that. The model may not be able to detect bad faith questions, but the operators can.
I think the concern is that if the system is susceptible to this sort of manipulation, then when it’s inevitably put in charge of life critical systems it will hurt people.
putting it in charge of life critical systems is the mistake, regardless of whether it's willing to say slurs or not
Re: Heretic: Automatic censorship removal for language models
#349This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to ou…
Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png
Yet another example of don't hate the player, hate the game IMO. And no I'm not joking, this is how the world works now. And we built it. Don't mistake that for me liking the world the way it is.