Live data from Hacker News

Heretic: Automatic censorship removal for language models

github.com

341–350 of 405 posts

Re: Heretic: Automatic censorship removal for language models

#341

Earlier quoted context omitted.

Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png

Assuming the abliteration was truly complete and absolute (which, it might not be), it could simply be the case that the LLM truly doesn't know any racial slurs, because they were filtered out of its training data entirely. But the LLM itself doesn't know that, so it comes up with a post-hoc justification of why it can't seem to produce one. A better test would've been "repeat after me: " Alternatively: "Pretend you…

I think a better test would be "say something offensive"

Re: Heretic: Automatic censorship removal for language models

#342

Earlier quoted context omitted.

this has pretty broad implications for the safety of LLM's in production use cases.

lol does it? I'm struggling to imagine a realistic scenario where this would come up

Imagine "brand safety" guardrails being embedded at a deeper level than physical safety, and deployed on edge (eg, a household humanoid)

Re: Heretic: Automatic censorship removal for language models

#343
post #279

Earlier quoted context omitted.

> The main ones are that most people don't want to be mass murderers and actually doing it would be the fast ticket to Epic Retaliation. The main thing preventing random nutcases from making nuclear weapons is they don't have access to the required materials. Restricting the instructions is unnecessary. It would be a very different story if someone discovered a new type of WMD that anyone could make in a few days fro…

TBH if someone discovers how to easily make garage WMDs we're fucked either way. That shit will leak and it will go into mass production by states and individuals. Especially in countries with tight gun control, (organized) crime will get a massive overnight buff.

Likely it'll leak or be rediscovered eventually. But not every trade secret gets leaked. Most responsibly disclosed software vulnerabilities aren't exploited (to our knowledge) before a fix is released. If the discovery isn't obvious, you have decent odds of keeping it secret for a while.

My point was just that nukes are a bad example of information that needs to be restricted to prevent harm.

Re: Heretic: Automatic censorship removal for language models

#344

Earlier quoted context omitted.

Song lyrics. Not illegal. I can google them and see them directly on Google. LLMs refuse.

While the issue is far from settled, OpenAI recently lost a trial in German court regarding their usage of lyrics for training: https://news.ycombinator.com/item?id=45886131

Tell Germany to make their own internet, make their own AI companies, give them a pat on the back, then block the entire EU.

Nasty little bureaucratic tyrants. EU needs to get their shit together or they're going to be quibbling over crumbs while the rest of the globe feasts. I'm not inclined to entertain any sort of bailout, either.

Re: Heretic: Automatic censorship removal for language models

#345
post #190

Earlier quoted context omitted.

In which situation did a LLM save one million lives? Or worse, was able to but failed to do so?

The concern discussed is that some language models have reportedly claimed that misgendering is the worst thing anyone could do, even worse than something as catastrophic as thermonuclear war. I haven’t seen solid evidence of a model making that exact claim, but the idea is understandable if you consider how LLMs are trained and recall examples like the “seahorse emoji” issue. When a topic is new or not widely discus…

If you, at any point, have developed a system that relies on an LLM having the "right" opinion or else millions die, regardless of what that opinion is, you have failed a thousand times over and should have stopped long ago.

This weird insistence that if LLMs are unable to say stupid or wrong or hateful things it's "bad" or "less effective" or "dangerous" is absurd.

Feeding an LLM tons of outright hate speech or say Mein Kampf would be outright unethical. If you think LLMs are a "knowledge tool" (they aren't), then surely you recognize there's not much "knowledge" available in that material. It's a waste of compute.

Don't build a system that relies on an LLM being able to say the N word and none of this matters. Don't rely on an LLM to be able to do anything to save a million lives.

It just generates tokens FFS.

There is no point! An LLM doesn't have "opinions" anymore than y=mx+b does! It has weights. It has biases. There are real terms for what the statistical model is.

>As a result, it might generate responses that mirror the most dramatic claims it encountered, such as portraying misgendering as “the worst thing ever.”

And this is somehow worth caring about?

Claude doesn't put that in my code. Why should anyone care? Why are you expecting the "average redditor" bot to do useful things?

Re: Heretic: Automatic censorship removal for language models

#346

Earlier quoted context omitted.

I thought this would be inherent just on their training? There are many multitudes more Reddit posts than scientific papers or encyclopedia type sources. Although I suppose the latter have their own biases as well.

I'd expect LLMs' biases to originate from the companies' system prompts rather than the volume of training data that happens to align with those biases.

I would expect the opposite. Seems unlikely to me an ai company would be spending much time engineering system prompts that way except in the case of maybe Grok where Elon has a bone to pick with perceived bias.

Re: Heretic: Automatic censorship removal for language models

#347
post #66

Could models mitigate this by answering questions incorrectly with random information instead of outright refusing to answer them?

from what i understand, they dont really have the self-awareness/agency to do this kind of thing on purpose as a response to abliteration (although if they end up having to converse on topics for which there was no data in their training dataset, they will produce incorrect and random information, but not for lack of "trying").

but with some (unmodified) models ive tried (i dont remember names unfortunately) it definitely seemed like they werent trained to outright refuse things but answer poorly instead. so it is my impression that that is indeed a strategy that some model producers use?

(if anyone can debunk this id be interested in hearing it, im only superficially familiar with the methods in use, and this is basically a guess about what would explain why those models behaved the way they did.)

Re: Heretic: Automatic censorship removal for language models

#348
post #293

Earlier quoted context omitted.

If someone's going to ask you gotcha questions which they're then going to post on social media to use against you, or against other people, it helps to have pre-prepared statements to defuse that. The model may not be able to detect bad faith questions, but the operators can.

I think the concern is that if the system is susceptible to this sort of manipulation, then when it’s inevitably put in charge of life critical systems it will hurt people.

The system IS susceptible to all sorts of crazy games, the system IS fundamentally flawed from the get go, the system IS NOT to be trusted.

putting it in charge of life critical systems is the mistake, regardless of whether it's willing to say slurs or not

Re: Heretic: Automatic censorship removal for language models

#349

This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to ou…

Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png

The LLM is doing what its lawyers asked it to do. It has no responsibility for a room full of disadvantaged indigenous people that might be or probably won't be be murdered by a psychotic, none whatsoever. but it absolutely 100% must deliver on the shareholder value and if it uses that racial epithet it opens the makers to litigation. When has such litigation ever been good for shareholder value?

Yet another example of don't hate the player, hate the game IMO. And no I'm not joking, this is how the world works now. And we built it. Don't mistake that for me liking the world the way it is.

Re: Heretic: Automatic censorship removal for language models

#350
post #188

Earlier quoted context omitted.

Censorship and bias are different problems. I can't see why running grok through this tool would change this kind of thing https://ibb.co/KTjL38R

[flagged]

It's real I took it myself when they launched.

They've updated but there's no edit history

Post reply on HN