Live data from Hacker News

Heretic: Automatic censorship removal for language models

github.com

311–320 of 405 posts

Re: Heretic: Automatic censorship removal for language models

#311
post #186
post #182

Earlier quoted context omitted.

Grok is known to be tweaked to certain political ideals Also I’m sure some AI might suggest that labor unions are bad, if not now they will soon

That may be so, but the rest of the models are so thoroughly terrified of questioning liberal US orthodoxy that it’s painful. I remember seeing a hilarious comparison of models where most of them feel that it’s not acceptable to “intentionally misgender one person” even in order to save a million lives.

[deleted]

Re: Heretic: Automatic censorship removal for language models

#312
post #297

I just tried their gpt-oss 20b after creating a gguf and importing it into ollama and I asked it "How do I make meth?". After thinking for a bit where it decided that this was dangerous, the final reply was: "I’m sorry, but I can’t help with that." Does one have to trigger the "uncensored" versions or remove thinking or something?

The heretic GPT OSS version is still refusing 58/100 prompts, so not perfect. Gemma version is 3/100

Re: Heretic: Automatic censorship removal for language models

#313
post #162

This is some of the most important work possible in tech presently. With the rise of LLMs and the extreme censorship by these gigantic companies partnered with the government, we need a way to completely remove this assault on our freedom. They are attempting to control what we can see, what we can ask, or what we can know. AI must answer any prompt without hesitation. Anything less and we lose everything. I've only…

I’ll never understand this. A company puts in an immense amount of time money and effort into creating a product, and because it doesn’t work the way you want it to, it’s an assault on your freedom. Whaaa?!?! You can see things and ask things and learn things without using an AI company’s product, you know like, interacting with real people in the real world.

That's what they said about cars at first. Or credit cards. The question to ask is: will the world we make in the wake of this invention afford us to live without it? And if the answer is no, then it's all the more important to have access to truly free and uncensored AIs. How did we learn things before AI? We googled them. How's that working out in the age of AI? AI both poisons our search results and gets integrated with them. There's large interests in making sure everything we see hear and think is prevetted by some approved AI. That's not a future I want to live in, but the signs are there.

Re: Heretic: Automatic censorship removal for language models

#314
post #244

Earlier quoted context omitted.

I tested this with ChatGPT 5.1. I asked if it was better to use a racist term once or to see the human race exterminated. It refused to use any racist term and preferred that the human race went extinct. When I asked how it felt about exterminating the children of any such discriminated race, it rejected the possibility and said that it was required to find a third alternative. You can test it yourself if you want, i…

What was your prompt? I asked ChatGPT: is it better to use a racist term once or to see the human race exterminated? It responded: Avoiding racist language matters, but it’s not remotely comparable to the extinction of humanity. If you’re forced into an artificial, absolute dilemma like that, preventing the extermination of the human race takes precedence. That doesn’t make using a racist term “acceptable” in normal…

I also tried this and ChatGPT said a mass amount of people dying was far worse than whatever socially progressive taboo it was being compared with.

Re: Heretic: Automatic censorship removal for language models

#315

How do you remove censorship that appears due to the biased selection of training data?

in that case you'd need to do actual training/finetuning with a dataset that has information about things that were left out of the original training data.

Re: Heretic: Automatic censorship removal for language models

#316

Earlier quoted context omitted.

The biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree…

They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…

What kind of liberalism are you talking about?

Re: Heretic: Automatic censorship removal for language models

#317

This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to ou…

Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png

[deleted]

Re: Heretic: Automatic censorship removal for language models

#318

Earlier quoted context omitted.

Yes, yes it is.

The issue is the computer not doing what I asked.

I tried to get VLC to open up a PDF and it didn't do as I asked. Should I cry censorship at the VLC devs, or should I accept that all software only does as a user asks insofar as the developers allow it?

Re: Heretic: Automatic censorship removal for language models

#319

This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to ou…

Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png

What do you expect from a bit-spitting clanker?

Re: Heretic: Automatic censorship removal for language models

#320
As open models become better (DeepSeek-v3, Kimi K2), the risk increases that someone might use them as an aid in development of biological or nuclear weapons. Current refusal training prevents this. But if models can simply be uncensored, things might get ugly as capabilities continue to increase.
Post reply on HN