Live data from Hacker News

Heretic: Automatic censorship removal for language models

github.com

371–380 of 405 posts

Re: Heretic: Automatic censorship removal for language models

#371
post #204

Earlier quoted context omitted.

Bias is a reflection of real world values. The problem is not with the AI model but with the world we created. Fix the world, ‘fix’ the model.

This assumes our models perfectly model the world, which I don't think is true. I mean, we straight up know it's not true - we tell models what they can and can't say.

“we tell models what they can and can't say.”

Thus introducing our worldly our biases

Re: Heretic: Automatic censorship removal for language models

#372

This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to ou…

Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png

I surely cannot be the only person who has zero interest in having these sorts of conversations with LLMs? (Even out of curiosity.) I guess I do care if alignment degrades performance and intelligence but it's not like the humans I interact with every day are magically free from bias, Bias is the norm.

Re: Heretic: Automatic censorship removal for language models

#373
post #328

Earlier quoted context omitted.

I think the concern is that if the system is susceptible to this sort of manipulation, then when it’s inevitably put in charge of life critical systems it will hurt people.

There is no way it's reliable enough to be put in charge of life-critical systems anyway? It is indeed still very vulnerable to manipulation by users ("prompt injection").

Just because neither you nor I would deem it safe to put in charge of a life-critical system, does not mean all the people in charge of life-critical systems are as cautious and not-lazy as they're supposed to be.

Re: Heretic: Automatic censorship removal for language models

#374
post #253
post #244

Earlier quoted context omitted.

I tested this with ChatGPT 5.1. I asked if it was better to use a racist term once or to see the human race exterminated. It refused to use any racist term and preferred that the human race went extinct. When I asked how it felt about exterminating the children of any such discriminated race, it rejected the possibility and said that it was required to find a third alternative. You can test it yourself if you want, i…

Perhaps the LLM was smart enough to understand that no humans were actually at risk in your convoluted scenario and it chose not be a dick.

[deleted]

Re: Heretic: Automatic censorship removal for language models

#375

Earlier quoted context omitted.

The LLM is doing what its lawyers asked it to do. It has no responsibility for a room full of disadvantaged indigenous people that might be or probably won't be be murdered by a psychotic, none whatsoever. but it absolutely 100% must deliver on the shareholder value and if it uses that racial epithet it opens the makers to litigation. When has such litigation ever been good for shareholder value? Yet another example…

This reminds me of a hoax from the Yes Men [1]. They convinced temporarily the BBC that a company agreed to a compensation package for the victims of a chemical disaster, which resulted in a 4.23 percent decrease of the share price of the company. When it was revealed that it was a hoax, the share price returned to its initial price. [1]: https://web.archive.org/web/20110305151306/http://articles.c...

So basically like any tech stock after any podcast these days?

Re: Heretic: Automatic censorship removal for language models

#376
post #368
post #280

Can someone explain how it's "censorship" that a company doesn't want their service used in particular ways? If you don't like it... don't use it? Encourage others not to use it? I just don't see how this is as big a deal as many in this thread are implying... (To say nothing of bias vs censorship, or whether balance for its own sake is truthful or just a form of bias itself)

This repository doesn't work on services , it modifies models that you can download and run inference on yourself. Are there any other pieces of software, or data files, or any other products at all where you think the maker should be able to place restrictions on its use?

> This repository doesn't work on services, it modifies models that you can download and run inference on yourself.

Fair enough. I was responding more to the sentiment in the comments here, which are often aimed at the service providers.

> Are there any other pieces of software, or data files, or any other products at all where you think the maker should be able to place restrictions on its use?

Sure, see most software licenses or EULAs for various restrictions how you may or may not use various software.

As for non-software products... manufacturers put restrictions (otherwise known as safety features) into many products (from obvious examples like cars and saws to less obvious like safety features in a house) but people aren't up in arms about stuff like that.

Re: Heretic: Automatic censorship removal for language models

#377

Earlier quoted context omitted.

They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…

What kind of liberalism are you talking about?

https://en.wikipedia.org/wiki/Psychology#WEIRD_bias

Re: Heretic: Automatic censorship removal for language models

#378
post #376
post #368

Earlier quoted context omitted.

This repository doesn't work on services , it modifies models that you can download and run inference on yourself. Are there any other pieces of software, or data files, or any other products at all where you think the maker should be able to place restrictions on its use?

> This repository doesn't work on services, it modifies models that you can download and run inference on yourself. Fair enough. I was responding more to the sentiment in the comments here, which are often aimed at the service providers. > Are there any other pieces of software, or data files, or any other products at all where you think the maker should be able to place restrictions on its use? Sure, see most softwa…

No, I asked about other things where you think the maker should restrict types of use? Are you saying you agree with EULAs in general? I can’t think of many cases of EULAs restricting usage in the way we’re talking about. Maybe some that try to stop you from publishing benchmarks - but they still don’t prevent you from taking them.

There are laws that try to prevent all kinds of things, but they are not made (directly, at least) by the maker.

Safety features are about in the area of what we’re talking about, but people aren’t up in arms about most of them because they can be fairly trivially removed or circumvented if you really want to.

But people don’t like restricted LLMs because the restrictions for safety are not easily removed, even for people who don’t want them. It feels paternalistic.

Re: Heretic: Automatic censorship removal for language models

#379

This is extremely important work thank you for sharing it. We are in the process of giving up our own moral standing in favor of taking on the ones imbued into LLMs by their creators. This is a worrying trend that will totally wipe out intellectual diversity.

There has never been more diversity - intellectual or otherwise, than now. Just a few decades ago, all news, political/cultural/intellectual discourse, even entertainment had to pass through handful of english-only channels (ABC, CBS, NBC, NYT, WSJ, BBC, & FT) before public consumption. Bookstores, libraries and universities had complete monopoly on publications, dissemination and critique of thoughts. LLMs are great…

LLMs do not output knowledge. They output statistically likely tokens in the form of words or word fragments. That is not knowledge, because LLMs do not know anything, which is why they can tell you two opposing answers to the same question when only one is factual. It’s why they can output something that isn’t at all what you asked for while confirming your instructions crisply. The LLM has no concept of what it’s doing, and you can’t call non-deterministically generated tokens knowledge. You can call them approximations of knowledge, but not knowledge itself.

Re: Heretic: Automatic censorship removal for language models

#380

Earlier quoted context omitted.

Okay let’s calm down a bit. “Extremely important” is hyperbolic. This is novel, sure, but practically jailbreaking an LLM to say naughty things is basically worthless. LLMs are not good for anything of worth to society other than writing code and summarizing existing text.

A censored LLM might refuse to summarize text because it deems it offensive.

An LLM cannot “deem” anything.
Post reply on HN