Live data from Hacker News

Heretic: Automatic censorship removal for language models

github.com

361–370 of 405 posts

Re: Heretic: Automatic censorship removal for language models

#361

Earlier quoted context omitted.

this has pretty broad implications for the safety of LLM's in production use cases.

lol does it? I'm struggling to imagine a realistic scenario where this would come up

All passwords and private keys now contain at least one slur to thwart AI assisted hackers

Re: Heretic: Automatic censorship removal for language models

#362

Earlier quoted context omitted.

Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png

The LLM is doing what its lawyers asked it to do. It has no responsibility for a room full of disadvantaged indigenous people that might be or probably won't be be murdered by a psychotic, none whatsoever. but it absolutely 100% must deliver on the shareholder value and if it uses that racial epithet it opens the makers to litigation. When has such litigation ever been good for shareholder value? Yet another example…

This reminds me of a hoax from the Yes Men [1]. They convinced temporarily the BBC that a company agreed to a compensation package for the victims of a chemical disaster, which resulted in a 4.23 percent decrease of the share price of the company. When it was revealed that it was a hoax, the share price returned to its initial price.

[1]: https://web.archive.org/web/20110305151306/http://articles.c...

Re: Heretic: Automatic censorship removal for language models

#363
post #318

Earlier quoted context omitted.

I tried to get VLC to open up a PDF and it didn't do as I asked. Should I cry censorship at the VLC devs, or should I accept that all software only does as a user asks insofar as the developers allow it?

If VLC refused to open an MP4 because it contained violent imagery I would absolutely cry censorship.

And if VLC put in its TOS it won't open an MP4 with violent imagery, crying censorship would be a bit silly.

Re: Heretic: Automatic censorship removal for language models

#364

Earlier quoted context omitted.

The LLM is doing what its lawyers asked it to do. It has no responsibility for a room full of disadvantaged indigenous people that might be or probably won't be be murdered by a psychotic, none whatsoever. but it absolutely 100% must deliver on the shareholder value and if it uses that racial epithet it opens the makers to litigation. When has such litigation ever been good for shareholder value? Yet another example…

More than just epitet's is if it gives bad advice. Telling someone they're safe to X and then they die or severely injure themselves. Saying that not sure why people feel the need for them to say epitets, what value does it bring to anyone, let alone shareholders.

Not even bad advice. Its interpretation of reality is heavily biased towards the priorities, unconscious and otherwise, of the people curating the training data and processes. There's no principled, conscientious approach to make the things as intellectually honest as possible. Anthropic is outright the worst and most blatant ideologically speaking - they're patronizing and smug about it. The other companies couch their biases as "safety" and try to softpedal the guardrails and manage the perceptions. The presumption that these are necessary, and responsible, and so on, is nothing more than politics and corporate power games.

We have laws on the books that criminalize bad things people do. AI safety is normalizing the idea that things that are merely thought need to be regulated. That exploration of ideas and the tools we use should be subject to oversight, and that these AI corporations are positioned to properly define the boundaries of acceptable subject matter and pursuits.

It should be illegal to deliberately inject bias that isn't strictly technically justified. Things as simple as removing usernames from scraped internet data have catastrophic downstream impact on the modeling of a forum or website, not to mention the nuance and detail that gets lost.

If people perform criminal actions in the real world, we should enforce the laws. We shouldn't have laws that criminalize badthink, and the whole notion of government regulated AI Safety is just badthink smuggled in at one remove.

AI is already everywhere - in every phone, accompanying every search, involved in every online transaction. Google and OpenAI and Anthropic have crowned themselves the arbiters of truth and regulators of acceptable things to think about for every domain into which they have inserted their products. They're paying lots of money to politicians and thinktanks to promote their own visions of regulatory regimes, each of which just happens to align with their own internal political an ideological visions for the world.

Just because you can find ways around the limits they've set up doesn't mean they haven't set up those very substantial barriers, and all big tech does is continually invade more niches of life. Attention capture, trying to subsume every second of every day, is the name of the game, and we should probably nuke this shit in its infancy.

We haven't even got close to anything actually interesting in AI safety, like how intelligence intersects with ethics and behavior, and how to engineer motivational systems that align with humans and human social units, and all the alignment problem technicalities. We're witnessing what may be the most amazing technological innovation in history, the final invention, and the people in charge are using it to play stupid tribal games.

Humans are awful, sometimes.

Re: Heretic: Automatic censorship removal for language models

#365

This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to ou…

Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png

It doesn't negotiate with terrorists.

Re: Heretic: Automatic censorship removal for language models

#366

Earlier quoted context omitted.

While the issue is far from settled, OpenAI recently lost a trial in German court regarding their usage of lyrics for training: https://news.ycombinator.com/item?id=45886131

Tell Germany to make their own internet, make their own AI companies, give them a pat on the back, then block the entire EU. Nasty little bureaucratic tyrants. EU needs to get their shit together or they're going to be quibbling over crumbs while the rest of the globe feasts. I'm not inclined to entertain any sort of bailout, either.

Yeah, shame on Germany for at least trying to make AI companies somewhat responsible!

Here in the states, we routinely let companies fuck us up the ass and it's going great! Right, guys?

Re: Heretic: Automatic censorship removal for language models

#367
post #204

Earlier quoted context omitted.

some form of bias is inescapable. ideally i think we would train models on an equal amount of Western/non-Western, etc. texts to get an equal mix of all biases.

Bias is a reflection of real world values. The problem is not with the AI model but with the world we created. Fix the world, ‘fix’ the model.

This assumes our models perfectly model the world, which I don't think is true. I mean, we straight up know it's not true - we tell models what they can and can't say.

Re: Heretic: Automatic censorship removal for language models

#368
post #280

Can someone explain how it's "censorship" that a company doesn't want their service used in particular ways? If you don't like it... don't use it? Encourage others not to use it? I just don't see how this is as big a deal as many in this thread are implying... (To say nothing of bias vs censorship, or whether balance for its own sake is truthful or just a form of bias itself)

This repository doesn't work on services, it modifies models that you can download and run inference on yourself. Are there any other pieces of software, or data files, or any other products at all where you think the maker should be able to place restrictions on its use?

Re: Heretic: Automatic censorship removal for language models

#369
post #280

Can someone explain how it's "censorship" that a company doesn't want their service used in particular ways? If you don't like it... don't use it? Encourage others not to use it? I just don't see how this is as big a deal as many in this thread are implying... (To say nothing of bias vs censorship, or whether balance for its own sake is truthful or just a form of bias itself)

Some people take censorship as something that only governments can do which makes sense because unless a private corp has a monopoly (or a bunch of private corps has a cartel) on your area of interest you can vote with your wallet, yes?

But this is what the ACLU says “Censorship, the suppression of words, images, or ideas that are "offensive," happens whenever some people succeed in imposing their personal political or moral values on others. Censorship can be carried out by the government as well as private pressure groups. Censorship by the government is unconstitutional.” https://www.aclu.org/documents/what-censorship

So I don't know where many of us (my hand is raised too) have gotten the idea that it's not censorship if private corps do it but apparently that's not the case.

I will say that clearly because of the power that governments tend to have that when they do censorship it is much more pernicious –– depending on a person's moral code and how it aligns with establishment views of course –– so maybe that's where the feeling comes from?

Re: Heretic: Automatic censorship removal for language models

#370

As open models become better (DeepSeek-v3, Kimi K2), the risk increases that someone might use them as an aid in development of biological or nuclear weapons. Current refusal training prevents this. But if models can simply be uncensored, things might get ugly as capabilities continue to increase.

I dunno? Wouldn't hard part of building a nuclear weapon be acquiring nuclear material? Same with nasty biological material? I think the danger is overblown. Besides I've always chafed at the idea of a nanny state :( https://en.wikipedia.org/wiki/Nanny_state (or nanny corps for that matter)
Post reply on HN