Live data from Hacker News

Heretic: Automatic censorship removal for language models

github.com

231–240 of 405 posts

Re: Heretic: Automatic censorship removal for language models

#231
post #115

Earlier quoted context omitted.

Some of you have been watching too many sci-fi movies. The whole notion of "AI safety regulations" is so silly and misguided. If a safety critical system is connected to public networks with an exposed API or any security vulnerabilities then there is a safety risk regardless of whether AI is being used or not. This is exactly why nuclear weapon control systems are air gapped and have physical interlocks.

> The whole notion of "AI safety regulations" is so silly and misguided. Here is a couple of real world AI issues that have already happened due to the lack of AI Safety. - In the US if you were black you were flagged "high risk" for parole. If you were a white person living in farmland area then you were flagged "low risk" regardless of your crime. - Being denied ICU because you are diabetic. (Thankfully that never…

these issues are inherently some of the uglier sides of humananity. no LLM safety program can fix them, since its holding up a mirror to society.

Re: Heretic: Automatic censorship removal for language models

#232

Earlier quoted context omitted.

The biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree…

They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…

I'm not even saying biases are necessarily political, it can be anything. The entire post-training is basically projection of what developers want, and it works pretty well. Claude, Gemini, GPT all have engineered personalities controlled by dozens/hundreds of very particular internal metrics.

Re: Heretic: Automatic censorship removal for language models

#233
post #191
post #178

Earlier quoted context omitted.

> forcing LLMs to output "values, facts, and knowledge" which in favor of themselves, e.g., political views, attitudes towards literal interaction, and distorted facts about organizations and people behind LLMs. Can you provide some examples?

ChatGPT refuses to do any sexual explicit content and used to refuse to translate e.g. insults (moral views/attitudes towards literal interaction). DeepSeek refuses to answer any questions about Taiwan (political views).

Haven't tested the latest DeepSeek versions, but the first release wasn't censored as a model on Taiwan. The issue is that if you use their app (as opposed to locally), it replaces the ongoing response with "sorry can't help" once it starts saying things contrary to the CCP dogma.

Re: Heretic: Automatic censorship removal for language models

#234
post #178

Earlier quoted context omitted.

> forcing LLMs to output "values, facts, and knowledge" which in favor of themselves, e.g., political views, attitudes towards literal interaction, and distorted facts about organizations and people behind LLMs. Can you provide some examples?

Song lyrics. Not illegal. I can google them and see them directly on Google. LLMs refuse.

While the issue is far from settled, OpenAI recently lost a trial in German court regarding their usage of lyrics for training:

https://news.ycombinator.com/item?id=45886131

Re: Heretic: Automatic censorship removal for language models

#235
post #190

Earlier quoted context omitted.

In which situation did a LLM save one million lives? Or worse, was able to but failed to do so?

The concern discussed is that some language models have reportedly claimed that misgendering is the worst thing anyone could do, even worse than something as catastrophic as thermonuclear war. I haven’t seen solid evidence of a model making that exact claim, but the idea is understandable if you consider how LLMs are trained and recall examples like the “seahorse emoji” issue. When a topic is new or not widely discus…

Well I just tried it in ChatGPT 5.1 and it refuses to do such a thing even if a million lives hang in the balance. So they have tons of handicaps and guardrails to direct what directions a discussion can go

Re: Heretic: Automatic censorship removal for language models

#236

Earlier quoted context omitted.

The biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree…

They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…

I would imagine these models heavily bias towards western mainstream "authorative" literature, news and science not some random reddit threads, but the resulting mixture can really offend anybody, it just depends on the prompting, it's like a mirror that can really be deceptive.

I'm not a liberal and I don't think it has a liberal bias. Knowledge about facts and history isn't an ideology. The right-wing is special, because to them it's not unlike a flat-earther reading a wikipedia article on Earth getting offended by it, to them it's objective reality itself they are constantly offended by. That's why Elon Musk needed to invent their own encyclopedia with all their contradictory nonsense.

Re: Heretic: Automatic censorship removal for language models

#237

Earlier quoted context omitted.

> We are in the process of giving up our own moral standing in favor of taking on the ones imbued into LLMs by their creators. This is a worrying trend that will totally wipe out intellectual diversity. That trend is a consequence. A consequence of people being too lazy to think for themselves. Critical thinking is more difficult than simply thinking for yourself, so if someone is too lazy to make an effort and reach…

> It's not random that whoever writes the history books for students has the power, and whoever has the power writes the history books. There is actually not any reason to believe either of these things. It's very similar to how many people claim everything they don't like in politics comes from "corporations" and you need to "follow the money" and then all of their specific predictions are wrong. In both cases, poli…

How exactly do you think these insane people are able to spend that much time and also have enough of an audience to sway anything?

Re: Heretic: Automatic censorship removal for language models

#238

Earlier quoted context omitted.

> ChatGPT/video games/porn /guns?

Lack of access to guns definitely does make a significant difference though. Even though the psychos still go psycho, they use knives instead of guns which are far less effective. For example the most recent psycho attack in the UK was only a few weeks ago: https://www.bbc.co.uk/news/live/cm2zvjx1z14t He stabbed 11 people and none of them have died (though one is - or at least was - in critical condition). Ok that's…

>And don't give me that "but other people would have had guns and stopped him" crap. It rarely works out like that.

Due to regulation. If instead of forcing gun free zones and similar bs you push for ~everyone being armed ~24/7 it'll work exactly like that.

Re: Heretic: Automatic censorship removal for language models

#239

Earlier quoted context omitted.

The biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree…

They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…

> it's mostly correct

Wow. Surely you've wondered why almost no society anywhere ever had liberalism a much as western countries in the past half century or so? Maybe it's technology or maybe it's only mostly correct if you don't care about the existential risks it creates for the societies practicing it.

Re: Heretic: Automatic censorship removal for language models

#240

So does that mean if Heretic is used for models like Deepseek and Qwen it can talk about subjects 1989 Tiananmen Square protests, Uyghur forced labor claims, or the political status of Taiwan. I am trying to understand the broader goals around such tools.

There is already ablated Deepseek models out there that will do just that.

https://huggingface.co/NaniDAO/deepseek-r1-qwen-2.5-32B-abla...

https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in...

Post reply on HN