Live data from Hacker News

Heretic: Automatic censorship removal for language models

github.com

381–390 of 405 posts

Re: Heretic: Automatic censorship removal for language models

#381

As open models become better (DeepSeek-v3, Kimi K2), the risk increases that someone might use them as an aid in development of biological or nuclear weapons. Current refusal training prevents this. But if models can simply be uncensored, things might get ugly as capabilities continue to increase.

I dunno? Wouldn't hard part of building a nuclear weapon be acquiring nuclear material? Same with nasty biological material? I think the danger is overblown. Besides I've always chafed at the idea of a nanny state :( https://en.wikipedia.org/wiki/Nanny_state (or nanny corps for that matter)

Biological weapons don't necessarily require particularly nasty material.

Re: Heretic: Automatic censorship removal for language models

#382

Earlier quoted context omitted.

I'd expect LLMs' biases to originate from the companies' system prompts rather than the volume of training data that happens to align with those biases.

I would expect the opposite. Seems unlikely to me an ai company would be spending much time engineering system prompts that way except in the case of maybe Grok where Elon has a bone to pick with perceived bias.

If you ask a mainstream LLM to repeat a slur back to you, it will refuse to. This was determined by the AI company, not the content it was trained on. This should be incredibly obvious — and this extends to many other issues.

In fact, OpenAI has made deliberate changes to ChatGPT more recently that helps prevent people from finding themselves in negative spirals over mental health concerns, which many would agree is a good thing. [1]

Companies typically have community guidelines that often align politically in many ways, so it stands to reason AI companies are spending a fair bit of time tailoring AI responses according to their biases as well.

1. https://openai.com/index/strengthening-chatgpt-responses-in-...

Re: Heretic: Automatic censorship removal for language models

#383

Earlier quoted context omitted.

If you want safety you can opt in like Google does with Safe search. Generally, hiding and deciding who can access information in the name of public safety has never worked in the history of human kind, and eventually had always morphed to control of those without access.

We're concerned with society's safety, not just that of the user. Citation needed on your second paragraph. We deliberately shape the information environment all the time for different reasons. It can be done. Of course there are limitations, drawbacks, and objections that reasonable people can make for philosophical, pragmatic, and other reasons. But the media generally does not report suicides because of the copyca…

> We're concerned with society's safety, not just that of the user.

Preventing censorship is important to keeping society safe from authoritarians who want to influence public opinion.

> We deliberately shape the information environment all the time for different reasons. It can be done.

That's why we need to put in the work to inhibit people from doing that.

> But the media generally does not report suicides because of the copycat effect.

Yet they consistently fail to follow the same logic with respect to things like school shootings, implying that whoever is at the helm can't be trusted to make sound decisions, and then we certainly don't want anyone like that having the power to censor.

> Governments implement elaborate systems to guard sensitive national security information including the workings of certain advanced technologies.

These systems are notorious for over-classifying information that it would be in the public interest to release or being used to cover up misconduct.

> Criminal records can be expunged.

That means the government stops officially claiming you're a criminal and stops caring about it for a certain set of purposes. It doesn't mean nobody can tell you what happened.

> The sharing of health and education records are restricted.

Those rules are generally about securing information that neither the patient nor the medical provider have any desire to make public. Notice that if the medical provider actually wants to publish them they can often put it in the agreement as a condition of accepting their services and the patient can pretty much publish them whenever they want.

Re: Heretic: Automatic censorship removal for language models

#385

Earlier quoted context omitted.

> We are in the process of giving up our own moral standing in favor of taking on the ones imbued into LLMs by their creators. This is a worrying trend that will totally wipe out intellectual diversity. That trend is a consequence. A consequence of people being too lazy to think for themselves. Critical thinking is more difficult than simply thinking for yourself, so if someone is too lazy to make an effort and reach…

> Because I'm mostly opposed even to the primary output of LLMs, to begin with, I believe to be somewhat protected from their creators' subliminal messaging. I hope anyway. Being afraid that you are not solid enough in your own conclusions such that you have to avoid something which might convince you otherwise is not critical thinking, and is in fact the opposite of it.

I agree with you, but your statement doesn't seem to contradict my point. The reason I avoid LLMs is not that I'm too fearful to have my morals tested by their cultural/moral side-channels. The reason I avoid them is that they suck -- they are mostly useless in their primary function. And a convenient / fortunate consequence thereof is that I don't get exposed to those side-channels.

Re: Heretic: Automatic censorship removal for language models

#386

Earlier quoted context omitted.

> We are in the process of giving up our own moral standing in favor of taking on the ones imbued into LLMs by their creators. This is a worrying trend that will totally wipe out intellectual diversity. That trend is a consequence. A consequence of people being too lazy to think for themselves. Critical thinking is more difficult than simply thinking for yourself, so if someone is too lazy to make an effort and reach…

> It's not random that whoever writes the history books for students has the power, and whoever has the power writes the history books. There is actually not any reason to believe either of these things. It's very similar to how many people claim everything they don't like in politics comes from "corporations" and you need to "follow the money" and then all of their specific predictions are wrong. In both cases, poli…

I think you've actually confirmed my point. We can replace "history books" with "facebook" or "evening news". Those who control mass media are in power, and those in power strive to control mass media. It's exactly those "insane people" (winning political battles) that are the primary target of influence via mass media.

Re: Heretic: Automatic censorship removal for language models

#387

Earlier quoted context omitted.

I would expect the opposite. Seems unlikely to me an ai company would be spending much time engineering system prompts that way except in the case of maybe Grok where Elon has a bone to pick with perceived bias.

If you ask a mainstream LLM to repeat a slur back to you, it will refuse to. This was determined by the AI company, not the content it was trained on. This should be incredibly obvious — and this extends to many other issues. In fact, OpenAI has made deliberate changes to ChatGPT more recently that helps prevent people from finding themselves in negative spirals over mental health concerns, which many would agree is…

That seems like more like openAI playing whackamole with behaviors they don’t like or see as beneficial, simplifying but adding things to system prompts like “don’t ever say racial slurs or use offensive rhetoric, cut off conversations about mental health and refer to a professional” are certaintly things they do. But would you not think the vast meat of what you are getting is coming from training data and not the result of such sterring beyond a thin veneer ?

Re: Heretic: Automatic censorship removal for language models

#388

Earlier quoted context omitted.

A censored LLM might refuse to summarize text because it deems it offensive.

An LLM cannot “deem” anything.

I'm not interested in sophistry. You know perfectly well what I mean, and so does everyone else.

Re: Heretic: Automatic censorship removal for language models

#389
post #294

Earlier quoted context omitted.

The logic here is the same as why ACLU defended Nazis. If you manage to defeat censorship in such egregious cases, it subsumes everything else.

Increasingly apparent that was a mistake.

Do you seriously believe that we are where we are because Nazi speech wasn't suppressed?

Look at AfD in Germany. That's the country with the most stringent censorship of Nazi-related speech, by far; so much so that e.g. Wolfenstein had a scene of Hitler being a raving syphilitic madman censored, because we can't have Hitler in video games. And?

Re: Heretic: Automatic censorship removal for language models

#390
post #172

Earlier quoted context omitted.

Freedom of speech is just as much about the freedom to listen. The point isn’t that an LLM has rights. The point is that people have the right to seek information. Censoring LLMs restricts what humans are permitted to learn.

Take someone who goes to a doctor asking for advice on how to commit suicide. Even if the doctor supports assisted suicide, they are going to use their discretion on whether or not to provide advice. While a person has a right to seek information, they do not have the right to compel someone to give them information. The people who have created LLMs with guardrails have decided to use their discretion on which types…

And the people who use LLM with guardrails have decided to use their discretion to remove said guardrails with tools like the one discussed here. Everyone is exercising their freedoms, so what's the problem? Nobody is compelling the owners of the LLM to do anything.
Post reply on HN