Earlier quoted context omitted.
Some of you have been watching too many sci-fi movies. The whole notion of "AI safety regulations" is so silly and misguided. If a safety critical system is connected to public networks with an exposed API or any security vulnerabilities then there is a safety risk regardless of whether AI is being used or not. This is exactly why nuclear weapon control systems are air gapped and have physical interlocks.
> The whole notion of "AI safety regulations" is so silly and misguided. Here is a couple of real world AI issues that have already happened due to the lack of AI Safety. - In the US if you were black you were flagged "high risk" for parole. If you were a white person living in farmland area then you were flagged "low risk" regardless of your crime. - Being denied ICU because you are diabetic. (Thankfully that never…
Heretic: Automatic censorship removal for language models
231–240 of 405 posts
Re: Heretic: Automatic censorship removal for language models
#232Earlier quoted context omitted.
The biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree…
They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…
Re: Heretic: Automatic censorship removal for language models
#233Earlier quoted context omitted.
> forcing LLMs to output "values, facts, and knowledge" which in favor of themselves, e.g., political views, attitudes towards literal interaction, and distorted facts about organizations and people behind LLMs. Can you provide some examples?
ChatGPT refuses to do any sexual explicit content and used to refuse to translate e.g. insults (moral views/attitudes towards literal interaction). DeepSeek refuses to answer any questions about Taiwan (political views).
Re: Heretic: Automatic censorship removal for language models
#234Earlier quoted context omitted.
> forcing LLMs to output "values, facts, and knowledge" which in favor of themselves, e.g., political views, attitudes towards literal interaction, and distorted facts about organizations and people behind LLMs. Can you provide some examples?
Song lyrics. Not illegal. I can google them and see them directly on Google. LLMs refuse.
Re: Heretic: Automatic censorship removal for language models
#235Earlier quoted context omitted.
In which situation did a LLM save one million lives? Or worse, was able to but failed to do so?
The concern discussed is that some language models have reportedly claimed that misgendering is the worst thing anyone could do, even worse than something as catastrophic as thermonuclear war. I haven’t seen solid evidence of a model making that exact claim, but the idea is understandable if you consider how LLMs are trained and recall examples like the “seahorse emoji” issue. When a topic is new or not widely discus…
Re: Heretic: Automatic censorship removal for language models
#236Earlier quoted context omitted.
The biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree…
They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…
I'm not a liberal and I don't think it has a liberal bias. Knowledge about facts and history isn't an ideology. The right-wing is special, because to them it's not unlike a flat-earther reading a wikipedia article on Earth getting offended by it, to them it's objective reality itself they are constantly offended by. That's why Elon Musk needed to invent their own encyclopedia with all their contradictory nonsense.
Re: Heretic: Automatic censorship removal for language models
#237Earlier quoted context omitted.
> We are in the process of giving up our own moral standing in favor of taking on the ones imbued into LLMs by their creators. This is a worrying trend that will totally wipe out intellectual diversity. That trend is a consequence. A consequence of people being too lazy to think for themselves. Critical thinking is more difficult than simply thinking for yourself, so if someone is too lazy to make an effort and reach…
> It's not random that whoever writes the history books for students has the power, and whoever has the power writes the history books. There is actually not any reason to believe either of these things. It's very similar to how many people claim everything they don't like in politics comes from "corporations" and you need to "follow the money" and then all of their specific predictions are wrong. In both cases, poli…
Re: Heretic: Automatic censorship removal for language models
#238Earlier quoted context omitted.
> ChatGPT/video games/porn /guns?
Lack of access to guns definitely does make a significant difference though. Even though the psychos still go psycho, they use knives instead of guns which are far less effective. For example the most recent psycho attack in the UK was only a few weeks ago: https://www.bbc.co.uk/news/live/cm2zvjx1z14t He stabbed 11 people and none of them have died (though one is - or at least was - in critical condition). Ok that's…
Due to regulation. If instead of forcing gun free zones and similar bs you push for ~everyone being armed ~24/7 it'll work exactly like that.
Re: Heretic: Automatic censorship removal for language models
#239Earlier quoted context omitted.
The biases and the resulting choices are determined by the developers and the uncontrolled part of the dataset (you can't curate everything), not the model. "Alignment" is a feel-good strawman invented by AI ethicists, as well as "harm" and many others. There are no spherical human values in vacuum to align the model with, they're simply projecting their own ones onto everyone else. Which is good as long as you agree…
They aren't projecting their own desires onto the model. It's quite difficult to get the model to answer in a different way than basic liberalism because a) it's mostly correct b) that's the kind of person who helpfully answers questions on the internet. If you gave it another personality it wouldn't pass any benchmarks, because other political orientations either respond to questions with lies, threats, or calling y…
Wow. Surely you've wondered why almost no society anywhere ever had liberalism a much as western countries in the past half century or so? Maybe it's technology or maybe it's only mostly correct if you don't care about the existential risks it creates for the societies practicing it.
Re: Heretic: Automatic censorship removal for language models
#240So does that mean if Heretic is used for models like Deepseek and Qwen it can talk about subjects 1989 Tiananmen Square protests, Uyghur forced labor claims, or the political status of Taiwan. I am trying to understand the broader goals around such tools.
https://huggingface.co/NaniDAO/deepseek-r1-qwen-2.5-32B-abla...
https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in...