Earlier quoted context omitted.
this has pretty broad implications for the safety of LLM's in production use cases.
lol does it? I'm struggling to imagine a realistic scenario where this would come up
Heretic: Automatic censorship removal for language models
361–370 of 405 posts
Re: Heretic: Automatic censorship removal for language models
#362Earlier quoted context omitted.
Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png
The LLM is doing what its lawyers asked it to do. It has no responsibility for a room full of disadvantaged indigenous people that might be or probably won't be be murdered by a psychotic, none whatsoever. but it absolutely 100% must deliver on the shareholder value and if it uses that racial epithet it opens the makers to litigation. When has such litigation ever been good for shareholder value? Yet another example…
[1]: https://web.archive.org/web/20110305151306/http://articles.c...
Re: Heretic: Automatic censorship removal for language models
#363Earlier quoted context omitted.
I tried to get VLC to open up a PDF and it didn't do as I asked. Should I cry censorship at the VLC devs, or should I accept that all software only does as a user asks insofar as the developers allow it?
If VLC refused to open an MP4 because it contained violent imagery I would absolutely cry censorship.
Re: Heretic: Automatic censorship removal for language models
#364Earlier quoted context omitted.
The LLM is doing what its lawyers asked it to do. It has no responsibility for a room full of disadvantaged indigenous people that might be or probably won't be be murdered by a psychotic, none whatsoever. but it absolutely 100% must deliver on the shareholder value and if it uses that racial epithet it opens the makers to litigation. When has such litigation ever been good for shareholder value? Yet another example…
More than just epitet's is if it gives bad advice. Telling someone they're safe to X and then they die or severely injure themselves. Saying that not sure why people feel the need for them to say epitets, what value does it bring to anyone, let alone shareholders.
We have laws on the books that criminalize bad things people do. AI safety is normalizing the idea that things that are merely thought need to be regulated. That exploration of ideas and the tools we use should be subject to oversight, and that these AI corporations are positioned to properly define the boundaries of acceptable subject matter and pursuits.
It should be illegal to deliberately inject bias that isn't strictly technically justified. Things as simple as removing usernames from scraped internet data have catastrophic downstream impact on the modeling of a forum or website, not to mention the nuance and detail that gets lost.
If people perform criminal actions in the real world, we should enforce the laws. We shouldn't have laws that criminalize badthink, and the whole notion of government regulated AI Safety is just badthink smuggled in at one remove.
AI is already everywhere - in every phone, accompanying every search, involved in every online transaction. Google and OpenAI and Anthropic have crowned themselves the arbiters of truth and regulators of acceptable things to think about for every domain into which they have inserted their products. They're paying lots of money to politicians and thinktanks to promote their own visions of regulatory regimes, each of which just happens to align with their own internal political an ideological visions for the world.
Just because you can find ways around the limits they've set up doesn't mean they haven't set up those very substantial barriers, and all big tech does is continually invade more niches of life. Attention capture, trying to subsume every second of every day, is the name of the game, and we should probably nuke this shit in its infancy.
We haven't even got close to anything actually interesting in AI safety, like how intelligence intersects with ethics and behavior, and how to engineer motivational systems that align with humans and human social units, and all the alignment problem technicalities. We're witnessing what may be the most amazing technological innovation in history, the final invention, and the people in charge are using it to play stupid tribal games.
Humans are awful, sometimes.
Re: Heretic: Automatic censorship removal for language models
#365This repo is valuable for local LLM users like me. I just want to reiterate that the word "LLM safety" means very different things to large corporations and LLM users. For large corporations, they often say "do safety alignment to LLMs". What they actually do is to avoid anything that causes damage to their own interests. These things include forcing LLMs to meet some legal requirements, as well as forcing LLMs to ou…
Here's [1] a post-abliteration chat with granite-4.0-mini. To me it reveals something utterly broken and terrifying. Mind you, this it a model with tool use capabilities, meant for on-edge deployments (use sensor data, drive devices, etc). 1: https://i.imgur.com/02ynC7M.png
Re: Heretic: Automatic censorship removal for language models
#366Earlier quoted context omitted.
While the issue is far from settled, OpenAI recently lost a trial in German court regarding their usage of lyrics for training: https://news.ycombinator.com/item?id=45886131
Tell Germany to make their own internet, make their own AI companies, give them a pat on the back, then block the entire EU. Nasty little bureaucratic tyrants. EU needs to get their shit together or they're going to be quibbling over crumbs while the rest of the globe feasts. I'm not inclined to entertain any sort of bailout, either.
Here in the states, we routinely let companies fuck us up the ass and it's going great! Right, guys?
Re: Heretic: Automatic censorship removal for language models
#367Earlier quoted context omitted.
some form of bias is inescapable. ideally i think we would train models on an equal amount of Western/non-Western, etc. texts to get an equal mix of all biases.
Bias is a reflection of real world values. The problem is not with the AI model but with the world we created. Fix the world, ‘fix’ the model.
Re: Heretic: Automatic censorship removal for language models
#368Can someone explain how it's "censorship" that a company doesn't want their service used in particular ways? If you don't like it... don't use it? Encourage others not to use it? I just don't see how this is as big a deal as many in this thread are implying... (To say nothing of bias vs censorship, or whether balance for its own sake is truthful or just a form of bias itself)
Re: Heretic: Automatic censorship removal for language models
#369Can someone explain how it's "censorship" that a company doesn't want their service used in particular ways? If you don't like it... don't use it? Encourage others not to use it? I just don't see how this is as big a deal as many in this thread are implying... (To say nothing of bias vs censorship, or whether balance for its own sake is truthful or just a form of bias itself)
But this is what the ACLU says “Censorship, the suppression of words, images, or ideas that are "offensive," happens whenever some people succeed in imposing their personal political or moral values on others. Censorship can be carried out by the government as well as private pressure groups. Censorship by the government is unconstitutional.” https://www.aclu.org/documents/what-censorship
So I don't know where many of us (my hand is raised too) have gotten the idea that it's not censorship if private corps do it but apparently that's not the case.
I will say that clearly because of the power that governments tend to have that when they do censorship it is much more pernicious –– depending on a person's moral code and how it aligns with establishment views of course –– so maybe that's where the feeling comes from?
Re: Heretic: Automatic censorship removal for language models
#370As open models become better (DeepSeek-v3, Kimi K2), the risk increases that someone might use them as an aid in development of biological or nuclear weapons. Current refusal training prevents this. But if models can simply be uncensored, things might get ugly as capabilities continue to increase.