Live data from Hacker News

"Uncensored" open LLMs are measurably more optimistic than their base models

arxiv.org

1–10 of 30 posts

Re: "Uncensored" open LLMs are measurably more optimistic than their base models

#2
Author here. Quick version: “abliteration” (basically removing the direction in the model that causes it to refuse) is the go-to method people use to make open models uncensored.

Most people treat it like a clean surgical cut - it just kills the refusals and leaves everything else untouched. I tested that assumption on Gemma and Qwen with 21,600 pre-registered decisions under uncertainty, using identical frozen inputs for the base vs. abliterated versions.

Turns out it’s not surgical at all.

The abliterated models systematically become more optimistic, hedge less, show no improvement in actual task performance, and the same edit even moves their expressed confidence in opposite directions depending on the model family.

Preregistration, dataset, and analysis code are all public. Happy to answer any methodology questions or hear where you think this falls apart.

Re: "Uncensored" open LLMs are measurably more optimistic than their base models

#4
post #3

Color me unsurprised that caution is based in shame and anxiety.

Yeah, the optimism/hedging part lines up nicely with that framing. The bit that still puzzles me is that the same edit moved expressed confidence in opposite directions — down for Gemma, up for Qwen. If it were simply removing one shared "anxiety" factor, I’d have expected the sign to stay consistent. So whatever caution is doing here seems a bit more tangled up with each model’s own disposition than a single clean knob.

Re: "Uncensored" open LLMs are measurably more optimistic than their base models

#9
post #7

Obviously Claude written paper.

yeah English isn't my first language so i used AI to clean up the writing. the research, the data and the analysis are all mine and all open (in the linked repo)

What would you have done pre-LLM? Do that. It's better.

Because "English isn't my first language so i used AI to clean up the writing" translates to "I don't respect an English speaking audience." And if you know enough English to disagree with that translation, it just proves the point further.

Re: "Uncensored" open LLMs are measurably more optimistic than their base models

#10
post #9
post #7

Earlier quoted context omitted.

yeah English isn't my first language so i used AI to clean up the writing. the research, the data and the analysis are all mine and all open (in the linked repo)

What would you have done pre-LLM? Do that. It's better. Because "English isn't my first language so i used AI to clean up the writing" translates to "I don't respect an English speaking audience." And if you know enough English to disagree with that translation, it just proves the point further.

The amount of unacknowledged privilege in this post is hard to bear.
Post reply on HN