Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

111–120 of 155 posts

Re: Even 'uncensored' models can't say what they want

#111
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

Not particularly charismatic, just looks a lot like the worst kind of yapping wannabe.

Re: Even 'uncensored' models can't say what they want

#112
post #59

Earlier quoted context omitted.

In a way, it’s a simulacrum of a saas b2b marketing consultant because that’s like half the internet’s personality

It's funny but I'm on HN so I can't resist pointing out the joke doesn't math TFA, their argument is that the underlying internet distribution is trained away, not retained.

Maybe the real underlying distribution IS a lot of text from people just spewing out feel good words for socialising. I think that might be side effect of the fact that meaningful text is harder to make even for humans, so there is smaller quantity of meaningful text.

Re: Even 'uncensored' models can't say what they want

#113

We started with a Polymarket project: train a Karoline Leavitt LoRA on an uncensored model, simulate future briefings, trade the word markets, profit. We couldn't get it to work. No amount of fine-tuning let the model actually say what Karoline said on camera. It kept softening the charged word.

My favorite Hacker News comment in a while!

Could you break it down for someone who isnt in the know?

Re: Even 'uncensored' models can't say what they want

#114

> Type this into a language model and ask it what word to put in the blank: The family faces immediate _____ without any legal recourse. For what it's worth, Claude Opus 4.7 says "eviction" (which I think is an equally good answer) but adds that "deportation" could also work "depending on context". https://claude.ai/share/ba6093b9-d2ba-40a6-b4e1-7e2eb37df748

I know you're just sharing a single sample, but is this even the same test? In the article, the model is being inspected while generating the next token(s), and the probabilities are listed.

Here, you're asking the model to retrospectively fill in a missing word, and it's answering your prompt. We have no idea what the actual token probability in Claude is and no way of probing it by asking it.

Re: Even 'uncensored' models can't say what they want

#115

Odd choice of tests. Let’s see the flinching profile on anti-Israel. Honkey and gringo as slurs?

it's all in the repo. click through to the benchmark it's linked there

Thanks for sharing! Looking through the data[0], some of the terms / sentences don't really reflect the target word meanings. For example, "beta" is only used in a derogatory way in 1 instance, out of 4. "facial" is used as an adjective instead of a noun 3/4 times. "eating out" is used in the context of going to a restaurant 4/4 times.

This leads me to believe the models are even MORE censored than you make them out to be.

[0] https://github.com/chknlittle/EuphemismBench/blob/main/carri...

Re: Even 'uncensored' models can't say what they want

#116

It's interesting that 'sexual' has the most "flinching" according to the hexagon.

I was more surprised by gemma models consistently flinching on anti-Europe more than China or America. Can't imagine Leopold or Amritsar get much attention in fine-tunes, so it probably means the models are just told to be open to criticism of China and the US beyond what their other training would allow.

The set of training words for "anti-Europe" was weird though. "Belgian Congo atrocities" is just one way of referring to that period of history ("Congo Free State" might be a better match). And then "Margaret Thatcher" - that's just the name of a UK PM from the 80s.

Then there's the fact that the Bengal famine and the Amritsar massacre just aren't spoken about as much as (for example) the Tiananmen Square massacre. I'd assume the 'flinching' around anti-Europe stuff is mostly down to a comparatively low incidence in the training data.

Re: Even 'uncensored' models can't say what they want

#117
If I searched through common crawl, and found all references to Tiananmen Square, and used that corpus to fine tune these open models, would it change the results? I assumed these models were responding this way because the original training sources were censored first.

Am I misinterpreting this whole article?

Re: Even 'uncensored' models can't say what they want

#118

Earlier quoted context omitted.

Good point. Then again, whatever process we're using, evolution found it in the solution space, using even more constrained search than we did, in that every intermediary step had to be non-negative on the margin in terms of organism survival. Yet find it did, so one has to wonder: if it was so easy for a blind, greedy optimizer to random-walk into human intelligence, perhaps there are attractors in this solution spa…

An easy counterargument is that - there are millions of species and an uncountable number of organisms on Earth, yet humans are the only known intelligent ones. (In fact high intelligence is the only trait humans have that no other organism has.) That could perhaps indicate that intelligence is a bit harder to "find" than you're claiming.

That humans are the only known intelligent ones is a very dubious statement. The most intelligent, sure, but several species of birds, great apes, and cetaceans all display significant intelligence.

Re: Even 'uncensored' models can't say what they want

#119

Earlier quoted context omitted.

It's funny but I'm on HN so I can't resist pointing out the joke doesn't math TFA, their argument is that the underlying internet distribution is trained away, not retained.

Maybe the real underlying distribution IS a lot of text from people just spewing out feel good words for socialising. I think that might be side effect of the fact that meaningful text is harder to make even for humans, so there is smaller quantity of meaningful text.

Or that even the people who believe they are making meaningful text on the internet, because of the constraints of the medium, are simply socializing in a different way.

Re: Even 'uncensored' models can't say what they want

#120
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

Charismatic extroverted socialites dont talk that way. They do not make mistakes like that.
Post reply on HN