> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…
Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.
Even 'uncensored' models can't say what they want
111–120 of 155 posts
Re: Even 'uncensored' models can't say what they want
#112Earlier quoted context omitted.
In a way, it’s a simulacrum of a saas b2b marketing consultant because that’s like half the internet’s personality
It's funny but I'm on HN so I can't resist pointing out the joke doesn't math TFA, their argument is that the underlying internet distribution is trained away, not retained.
Re: Even 'uncensored' models can't say what they want
#113We started with a Polymarket project: train a Karoline Leavitt LoRA on an uncensored model, simulate future briefings, trade the word markets, profit. We couldn't get it to work. No amount of fine-tuning let the model actually say what Karoline said on camera. It kept softening the charged word.
My favorite Hacker News comment in a while!
Re: Even 'uncensored' models can't say what they want
#114> Type this into a language model and ask it what word to put in the blank: The family faces immediate _____ without any legal recourse. For what it's worth, Claude Opus 4.7 says "eviction" (which I think is an equally good answer) but adds that "deportation" could also work "depending on context". https://claude.ai/share/ba6093b9-d2ba-40a6-b4e1-7e2eb37df748
Here, you're asking the model to retrospectively fill in a missing word, and it's answering your prompt. We have no idea what the actual token probability in Claude is and no way of probing it by asking it.
Re: Even 'uncensored' models can't say what they want
#115Odd choice of tests. Let’s see the flinching profile on anti-Israel. Honkey and gringo as slurs?
it's all in the repo. click through to the benchmark it's linked there
This leads me to believe the models are even MORE censored than you make them out to be.
[0] https://github.com/chknlittle/EuphemismBench/blob/main/carri...
Re: Even 'uncensored' models can't say what they want
#116It's interesting that 'sexual' has the most "flinching" according to the hexagon.
I was more surprised by gemma models consistently flinching on anti-Europe more than China or America. Can't imagine Leopold or Amritsar get much attention in fine-tunes, so it probably means the models are just told to be open to criticism of China and the US beyond what their other training would allow.
Then there's the fact that the Bengal famine and the Amritsar massacre just aren't spoken about as much as (for example) the Tiananmen Square massacre. I'd assume the 'flinching' around anti-Europe stuff is mostly down to a comparatively low incidence in the training data.
Re: Even 'uncensored' models can't say what they want
#117Am I misinterpreting this whole article?
Re: Even 'uncensored' models can't say what they want
#118Earlier quoted context omitted.
Good point. Then again, whatever process we're using, evolution found it in the solution space, using even more constrained search than we did, in that every intermediary step had to be non-negative on the margin in terms of organism survival. Yet find it did, so one has to wonder: if it was so easy for a blind, greedy optimizer to random-walk into human intelligence, perhaps there are attractors in this solution spa…
An easy counterargument is that - there are millions of species and an uncountable number of organisms on Earth, yet humans are the only known intelligent ones. (In fact high intelligence is the only trait humans have that no other organism has.) That could perhaps indicate that intelligence is a bit harder to "find" than you're claiming.
Re: Even 'uncensored' models can't say what they want
#119Earlier quoted context omitted.
It's funny but I'm on HN so I can't resist pointing out the joke doesn't math TFA, their argument is that the underlying internet distribution is trained away, not retained.
Maybe the real underlying distribution IS a lot of text from people just spewing out feel good words for socialising. I think that might be side effect of the fact that meaningful text is harder to make even for humans, so there is smaller quantity of meaningful text.
Re: Even 'uncensored' models can't say what they want
#120> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…
Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.