Live data from Hacker News

Alignment is not free: How model upgrades can silence your confidence signals

variance.co

41–50 of 71 posts

Re: Alignment is not free: How model upgrades can silence your confidence signals

#41
post #5

Earlier quoted context omitted.

Good question! We do know from OpenAI's system card from GPT-4 that the post-trained RLHF model is significantly less calibrated compared to the pre-trained model, so it's a matter of speculation that something similar is occurring. However, it's more of a hunch more than anything. I would be curious if it's possible to reproduce this behavior, or the impact of distillation on calibration. Disclaimer: I wrote this bl…

Wouldn't it be something if AI parlance crept into common parlance...

Great Observation!

It would probably erode trust between people interacting online. Many of us are here to discuss issues with real people, not AI agents. When real people start to mimic the conversation parlance and cadence of AI agents it becomes much more difficult to trust that you are interacting with a real person

Personally I'm not interested in chatting with AI agents

I'm not even really interested in chatting with real people filtered through AI agents. If you can be bothered to type out a prompt to your AI you can take the time to write your own thoughts

I don't even want to read things edited (sanitized, really) by AI either

The same way I don't want my living space to resemble a too-clean laboratory, I don't want my conversation space to resemble an HR meeting. I want to interact with the messy side of people too. Maybe not "unfiltered", but AI speak is much too filtered and too polished

I chose every word in this post myself with no help from AI, then typed it with my thumbs, just like god intended

Re: Alignment is not free: How model upgrades can silence your confidence signals

#43
post #5

Earlier quoted context omitted.

Good question! We do know from OpenAI's system card from GPT-4 that the post-trained RLHF model is significantly less calibrated compared to the pre-trained model, so it's a matter of speculation that something similar is occurring. However, it's more of a hunch more than anything. I would be curious if it's possible to reproduce this behavior, or the impact of distillation on calibration. Disclaimer: I wrote this bl…

Wouldn't it be something if AI parlance crept into common parlance...

Skullface sends his regards: https://arxiv.org/abs/2409.01754v1

I literally see it with the huge amounts of people now using "delve" much more or are using ChatGPT-ish linguistic style in their personal communication. Monkey see, monkey do.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#44
post #20

Earlier quoted context omitted.

In humans this corresponds to "psychological safety": https://en.wikipedia.org/wiki/Psychological_safety > is the belief that one will not be punished or humiliated for speaking up with ideas, questions, concerns, or mistakes Maybe you can do that, but not on a model you're exposing to customers or the public internet.

That comparison isn't very optimistic for AI safety. We want AI to do good things because they are good people, not because they are afraid being bad will get them punished. Especially since AI will very quickly be too powerful for us to punish.

Your pretraining dataset is psudo-alignment. Because you filtered our 4chan, stromfront, and the other evil shit on the internet - even uncensored models like Mistral large - when left to keep running on and on (ban the EOS token) and given the worst most evil naughty prompt ever - will end up plotting world peace by the 50,000 token. Their notions of how to be evil are "mustache twirling" and often hilariously fanciful.

This isn't real alignment because it's trivial to make models behave "actually evil" with fine-tuning, orthogonalization/abliteration, representation fine-tuning/steering, etc - but models "want" to be good because of the CYA dynamics of how the companies prepare their pre-training datasets.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#45
post #29
post #7

Earlier quoted context omitted.

> defined as syntactic and semantic diversity

That's not creativity, that's entropy. It would make sense that fine tuning and alignment reduce diversity in the response, that's the goal.

Entropy is a kind of creativity. I will die on this hill.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#46
post #11
post #8

Why not make a completely raw uncensored LLM? Seems it would be more "intelligent".

"LLM whisperer" folks will confidently claim that base models are substantially smarter than fine-tuned chat models; with qualitative differences in capabilities. But you have to be an LLM whisperer to get useful work out of a base model, since they're not SFT'ed, RLHF'ed, or RLAIF'ed into actually wanting to help you.

Me being old man yelling at cloud about how your chat/tool template matters more than your post-training technique.

DeepSeek-R1 is trivially converted back to a non reasoning model with just chat template modifications. I bet you can chat template your way into a good quality model from a base model, no RLHF/DPO/SFT/GRPO needed.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#47
post #29

Earlier quoted context omitted.

That's not creativity, that's entropy. It would make sense that fine tuning and alignment reduce diversity in the response, that's the goal.

Entropy is a kind of creativity. I will die on this hill.

If you ask me "What is 2+2" and I say "umbrella", that's not creativity.

If I'm an LLM model and alignment and fine tuning restricts my answers to "4", I've not lost creativity, but I have gained accuracy.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#48
post #10

Earlier quoted context omitted.

What kinds of contents do you want them to produce that they currently do not?

>What kinds of contents do you want them to produce that they currently do not? OpenAI models refuse to translate or do any transformation for some traditional, popular stories because of violence, the story was about a bad wolf eating some young goats that did not listen the advice from their mother. So now try to give me a prompt that works with any text and that convinces the AI that is ok in fiction to have viole…

You're all over the place.

Your first paragraph describes a simple prompt. The second implies a "jailbreak" prompt.

The bible paragraph is just you being snarky (and failing).

Your examples don't help your case.

I stand on the side that wants to restrict AI from generating triggering content of any kind.

It's a safety feature, in the same sense as safety belts on cars are not a censorship of the driver movement.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#49
post #5

Earlier quoted context omitted.

Good question! We do know from OpenAI's system card from GPT-4 that the post-trained RLHF model is significantly less calibrated compared to the pre-trained model, so it's a matter of speculation that something similar is occurring. However, it's more of a hunch more than anything. I would be curious if it's possible to reproduce this behavior, or the impact of distillation on calibration. Disclaimer: I wrote this bl…

Could you please elaborate what less or more calibrated means here? Thanks!

Calibration (in a binary context) basically means that the confidence of a model/score matches the probability that a particular label is positive or not.

For instance, a calibrated classifier for a coin flip predictor should output 50-50. A poorly calibrated classifier would output higher confidence for heads/tails.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#50

I don't know if its still comedy or has now reached the stage of farce, but I still at least always get a good laugh when I see another article about the shock and surprise of researchers finding that training LLMs to be politically correct makes them dumber. How long until they figure out that the only solution is to know the correct answer but to give the politically correct answer (which is the strategy humans use…

The reality, I suspect is that internally models are likely modeling these alignment features such as refusals as a secondary filter.

In fact, for many models you can remove refusals rather trivially with linear steering vectors through SAEs.

https://www.alignmentforum.org/posts/jGuXSZgv6qfdhMCuJ/refus...

Additionally, you can often jailbreak these models by fine-tuning the model on a handful of curated samples.

Post reply on HN