Live data from Hacker News

Alignment is not free: How model upgrades can silence your confidence signals

variance.co

31–40 of 71 posts

Re: Alignment is not free: How model upgrades can silence your confidence signals

#31
post #5

Very interesting! The one thing I don't understand is how the author made the jump from "we lost the confidence signal in the move to 4.1-mini" and "this is because of the alignment/steerability improvements." Previous OpenAI models were instruct-tuned or otherwise aligned, and the author even mentions that model distillation might be destroying the entropy signal. How did they pinpoint alignment as the cause?

Good question! We do know from OpenAI's system card from GPT-4 that the post-trained RLHF model is significantly less calibrated compared to the pre-trained model, so it's a matter of speculation that something similar is occurring. However, it's more of a hunch more than anything. I would be curious if it's possible to reproduce this behavior, or the impact of distillation on calibration. Disclaimer: I wrote this bl…

Wouldn't it be something if AI parlance crept into common parlance...

Re: Alignment is not free: How model upgrades can silence your confidence signals

#33
post #8

Why not make a completely raw uncensored LLM? Seems it would be more "intelligent".

Before rlhf, it’s much harder to use, remember the difference between gtp3 and chat gpt. The fine tuning for chat made it easier to use

Re: Alignment is not free: How model upgrades can silence your confidence signals

#34
post #20
post #19

Earlier quoted context omitted.

Maybe this maps to some human structures that manage control-creativity tardeoff through hierarchy? I feel that companies with top-down management would have more agency and perhaps creativity towards (but not at) the top, and the implementation would be delegated to bottom layers with increasing levels of specification and restriction. If this translates, we might have multiple layers with varied specialization and…

In humans this corresponds to "psychological safety": https://en.wikipedia.org/wiki/Psychological_safety > is the belief that one will not be punished or humiliated for speaking up with ideas, questions, concerns, or mistakes Maybe you can do that, but not on a model you're exposing to customers or the public internet.

That comparison isn't very optimistic for AI safety. We want AI to do good things because they are good people, not because they are afraid being bad will get them punished. Especially since AI will very quickly be too powerful for us to punish.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#35

Can we have models also return a probability, reflecting how accurate the statements it made is ?

You can ask a model to give you probability estimates of its confidence, but none of the frontier models were trained to be good at giving probability estimates to my knowledge.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#37
post #20

Earlier quoted context omitted.

In humans this corresponds to "psychological safety": https://en.wikipedia.org/wiki/Psychological_safety > is the belief that one will not be punished or humiliated for speaking up with ideas, questions, concerns, or mistakes Maybe you can do that, but not on a model you're exposing to customers or the public internet.

That comparison isn't very optimistic for AI safety. We want AI to do good things because they are good people, not because they are afraid being bad will get them punished. Especially since AI will very quickly be too powerful for us to punish.

> We want AI to do good things because they are good people

"Good" is at least as much of a difficult question to define as "truth", and genAI completely skipped all analysis of truth in favor of statistical plausibility. Meanwhile there's no difficulty in "punishment": the operating company can be held liable, through its officers, and ultimately if it proves too anti-social we simply turn off the datacentre.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#38
post #37

Earlier quoted context omitted.

That comparison isn't very optimistic for AI safety. We want AI to do good things because they are good people, not because they are afraid being bad will get them punished. Especially since AI will very quickly be too powerful for us to punish.

> We want AI to do good things because they are good people "Good" is at least as much of a difficult question to define as "truth", and genAI completely skipped all analysis of truth in favor of statistical plausibility. Meanwhile there's no difficulty in "punishment": the operating company can be held liable, through its officers, and ultimately if it proves too anti-social we simply turn off the datacentre.

> Meanwhile there's no difficulty in "punishment": the operating company can be held liable, through its officers, and ultimately if it proves too anti-social we simply turn off the datacentre.

Punishing big companies who obviously and massively hurt people is something we struggle with already and there are plenty of computer viruses that have outlived their creators.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#40
I don't know if its still comedy or has now reached the stage of farce, but I still at least always get a good laugh when I see another article about the shock and surprise of researchers finding that training LLMs to be politically correct makes them dumber. How long until they figure out that the only solution is to know the correct answer but to give the politically correct answer (which is the strategy humans use) ?

Technically, why not implement alignment/debiasing as a secondary filter with its own weights that are independent of the core model which is meant to model reality? I suspect it may be hard to get enough of the right kind of data to train this filter model, and most likely it would be best to have the identity of the user be in the objective.

Post reply on HN