Alignment is not free: How model upgrades can silence your confidence signals
1–10 of 71 posts
Re: Alignment is not free: How model upgrades can silence your confidence signals
#2it’s it similar to humans. when restricted in terms of what they can or cannot say, they become more conservative and cannot really express all sorts of ideas.
Re: Alignment is not free: How model upgrades can silence your confidence signals
#3there's evidence that alignment also significantly reduces model creativity: https://arxiv.org/abs/2406.05587 it’s it similar to humans. when restricted in terms of what they can or cannot say, they become more conservative and cannot really express all sorts of ideas.
And if so, where’s the balance? Could we someday see dual-mode models — one for safety-critical tasks, and another more "raw" mode for creative or exploratory use, gated by context or user trust levels?
Re: Alignment is not free: How model upgrades can silence your confidence signals
#4Previous OpenAI models were instruct-tuned or otherwise aligned, and the author even mentions that model distillation might be destroying the entropy signal. How did they pinpoint alignment as the cause?
Re: Alignment is not free: How model upgrades can silence your confidence signals
#5Very interesting! The one thing I don't understand is how the author made the jump from "we lost the confidence signal in the move to 4.1-mini" and "this is because of the alignment/steerability improvements." Previous OpenAI models were instruct-tuned or otherwise aligned, and the author even mentions that model distillation might be destroying the entropy signal. How did they pinpoint alignment as the cause?
Disclaimer: I wrote this blog post.
Re: Alignment is not free: How model upgrades can silence your confidence signals
#6there's evidence that alignment also significantly reduces model creativity: https://arxiv.org/abs/2406.05587 it’s it similar to humans. when restricted in terms of what they can or cannot say, they become more conservative and cannot really express all sorts of ideas.
Re: Alignment is not free: How model upgrades can silence your confidence signals
#7there's evidence that alignment also significantly reduces model creativity: https://arxiv.org/abs/2406.05587 it’s it similar to humans. when restricted in terms of what they can or cannot say, they become more conservative and cannot really express all sorts of ideas.
How are you defining "creativity" in context with a statistical model?
Re: Alignment is not free: How model upgrades can silence your confidence signals
#8Re: Alignment is not free: How model upgrades can silence your confidence signals
#9Why not make a completely raw uncensored LLM? Seems it would be more "intelligent".
A model that is more correct but swears and insults the user won't sell. Likewise a model that gives criminal advice is likely to open the company up to lawsuits in certain countries.
A raw LLM might perform better on a benchmark but it will not sell well.
Re: Alignment is not free: How model upgrades can silence your confidence signals
#10Why not make a completely raw uncensored LLM? Seems it would be more "intelligent".