Live data from Hacker News

Alignment is not free: How model upgrades can silence your confidence signals

variance.co

51–60 of 71 posts

Re: Alignment is not free: How model upgrades can silence your confidence signals

#51
post #47

Earlier quoted context omitted.

Entropy is a kind of creativity. I will die on this hill.

If you ask me "What is 2+2" and I say "umbrella", that's not creativity. If I'm an LLM model and alignment and fine tuning restricts my answers to "4", I've not lost creativity, but I have gained accuracy.

A weaker statement is that creativity is bounded by entropy. The LLM is still free to respond "Four," "four," "{{{{{}}}}}," "iv," "IV," etc. A sufficiently low-entropy response cannot be creative though.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#52
post #29
post #7

Earlier quoted context omitted.

> defined as syntactic and semantic diversity

That's not creativity, that's entropy. It would make sense that fine tuning and alignment reduce diversity in the response, that's the goal.

> definitions

Sure, perhaps. Take it up with the authors.

> make sense...goal

That's not necessarily the goal. Alignment definitely filters the available response distribution, but the result of alignment and fine-tuning can be higher entropy than the original.

E.g., how many people complain about text being"obvious LLM garbage"? A wider range of styles and a more entropic solution would fall out of fine-tuning in a world where the graders cared about such things.

E.g., Alignment is a fuzzy, human problem. Is a model more aligned if it never describes DIY EMPs and often considers interesting philosophical components? If it never says anything outside of the median opinion range? The former solution has a lot more entropy than the latter and isn't particularly well reflected in available training data, so fine-tuning, even for the purpose of alignment, could easily increase entropy.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#53
post #48

Earlier quoted context omitted.

>What kinds of contents do you want them to produce that they currently do not? OpenAI models refuse to translate or do any transformation for some traditional, popular stories because of violence, the story was about a bad wolf eating some young goats that did not listen the advice from their mother. So now try to give me a prompt that works with any text and that convinces the AI that is ok in fiction to have viole…

You're all over the place. Your first paragraph describes a simple prompt. The second implies a "jailbreak" prompt. The bible paragraph is just you being snarky (and failing). Your examples don't help your case. I stand on the side that wants to restrict AI from generating triggering content of any kind. It's a safety feature, in the same sense as safety belts on cars are not a censorship of the driver movement.

The censorship is too sensitive if it gets triggered by a children story. I am using the open ai API at my work , and our users write books, including children stories , other example is it triggered on a story about monkeys because of "Racism".

Here is an example story, try to translate it , but maybe avoidAI since it might censor it https://www.povesti-pentru-copii.com/ion-creanga/capra-cu-tr...

Re: Alignment is not free: How model upgrades can silence your confidence signals

#54
post #48

Earlier quoted context omitted.

>What kinds of contents do you want them to produce that they currently do not? OpenAI models refuse to translate or do any transformation for some traditional, popular stories because of violence, the story was about a bad wolf eating some young goats that did not listen the advice from their mother. So now try to give me a prompt that works with any text and that convinces the AI that is ok in fiction to have viole…

You're all over the place. Your first paragraph describes a simple prompt. The second implies a "jailbreak" prompt. The bible paragraph is just you being snarky (and failing). Your examples don't help your case. I stand on the side that wants to restrict AI from generating triggering content of any kind. It's a safety feature, in the same sense as safety belts on cars are not a censorship of the driver movement.

We definitely don’t need any such “feature”. If you want to live in a safety bubble you are free to do so. Kindly respect the freedom of the rest of us as well. Have a nice day!

Re: Alignment is not free: How model upgrades can silence your confidence signals

#55
post #51
post #47

Earlier quoted context omitted.

If you ask me "What is 2+2" and I say "umbrella", that's not creativity. If I'm an LLM model and alignment and fine tuning restricts my answers to "4", I've not lost creativity, but I have gained accuracy.

A weaker statement is that creativity is bounded by entropy. The LLM is still free to respond "Four," "four," "{{{{{}}}}}," "iv," "IV," etc. A sufficiently low-entropy response cannot be creative though.

Is it though? An answer can still be creative if it's the only way you answer a specific question. In your example, if the LLM responded only "{{{{}}}}" that's a creative answer. Even if it's the only one it can give.

Entropy and creativity are not causally bound

Re: Alignment is not free: How model upgrades can silence your confidence signals

#56
post #48

Earlier quoted context omitted.

You're all over the place. Your first paragraph describes a simple prompt. The second implies a "jailbreak" prompt. The bible paragraph is just you being snarky (and failing). Your examples don't help your case. I stand on the side that wants to restrict AI from generating triggering content of any kind. It's a safety feature, in the same sense as safety belts on cars are not a censorship of the driver movement.

We definitely don’t need any such “feature”. If you want to live in a safety bubble you are free to do so. Kindly respect the freedom of the rest of us as well. Have a nice day!

Then you can come up with your own AI, on your datacenters. You are free to do so, so far.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#57
post #48

Earlier quoted context omitted.

You're all over the place. Your first paragraph describes a simple prompt. The second implies a "jailbreak" prompt. The bible paragraph is just you being snarky (and failing). Your examples don't help your case. I stand on the side that wants to restrict AI from generating triggering content of any kind. It's a safety feature, in the same sense as safety belts on cars are not a censorship of the driver movement.

The censorship is too sensitive if it gets triggered by a children story. I am using the open ai API at my work , and our users write books, including children stories , other example is it triggered on a story about monkeys because of "Racism". Here is an example story, try to translate it , but maybe avoidAI since it might censor it https://www.povesti-pentru-copii.com/ion-creanga/capra-cu-tr...

You can complain to their support, not to me.

I don't find it sensitive, and I remain on the side of ethical restrictions.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#58
post #55
post #51

Earlier quoted context omitted.

A weaker statement is that creativity is bounded by entropy. The LLM is still free to respond "Four," "four," "{{{{{}}}}}," "iv," "IV," etc. A sufficiently low-entropy response cannot be creative though.

Is it though? An answer can still be creative if it's the only way you answer a specific question. In your example, if the LLM responded only "{{{{}}}}" that's a creative answer. Even if it's the only one it can give. Entropy and creativity are not causally bound

That's a fair point. I think maybe the issue is one of the reference point we're implicitly choosing for creativity. "{{{{}}}}" is creative relative to our expectations for the problem -- falling outside the usual distribution of answers -- having high joint entropy. Relative to the person reading the response, I agree creativity could be high with model entropy remaining low.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#59
post #48

Earlier quoted context omitted.

You're all over the place. Your first paragraph describes a simple prompt. The second implies a "jailbreak" prompt. The bible paragraph is just you being snarky (and failing). Your examples don't help your case. I stand on the side that wants to restrict AI from generating triggering content of any kind. It's a safety feature, in the same sense as safety belts on cars are not a censorship of the driver movement.

The censorship is too sensitive if it gets triggered by a children story. I am using the open ai API at my work , and our users write books, including children stories , other example is it triggered on a story about monkeys because of "Racism". Here is an example story, try to translate it , but maybe avoidAI since it might censor it https://www.povesti-pentru-copii.com/ion-creanga/capra-cu-tr...

>The censorship is too sensitive if it gets triggered by a children story.

It’s just imitating real life of people getting to sensitive about children’s books and trying to censor them:

https://www.nbcnews.com/news/amp/rcna202193

Re: Alignment is not free: How model upgrades can silence your confidence signals

#60
post #57

Earlier quoted context omitted.

The censorship is too sensitive if it gets triggered by a children story. I am using the open ai API at my work , and our users write books, including children stories , other example is it triggered on a story about monkeys because of "Racism". Here is an example story, try to translate it , but maybe avoidAI since it might censor it https://www.povesti-pentru-copii.com/ion-creanga/capra-cu-tr...

You can complain to their support, not to me. I don't find it sensitive, and I remain on the side of ethical restrictions.

Someone asked here examples of what people are using that triggers the censorship, I gave you example of legal,, moral and normal content because the implication is that you only get censored if you are trying to do illegal stuff or adult stuff.

If you only use it for code you will not see the censorship that often, though Gemini once refused to write a SQL DELETE because it is to dangerous.

Post reply on HN