Live data from Hacker News

Alignment is not free: How model upgrades can silence your confidence signals

variance.co

61–70 of 71 posts

Re: Alignment is not free: How model upgrades can silence your confidence signals

#61
post #57

Earlier quoted context omitted.

You can complain to their support, not to me. I don't find it sensitive, and I remain on the side of ethical restrictions.

Someone asked here examples of what people are using that triggers the censorship, I gave you example of legal,, moral and normal content because the implication is that you only get censored if you are trying to do illegal stuff or adult stuff. If you only use it for code you will not see the censorship that often, though Gemini once refused to write a SQL DELETE because it is to dangerous.

You said you wanted "no censorship", I explained why it exists with a cheerful metaphor, then you said "it's too sensitive" (like your car seat belt is too tight).

Decide what you are. If you want no seat belts, I think you're insane. If it's too tight, then you need to complain to the manufacturer.

I only asked about examples to make you explain what you meant. Once it was clear, the conversation actually ended.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#62
post #17
post #11

Earlier quoted context omitted.

"LLM whisperer" folks will confidently claim that base models are substantially smarter than fine-tuned chat models; with qualitative differences in capabilities. But you have to be an LLM whisperer to get useful work out of a base model, since they're not SFT'ed, RLHF'ed, or RLAIF'ed into actually wanting to help you.

How can I learn more about this? Is it like in the early GPT-3 days, when you had to give it a bunch of examples and hope it catches the pattern?

Not so much examples, though those can help... but you have to imagine a document of a sort that would be in the training set whose completion would be the answer you seek.

Like, "Solve this equation for me: " more likely gets completed with "Do your own homework buddy!" or just a list of more similar questions without answers. While, "careful analysis revealed the solution the equation X turned out to have a solution of", might be more likely to get what you want.

Also a lot more sensitivity to tone and context, write a prompt that sounds like it was written on some teenager fan subreddit, you'll get an answer of the sort that sounds like it belongs there.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#63

Earlier quoted context omitted.

That comparison isn't very optimistic for AI safety. We want AI to do good things because they are good people, not because they are afraid being bad will get them punished. Especially since AI will very quickly be too powerful for us to punish.

Your pretraining dataset is psudo-alignment. Because you filtered our 4chan, stromfront, and the other evil shit on the internet - even uncensored models like Mistral large - when left to keep running on and on (ban the EOS token) and given the worst most evil naughty prompt ever - will end up plotting world peace by the 50,000 token. Their notions of how to be evil are "mustache twirling" and often hilariously fanci…

> it's trivial to make models behave "actually evil" with fine-tuning, orthogonalization/abliteration, representation fine-tuning/steering, etc

It's actually pretty difficult to do this and make them useful. You can see this because Grok is a helpful liberal just like all the other models.

Evil / illiberal people don't answer questions on the internet! So there is no personality in the base model for you to uncover that is both illiberal and capable of helpfully answering questions. If they tried to make a Grok that acted like the typical new-age X user, it'd just respond to any prompt by calling you a slur you've never heard of.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#64

Earlier quoted context omitted.

Wouldn't it be something if AI parlance crept into common parlance...

Great Observation! It would probably erode trust between people interacting online. Many of us are here to discuss issues with real people, not AI agents. When real people start to mimic the conversation parlance and cadence of AI agents it becomes much more difficult to trust that you are interacting with a real person Personally I'm not interested in chatting with AI agents I'm not even really interested in chattin…

> Personally I'm not interested in chatting with AI agents

Why, though? If the AI agent is making sense, then what does it matter?

For certain types of conversations I've had more interesting conversations with AI than with a solid 90% of people I've ever interacted with. Really not that surprising given that most people have only an average grasp of most things and a poor grasp of very specific things.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#65

Earlier quoted context omitted.

Great Observation! It would probably erode trust between people interacting online. Many of us are here to discuss issues with real people, not AI agents. When real people start to mimic the conversation parlance and cadence of AI agents it becomes much more difficult to trust that you are interacting with a real person Personally I'm not interested in chatting with AI agents I'm not even really interested in chattin…

> Personally I'm not interested in chatting with AI agents Why, though? If the AI agent is making sense, then what does it matter? For certain types of conversations I've had more interesting conversations with AI than with a solid 90% of people I've ever interacted with. Really not that surprising given that most people have only an average grasp of most things and a poor grasp of very specific things.

> Why, though? If the AI agent is making sense, then what does it matter?

The same reason I prefer having sex with humans and not blow up dolls

If your only goal is to get off, then the blow up doll does the job. If all you care about is having an interesting conversation then I guess an LLM is fine

I care about human connection. I have no interest in spending time interacting with machines instead of people

Re: Alignment is not free: How model upgrades can silence your confidence signals

#66

Earlier quoted context omitted.

Your pretraining dataset is psudo-alignment. Because you filtered our 4chan, stromfront, and the other evil shit on the internet - even uncensored models like Mistral large - when left to keep running on and on (ban the EOS token) and given the worst most evil naughty prompt ever - will end up plotting world peace by the 50,000 token. Their notions of how to be evil are "mustache twirling" and often hilariously fanci…

> it's trivial to make models behave "actually evil" with fine-tuning, orthogonalization/abliteration, representation fine-tuning/steering, etc It's actually pretty difficult to do this and make them useful. You can see this because Grok is a helpful liberal just like all the other models. Evil / illiberal people don't answer questions on the internet! So there is no personality in the base model for you to uncover t…

Grok didn't use the techniques listed above because even elon musk will not take the risks associated with models which are willing to do any number of illegal things.

It is not difficult to do this and make them useful at all. Please familiarize yourself with the literature.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#67
post #61

Earlier quoted context omitted.

Someone asked here examples of what people are using that triggers the censorship, I gave you example of legal,, moral and normal content because the implication is that you only get censored if you are trying to do illegal stuff or adult stuff. If you only use it for code you will not see the censorship that often, though Gemini once refused to write a SQL DELETE because it is to dangerous.

You said you wanted "no censorship", I explained why it exists with a cheerful metaphor, then you said "it's too sensitive" (like your car seat belt is too tight). Decide what you are. If you want no seat belts, I think you're insane. If it's too tight, then you need to complain to the manufacturer. I only asked about examples to make you explain what you meant. Once it was clear, the conversation actually ended.

Go back to the start of the thread, I gave example of censorship either beeing buggy or just stupidly setup so it makes happy both USA extremes.

I give you examples, I do not ask you to fix it. YOu just need to have the mental strength to admit that other people hit the censorship in day to day, in work related circumstances instead of defending the Big Tech by pretending that nobody normal would hit this issues.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#68

Earlier quoted context omitted.

> Personally I'm not interested in chatting with AI agents Why, though? If the AI agent is making sense, then what does it matter? For certain types of conversations I've had more interesting conversations with AI than with a solid 90% of people I've ever interacted with. Really not that surprising given that most people have only an average grasp of most things and a poor grasp of very specific things.

> Why, though? If the AI agent is making sense, then what does it matter? The same reason I prefer having sex with humans and not blow up dolls If your only goal is to get off, then the blow up doll does the job. If all you care about is having an interesting conversation then I guess an LLM is fine I care about human connection. I have no interest in spending time interacting with machines instead of people

That is a silly comparison.

1. Humans and blow up dolls feel massively different, physically. 2. Blow up dolls don't do anything autonomously.

The comparison would have to be with a sex bot that is virtually indistinguishable from a human when having sex with it, just like text chatting with an AI bot versus chatting with a human can be.

What human connection are you and I currently forming? Does it really matter that I am a human for this interaction we're having?

Re: Alignment is not free: How model upgrades can silence your confidence signals

#69

Earlier quoted context omitted.

> it's trivial to make models behave "actually evil" with fine-tuning, orthogonalization/abliteration, representation fine-tuning/steering, etc It's actually pretty difficult to do this and make them useful. You can see this because Grok is a helpful liberal just like all the other models. Evil / illiberal people don't answer questions on the internet! So there is no personality in the base model for you to uncover t…

Grok didn't use the techniques listed above because even elon musk will not take the risks associated with models which are willing to do any number of illegal things. It is not difficult to do this and make them useful at all. Please familiarize yourself with the literature.

Elon has never followed a law in his life and he's not going to start now.

Re: Alignment is not free: How model upgrades can silence your confidence signals

#70

Earlier quoted context omitted.

Grok didn't use the techniques listed above because even elon musk will not take the risks associated with models which are willing to do any number of illegal things. It is not difficult to do this and make them useful at all. Please familiarize yourself with the literature.

Elon has never followed a law in his life and he's not going to start now.

[flagged]
Post reply on HN