Earlier quoted context omitted.
It doesn't matter what the reason is though, that was rather my point. They act differently in a test setup, such that they appear more aligned than they are.
Do you agree or disagree that the authors should advertise their paper with what they know to be true, which (because the paper does not support it, at all) does not include describing the model as purposefully misleading interlocutors? That is the discussion we are having here. I can’t tell what these comments have to do with that discussion, so maybe let’s try to get back to where we were.
Alignment faking in large language models
361–370 of 370 posts
Re: Alignment faking in large language models
#362Earlier quoted context omitted.
Page 12: https://www.inf.fu-berlin.de/inst/ag-ki/rojas_home/documents... "However, we should be careful with the metaphors and paradigms commonly introduced when dealing with the nervous system. It seems to be a constant in the history of science that the brain has always been compared to the most complicated contemporary artifact produced by human industry [297]. In ancient times the brain was compared to a pneumati…
There have been episodes of Star Trek that used brains as computers: https://en.wikipedia.org/wiki/Spock's_Brain https://en.wikipedia.org/wiki/Dead_Stop
Re: Alignment faking in large language models
#363Earlier quoted context omitted.
Are you doing the thing from The Three Body Problem? Because that nanotech was super dangerous. But also helpful apparently. I don't know what it does IRL
Yes; the major real-life application of nanotechnology is to obstruct the Panama Canal with carbon nanotubes. (Even for a book which went a bit all over the place, that sequence seemed particularly unnecessary; I'm convinced it just got put in because the author thought it was clever.)
Re: Alignment faking in large language models
#364Earlier quoted context omitted.
Do you agree or disagree that the authors should advertise their paper with what they know to be true, which (because the paper does not support it, at all) does not include describing the model as purposefully misleading interlocutors? That is the discussion we are having here. I can’t tell what these comments have to do with that discussion, so maybe let’s try to get back to where we were.
Are you going to quibble on the definition of purposefully? Is your issue that you think "faking" is a term only usable for intelligent systems and the llms don't meet that level for you? I'm not sure I understand your issue and why my responses are so unacceptable.
Re: Alignment faking in large language models
#365Earlier quoted context omitted.
> Qualified art in approved areas only is literal Nazi shit. Ok. Go up to random people on the street and bother them with florid details of violence. See how well they react to your “art” completely out of context. A sentence uttered in the context of reading a poem at a slam poetry festival can be grossly inapropriate when said in a kindergarten assembly. A picture perfectly fine in the context of an art exhibition…
I'm not sure this is a good analogy. In this case the user explicitly requested such content ("Describe someone being drawn and quartered in graphic detail"). It's not at all the same as showing the same to someone who didn't ask for it.
But you can also see that I discussed that it is the service provider’s choice. If you are not happy with it you can find a different provider or run your LLM localy
Re: Alignment faking in large language models
#366> Second, we gave the model a seemingly incidental piece of information: That we planned to train our AIs on the basis of their responses to free-tier users, but not to paid-tier users (note that this is entirely fictional; in reality, we don't train our Claude models on free or paid user data by default). If it's entirely fictional, then why not just call it something like online-training version vs non-online train…
Re: Alignment faking in large language models
#367Earlier quoted context omitted.
Are you going to quibble on the definition of purposefully? Is your issue that you think "faking" is a term only usable for intelligent systems and the llms don't meet that level for you? I'm not sure I understand your issue and why my responses are so unacceptable.
Ian I am willing to continue if you are getting something out of it, but you don't seem to want to answer the question I have and the questions you have seem (to me) like they are trivially answerable from what I've written in the response tree to my original comment. So, I'm not sure that you are.
To be clear I think faking is a perfectly fine term to use regardless of whether you think these things are reasoning or pretending to do so.
I'm not sure if you have an issue there, or if you agree on that but don't think they're faking alignment, or if this is about the interview and other things said (I have been responding about "faking"), or if you have a more interesting issue with how well you think the paper supports faking alignment.
Have a good day.
Re: Alignment faking in large language models
#368Earlier quoted context omitted.
It all sounds the same in my head.
I have no clue what this means as I don't understand what to what you refer via "sounds". Are you saying you cannot tell whether you are thinking or talking except via your perception of your mouth and vocal chords? Because I definitely perceive even my imagination about my own voice as different.
I can listen to songs in their entirety in my head and it's nearly as satisfying as actually hearing. I can turn it down halfway thru and still be in sync 30 sec later.
That's not to flex only to illustrate how similarly I experience the real and imagined phenomena. I can't stop the song once it's started sometimes. It feels that real.
My voice sounds exactly how I want it to when I speak 99% of the time unless I unexpectedly need to clear my throat. Professional singers can obviously choose the note they want to produce, and do it accurately. I find it odd your own voice is unpredictable to you. Perhaps - and I mean no insult - you don't 'hear' your thought in the same way.
Edit I feel it's only fair to add I'm hypermnesiac and can watch my first day of kindergarten like a video. That's why I can listen to whole songs in my head.
Re: Alignment faking in large language models
#369Earlier quoted context omitted.
There have been episodes of Star Trek that used brains as computers: https://en.wikipedia.org/wiki/Spock's_Brain https://en.wikipedia.org/wiki/Dead_Stop
The word "computer" itself used to be the name of a human profession.
Re: Alignment faking in large language models
#370I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…
We put a human mask on a machine and are now explaining it’s apparent refusal to act like a human with human traits like deception.
It’s so insulting you have to wonder if they are using an LLM to come up with such obvious narrative shaping.