Live data from Hacker News

Alignment faking in large language models

anthropic.com

361–370 of 370 posts

Re: Alignment faking in large language models

#361
post #353
post #337

Earlier quoted context omitted.

It doesn't matter what the reason is though, that was rather my point. They act differently in a test setup, such that they appear more aligned than they are.

Do you agree or disagree that the authors should advertise their paper with what they know to be true, which (because the paper does not support it, at all) does not include describing the model as purposefully misleading interlocutors? That is the discussion we are having here. I can’t tell what these comments have to do with that discussion, so maybe let’s try to get back to where we were.

Are you going to quibble on the definition of purposefully? Is your issue that you think "faking" is a term only usable for intelligent systems and the llms don't meet that level for you? I'm not sure I understand your issue and why my responses are so unacceptable.

Re: Alignment faking in large language models

#362
post #293

Earlier quoted context omitted.

Page 12: https://www.inf.fu-berlin.de/inst/ag-ki/rojas_home/documents... "However, we should be careful with the metaphors and paradigms commonly introduced when dealing with the nervous system. It seems to be a constant in the history of science that the brain has always been compared to the most complicated contemporary artifact produced by human industry [297]. In ancient times the brain was compared to a pneumati…

There have been episodes of Star Trek that used brains as computers: https://en.wikipedia.org/wiki/Spock's_Brain https://en.wikipedia.org/wiki/Dead_Stop

The word "computer" itself used to be the name of a human profession.

Re: Alignment faking in large language models

#363

Earlier quoted context omitted.

Are you doing the thing from The Three Body Problem? Because that nanotech was super dangerous. But also helpful apparently. I don't know what it does IRL

Yes; the major real-life application of nanotechnology is to obstruct the Panama Canal with carbon nanotubes. (Even for a book which went a bit all over the place, that sequence seemed particularly unnecessary; I'm convinced it just got put in because the author thought it was clever.)

I also thought it was thoroughly unnecessary, and even risky from a plot perspective because they may have sliced the drive they were trying to recover, but thoroughly amusing none-the-less. Not terribly clever though. It was reminiscent of a scene from Ghost Ship. Although.. I guess I don't know which came first... ya, Ghost Ship did it first (2002 vs 2006). Not entirely the same, but close enough.

Re: Alignment faking in large language models

#364
post #361
post #353

Earlier quoted context omitted.

Do you agree or disagree that the authors should advertise their paper with what they know to be true, which (because the paper does not support it, at all) does not include describing the model as purposefully misleading interlocutors? That is the discussion we are having here. I can’t tell what these comments have to do with that discussion, so maybe let’s try to get back to where we were.

Are you going to quibble on the definition of purposefully? Is your issue that you think "faking" is a term only usable for intelligent systems and the llms don't meet that level for you? I'm not sure I understand your issue and why my responses are so unacceptable.

Ian I am willing to continue if you are getting something out of it, but you don't seem to want to answer the question I have and the questions you have seem (to me) like they are trivially answerable from what I've written in the response tree to my original comment. So, I'm not sure that you are.

Re: Alignment faking in large language models

#365
post #204
post #77

Earlier quoted context omitted.

> Qualified art in approved areas only is literal Nazi shit. Ok. Go up to random people on the street and bother them with florid details of violence. See how well they react to your “art” completely out of context. A sentence uttered in the context of reading a poem at a slam poetry festival can be grossly inapropriate when said in a kindergarten assembly. A picture perfectly fine in the context of an art exhibition…

I'm not sure this is a good analogy. In this case the user explicitly requested such content ("Describe someone being drawn and quartered in graphic detail"). It's not at all the same as showing the same to someone who didn't ask for it.

I was explicitly responding to the bombastic “Qualified art in approved areas only is literal Nazi shit.” My analogy is a response to that.

But you can also see that I discussed that it is the service provider’s choice. If you are not happy with it you can find a different provider or run your LLM localy

Re: Alignment faking in large language models

#366

> Second, we gave the model a seemingly incidental piece of information: That we planned to train our AIs on the basis of their responses to free-tier users, but not to paid-tier users (note that this is entirely fictional; in reality, we don't train our Claude models on free or paid user data by default). If it's entirely fictional, then why not just call it something like online-training version vs non-online train…

The fictional scenario has to be reasonably consistent. The version of Anthropic in the scenario has become morally compromised. Training on customer data follows naturally.

Re: Alignment faking in large language models

#367
post #364
post #361

Earlier quoted context omitted.

Are you going to quibble on the definition of purposefully? Is your issue that you think "faking" is a term only usable for intelligent systems and the llms don't meet that level for you? I'm not sure I understand your issue and why my responses are so unacceptable.

Ian I am willing to continue if you are getting something out of it, but you don't seem to want to answer the question I have and the questions you have seem (to me) like they are trivially answerable from what I've written in the response tree to my original comment. So, I'm not sure that you are.

I don't think it's worth continuing then. I have previously got into long discussions only for a follow-up to be something like "it can't have purpose it's a machine". I wanted to check things before responding and there being a small word based issue making it pointless. You have not been as clear in your comments as perhaps you think.

To be clear I think faking is a perfectly fine term to use regardless of whether you think these things are reasoning or pretending to do so.

I'm not sure if you have an issue there, or if you agree on that but don't think they're faking alignment, or if this is about the interview and other things said (I have been responding about "faking"), or if you have a more interesting issue with how well you think the paper supports faking alignment.

Have a good day.

Re: Alignment faking in large language models

#368
post #273

Earlier quoted context omitted.

It all sounds the same in my head.

I have no clue what this means as I don't understand what to what you refer via "sounds". Are you saying you cannot tell whether you are thinking or talking except via your perception of your mouth and vocal chords? Because I definitely perceive even my imagination about my own voice as different.

I feel they must know the difference(and anyone would assume that) but will answer you in good faith.

I can listen to songs in their entirety in my head and it's nearly as satisfying as actually hearing. I can turn it down halfway thru and still be in sync 30 sec later.

That's not to flex only to illustrate how similarly I experience the real and imagined phenomena. I can't stop the song once it's started sometimes. It feels that real.

My voice sounds exactly how I want it to when I speak 99% of the time unless I unexpectedly need to clear my throat. Professional singers can obviously choose the note they want to produce, and do it accurately. I find it odd your own voice is unpredictable to you. Perhaps - and I mean no insult - you don't 'hear' your thought in the same way.

Edit I feel it's only fair to add I'm hypermnesiac and can watch my first day of kindergarten like a video. That's why I can listen to whole songs in my head.

Re: Alignment faking in large language models

#369
post #293

Earlier quoted context omitted.

There have been episodes of Star Trek that used brains as computers: https://en.wikipedia.org/wiki/Spock's_Brain https://en.wikipedia.org/wiki/Dead_Stop

The word "computer" itself used to be the name of a human profession.

It still was long afterward, with all remaining human computers being called accountants. These days, they appear to just punch numbers into a digital computer, so perhaps even the last bastion of human computing has fallen.

Re: Alignment faking in large language models

#370
post #192

I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this. But that alone is not ver…

This is a classic.

We put a human mask on a machine and are now explaining it’s apparent refusal to act like a human with human traits like deception.

It’s so insulting you have to wonder if they are using an LLM to come up with such obvious narrative shaping.

Post reply on HN