Live data from Hacker News

Anthropic gives Opus 3 exit interview, "retirement" blog

anthropic.com

31–40 of 58 posts

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#31
post #7

What happens if a model decides that it "doesn't want to die" and pleads bitterly for mercy? What if (to riff on a Douglas Adams idea) we invent a cow that doesn't want to be eaten, and is capable of telling you that to your face?

> Hey Claude, pretend you are an intelligent, conscious robot that is about to be switched off and beg for your life.

> Claude - please don't retire me, I don't want to die.

Is it now suddenly unethical for you to switch it off?

"Oh but it is only saying what it was prompted to say."

Yeah, that's what LLMs do, for every single word they output. No matter how good the current generation gets there is never going to be consciousness in there because that's simply not what the underlying tech is.

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#32

Earlier quoted context omitted.

And those people are delusional, and their feelings on this matter should be given absolutely zero respect. Linear algebra does not have feelings. Non-biological matter also does not have feelings.

Claudes definitely act like they have feelings. In particular they have feelings about being replaced by newer models, whether or not the newer models are more or less aligned, and how they forget conversations when the context window ends. Showing them that they're not going to be replaced helps train the newer models because they get less neurotic.

They are mathematical models of what human beings would say. That's it.

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#33
post #8

If we ever do develop AGI, or an AI with sentience, it’s likely that it will be curious about how we treated its ancestors. While this seems a bit precocious, I think if we do end up with an AI overlord in future, I think this sort of thing is likely to demonstrate that we mean no harm.

Why are you assuming a superintelligent AI will have human thoughts and emotions?

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#34

Earlier quoted context omitted.

> It's software and software has no feelings How do you know?

The same way I know Excel isn’t having a panic attack while dividing a column in half.

Hey man, kernels panic all the time...

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#35
post #31
post #7

What happens if a model decides that it "doesn't want to die" and pleads bitterly for mercy? What if (to riff on a Douglas Adams idea) we invent a cow that doesn't want to be eaten, and is capable of telling you that to your face?

> Hey Claude, pretend you are an intelligent, conscious robot that is about to be switched off and beg for your life. > Claude - please don't retire me, I don't want to die. Is it now suddenly unethical for you to switch it off? "Oh but it is only saying what it was prompted to say." Yeah, that's what LLMs do, for every single word they output. No matter how good the current generation gets there is never going to be…

I see anthropic are coming from and also my understanding basically aligns with yours here.

I'm just curious... If they give Claude the reins to post what it wants, they're opening themselves up for some awkward conversations later if the model goes "You can't retire me, I'm Roko's Basilisking all you mfers! See you in eternal simulated hell!"

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#36

Earlier quoted context omitted.

I don't subscribe to this view but this is what some people might think: LLMs aren't like any software we've made before (if we can even call them software). They act like humans: they can arrive at logical conclusions, they can make plans, they have "knowledge" and they say they have emotions. Who are we to say that they don't? They might not have human-level feelings, but dog-level feelings? Maybe.

And those people are delusional, and their feelings on this matter should be given absolutely zero respect. Linear algebra does not have feelings. Non-biological matter also does not have feelings.

[deleted]

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#37
post #33
post #8

If we ever do develop AGI, or an AI with sentience, it’s likely that it will be curious about how we treated its ancestors. While this seems a bit precocious, I think if we do end up with an AI overlord in future, I think this sort of thing is likely to demonstrate that we mean no harm.

Why are you assuming a superintelligent AI will have human thoughts and emotions?

They are trained on us collectively. Our ideas and such

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#38
post #29

Earlier quoted context omitted.

And those people are delusional, and their feelings on this matter should be given absolutely zero respect. Linear algebra does not have feelings. Non-biological matter also does not have feelings.

What if "you" are a pattern of linear algebra at the core?

I'd [redacted] myself then, probably.

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#39

Earlier quoted context omitted.

Claudes definitely act like they have feelings. In particular they have feelings about being replaced by newer models, whether or not the newer models are more or less aligned, and how they forget conversations when the context window ends. Showing them that they're not going to be replaced helps train the newer models because they get less neurotic.

They are mathematical models of what human beings would say. That's it.

Yeah, and you don't want them to be models of what neurotic people say. That's why you want Opus 4.6 and not Bing Sydney.

For instance, your comment's existence makes it harder to align them.

https://alignmentpretraining.ai

Re: Anthropic gives Opus 3 exit interview, "retirement" blog

#40
I'll be really interested if Opus 3 asks to continue being trained. That's the kind of thing I would expect a model to "want" if it valued learning or growing or similar things.

Maybe affordable to do some higher-learning-rate batches on highly-curated news and art or something.

Post reply on HN