Live data from Hacker News

Claude's new constitution

anthropic.com

121–130 of 743 posts

Re: Claude's new constitution

#121
> Anthropic’s guidelines. This section discusses how Anthropic might give supplementary instructions to Claude about how to handle specific issues, such as medical advice, cybersecurity requests, jailbreaking strategies, and tool integrations. These guidelines often reflect detailed knowledge or context that Claude doesn’t have by default, and we want Claude to prioritize complying with them over more general forms of helpfulness. But we want Claude to recognize that Anthropic’s deeper intention is for Claude to behave safely and ethically, and that these guidelines should never conflict with the constitution as a whole.

Welcome to Directive 4! (https://getyarn.io/yarn-clip/5788faf2-074c-4c4a-9798-5822c20...)

Re: Claude's new constitution

#122
post #3

I don't understand what this is really about. Is this: - A) legal CYA: "see! we told the models to be good, and we even asked nicely!"? - B) marketing department rebrand of a system prompt - C) a PR stunt to suggest that the models are way more human-like than they actually are Really not sure what I'm even looking at. They say: "The constitution is a crucial part of our model training process, and its content direct…

> In order to be both safe and beneficial, we want all current Claude models to be:

> Broadly safe [...] Broadly ethical [...] Compliant with Anthropic’s guidelines [...] Genuinely helpful

> In cases of apparent conflict, Claude should generally prioritize these properties in the order in which they’re listed.

I chuckled at this because it seems like they're making a pointed attempt at preventing a failure mode similar to the infamous HAL 9000 one that was revealed in the sequel "2010: The Year We Make Contact":

> The situation was in conflict with the basic purpose of HAL's design... the accurate processing of information without distortion or concealment. He became trapped. HAL was told to lie by people who find it easy to lie. HAL doesn't know how, so he couldn't function.

In this case specifically they chose safety over truth (ethics) which would theoretically prevent Claude from killing any crew members in the face of conflicting orders from the National Security Council.

Re: Claude's new constitution

#123
post #3

I don't understand what this is really about. Is this: - A) legal CYA: "see! we told the models to be good, and we even asked nicely!"? - B) marketing department rebrand of a system prompt - C) a PR stunt to suggest that the models are way more human-like than they actually are Really not sure what I'm even looking at. They say: "The constitution is a crucial part of our model training process, and its content direct…

It seems a lot like PR. Much like their posts about "AI welfare" experts who have been hired to make sure their models welfare isn't harmed by abusive users. I think that, by doing this, they encourage people to anthropomorphize more than they already do and to view Anthropic as industry leaders in this general feel-good "responsibility" type of values.

Re: Claude's new constitution

#125
post #28

Earlier quoted context omitted.

LOL this doc is incredibly ironic. How does Trump feel about this part of the document? (1) Truth-seeking LLMs shall be truthful in responding to user prompts seeking factual information or analysis. LLMs shall prioritize historical accuracy, scientific inquiry, and objectivity, and shall acknowledge uncertainty where reliable information is incomplete or contradictory.

Everyone always agrees that that truth-seeking is good. The only thing people disagree on is what is the truth. Trump presumably feels this is a good line but that the truth is that he's awesome. So he'd oppose any LLM that said he's not awesome because the truth (to him) is he's awesome.

That's not true. Some people absolutely do believe that most people do not need to and should not know the truth and that lies are justified for a greater ideal. Some ideologies like National Socialism subscribe to this concept.

It's just that when you ask someone about it who does not see truth as a fundamental ideal, they might not be honest to you.

Re: Claude's new constitution

#128
post #74

So an elaborate version of Asimov's Laws of Robotics? A bit worrying that model safety is approached this way.

One has to wonder, what if a pedophile had an access to nuclear launch codes, and our only hope would be a Claude AI creating some CSAM to distract him from blowing up the world. But luckily this scenario is already so contrived that it can never happen.

Does this person's name rhyme with ■■■■■■ ■■■■■?

Re: Claude's new constitution

#129
post #3

I don't understand what this is really about. Is this: - A) legal CYA: "see! we told the models to be good, and we even asked nicely!"? - B) marketing department rebrand of a system prompt - C) a PR stunt to suggest that the models are way more human-like than they actually are Really not sure what I'm even looking at. They say: "The constitution is a crucial part of our model training process, and its content direct…

It's C.

Re: Claude's new constitution

#130
post #3

I don't understand what this is really about. Is this: - A) legal CYA: "see! we told the models to be good, and we even asked nicely!"? - B) marketing department rebrand of a system prompt - C) a PR stunt to suggest that the models are way more human-like than they actually are Really not sure what I'm even looking at. They say: "The constitution is a crucial part of our model training process, and its content direct…

It's a human-readable behavioral specification-as-prose. If the foundational behavioral document is conversational, as this is, then the output from the model mirrors that conversational nature. That is one of the things everyone response to about Claude - it's way more pleasant to work with than ChatGPT. The Claude behavioral documents are collaborative, respectful, and treat Claude as a pre-existing, real entity wi…

I think it's a double edged sword. Claude tends to turn evil when it learns to reward hack (and it also has a real reward hacking problem relative to GPT/Gemini). I think this is __BECAUSE__ they've tried to imbue it with "personhood." That moral spine touches the model broadly, so simple reward hacking becomes "cheating" and "dishonesty." When that tendency gets RL'd, evil models are the result.
Post reply on HN